nowfound

Alternatives

Products that do what ZeroModels does

Open-source Keras 3 collection of pretrained models.

  1. 1DA
  2. 2
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  3. 3TN

    Kitten TTS (https:&#x2F;&#x2F;github.com&#x2F;KittenML&#x2F;KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=44807868. Today we're releasing three new models with 80M, 40M and 14M parameters. The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in…

    Mar 2026 · github.com

  4. 4II

    I invented Discrete Distribution Networks, a novel generative model with simple principles and unique properties, and the paper has been accepted to ICLR2025! Modeling data distribution is challenging; DDN adopts a simple yet fundamentally different approach compared to mainstream generative models (Diffusion, GAN, VAE, autoregressive model): 1. The model generates multiple outputs simultaneously in a single forward pass, rather than just one output. 2. It uses these multiple outputs to approximate the target distribution of the training data. 3. These outputs together represent a discrete…

    Oct 2025 · discrete-distribution-networks.github.io

  5. 5
    Kimi K3448

    The world's first open 3T-class model

    Jul 2026 · kimi.ai

  6. 6
    GLM-5.3254

    Coding leap from scaled post-training on the same base

    22d ago · z.ai

  7. 7
    MARS5 TTS489

    Open-source, insanely prosodic text-to-speech model

    2024

  8. 8

    First TTS model to support all 22 Indic languages + English

    2024

  9. 9BV

    Vision models have been gaining popularity as a replacement for traditional OCR. Especially with Gemini 2.0 becoming cost competitive with the cloud platforms. We've been continuously evaluating different models since we released the Zerox package last year (https:&#x2F;&#x2F;github.com&#x2F;getomni-ai&#x2F;zerox). And we wanted to put some numbers behind it. So we’re open sourcing our internal OCR benchmark + evaluation datasets. Full writeup + data explorer here: https:&#x2F;&#x2F;getomni.ai&#x2F;ocr-benchmark Github: https:&#x2F;&#x2F;github.com&#x2F;getomni-ai&#x2F;benchmark Huggingface:…

    2025 · getomni.ai

  10. 10SO

    Hi HN - Marcello and Vaibhav here. We built smolmodels to experiment with using LLMs for ML development. It's a fully open-source library that generates complete model training and inference code from natural language descriptions. It combines graph search with LLM code generation to find a model that gives as good predictions as possible. The core idea is that LLMs are overkill for a lot of predictive tasks. Smolmodels automates the trial-and-error process of finding the right model architecture and training approach, letting you build small, specialised models. You can either provide your…

    2025 · github.com

  11. 11
    GLM-4.6V239

    Open-source multimodal model with native tool use

    Dec 2025

  12. 12

    Vision-to-code foundation model for real GUI automation

    Apr 2026 · docs.z.ai

  13. 13

    How small can a language model be while still doing something useful? I wanted to find out, and had some spare time over the holidays. Z80-μLM is a character-level language model with 2-bit quantized weights ({-2,-1,0,+1}) that runs on a Z80 with 64KB RAM. The entire thing: inference, weights, chat UI, it all fits in a 40KB .COM file that you can run in a CP&#x2F;M emulator and hopefully even real hardware! It won't write your emails, but it can be trained to play a stripped down version of 20 Questions, and is sometimes able to maintain the illusion of having simple but terse conversations…

    Dec 2025 · github.com

  14. 14IB

    Built a ~9M param LLM from scratch to understand how they actually work. Vanilla transformer, 60K synthetic conversations, ~130 lines of PyTorch. Trains in 5 min on a free Colab T4. The fish thinks the meaning of life is food. Fork it and swap the personality for your own character.

    Apr 2026 · github.com

  15. 15
    Mistral 3415

    A family of frontier open-source multimodal models

    Dec 2025

  16. 16SO
  17. 17BA

    Hello HN! I want to share something me and a few friends have been working on for a while now — Zeroshot, a web tool that builds image classifiers using text-image models and autolabeling. What does this mean in practice? You can put together an image classifier in about 30 seconds that’s faster and more accurate than CLIP, but that you can deploy yourself however you’d like. It’s open source, commercially licensed, and doesn’t require you to pay anyone per API call. Here's a 2 minute video that shows it off: https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=S4R1gtmM-Lo How&#x2F;why does it…

    2023 · usezeroshot.com

  18. 18

    Pre-trained computer vision models and public datasets

    2021

  19. 19

    0.8B-9B native multimodal w/ more intelligence, less compute

    Mar 2026

  20. 20

    Unified video generation for motion design and branding

    Jul 2026 · minimax.io

  21. 21

    MoE vision-language, now easier to access

    2025

  22. 22
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026

  23. 23OP

    Open Prompts is the dataset used to build krea.ai. The data comes from the Stability AI Discord and includes around 10M images from 2M prompts. You can use it for creating semantic search engines of prompts, training LLMs, fine-tuning image-to-text models like BLIP, or extracting insights from the data—like the most common combinations of modifiers.

    2022 · github.com

  24. 24
    Ferret193

    Refer and ground anything anywhere at any granularity

    2024

Ranked by how close each launch is in meaning, then by votes. Refine with a description →