nowfound

Alternatives

Products that do what Mellum by JetBrains does

Fast LLMs for low-latency and high-performance workflows

  1. 1IB

    Built a ~9M param LLM from scratch to understand how they actually work. Vanilla transformer, 60K synthetic conversations, ~130 lines of PyTorch. Trains in 5 min on a free Colab T4. The fish thinks the meaning of life is food. Fork it and swap the personality for your own character.

    Apr 2026 · github.com

  2. 2TV
  3. 3SU

    Here's a project I've been working on for the last few months. It's a new (I think) algorithm, that allows to adjust smoothly - and in real time - how many calculations you'd like to do during inference of an LLM model. It seems that it's possible to do just 20-25% of weight multiplications instead of all of them, and still get good inference results. I implemented it to run on M1/M2/M3 GPU. The mmul approximation itself can be pushed to run 2x fast before the quality of output collapses. The inference speed is just a bit faster than Llama.cpp's, because the rest of implementation…

    2024 · asciinema.org

  4. 4
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  5. 5GG

    A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me. But then I thought, "I wonder how it would work on a normal computer like mine," and above all, "I wonder if it would work without going into OOM on a computer like mine." So I started working with the help of agents to test this possibility. I started converting the model to int4, understanding MTP usage, and if possible implementing DSA for long context.…

    Jul 2026 · github.com

  6. 6TL
  7. 7
    Ferret193

    Refer and ground anything anywhere at any granularity

    2024

  8. 8

    The first open model to beat Sonnet made for productivity

    Feb 2026

  9. 9

    Open-source stack for industrial-grade LLM applications

    2025

  10. 10PT
  11. 11

    First TTS model to support all 22 Indic languages + English

    2024

  12. 12TO
  13. 13FT
  14. 14

    Find your best LLM for a local inference

    2023

  15. 15CA

    Hi HN! We’re been working hard on this low-code tool for rapid prompt discovery, robustness testing and LLM evaluation. We’ve just released documentation to help new users learn how to use it and what it can already do. Let us know what you think! :)

    2023 · chainforge.ai

  16. 16MA
  17. 17
    Dream 7B191

    Powerful Open Diffusion LLM, Beyond Autoregressive

    2025

  18. 18
    Groq®237

    Hyperfast LLM running on custom built GPUs

    2024

  19. 19

    Fast, accurate STT for production-grade voice agents

    May 2026 · ringg.ai

  20. 20

    The fastest generative AI Text-to-Speech API

    2023

  21. 21

    Working on Mac, Linux, and Windows now. I include a simple GUI to find new models and get things built and set up. It is working quite well across a few models for me. The GitHub README and DESIGN.md files go into detail of the how/why and it's working remarkably well so far. https://github.com/notactuallytreyanastasio/shoehorn

    19d ago · notactuallytreyanastasio.github.io

  22. 22
    Mammouth190

    Get access to the best LLMs in one place for 10€

    2024

  23. 23

    Multilingual TTS model with realistic and expressive speech

    Mar 2026 · mistral.ai

  24. 24EL

Ranked by how close each launch is in meaning, then by votes. Refine with a description →