nowfound

Alternatives

Products that do what Mellum by JetBrains does

Fast LLMs for low-latency and high-performance workflows

  1. 1
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  2. 2
    Ferret193

    Refer and ground anything anywhere at any granularity

    2024

  3. 3

    The first open model to beat Sonnet made for productivity

    Feb 2026

  4. 4

    Find your best LLM for a local inference

    2023

  5. 5
    Dream 7B191

    Powerful Open Diffusion LLM, Beyond Autoregressive

    2025

  6. 6

    Fast, accurate STT for production-grade voice agents

    May 2026

  7. 7

    The fastest generative AI Text-to-Speech API

    2023

  8. 8
    Tiny Aya211

    Local, open-weight AI designed for real-world languages

    Apr 2026

  9. 9

    Multilingual TTS model with realistic and expressive speech

    Mar 2026

  10. 10

    Train and run LLMs on your device

    2025

  11. 11
    VELS187

    Voice-enabled learning simulations for learning and training

    2024

  12. 12

    Unlock your knowledge with 2000 LLM prompts

    2023

  13. 13
    Taylor AI118

    Fine-tune open source LLMs in minutes

    2023

  14. 14

    Go from dataset to custom RAG prototype in 5 minutes

    2024

  15. 15

    From English prompt to deployed ML model with human approval

    Jun 2026

  16. 16

    Fine-tuning, RL, and inference in one CLI

    Dec 2025

  17. 17AA

    An all-in-one blog for learning LLM ins and outs: tokenize, attention, PE, and more Project I've been diving deep into the internals of Large Language Models (LLMs) and started documenting my findings. My blog covers topics like: Tokenization techniques (e.g., BBPE) Attention mechanism (e.g. MHA, MQA, MLA) Positional encoding and extrapolation (e.g. RoPE, NTK-aware interpolation, YaRN) Architecture details of models like QWen, LLaMA Training methods including SFT and Reinforcement Learning If you're interested in the nuts and bolts of LLMs, feel free to check it out:…

    2025 · comfyai.app

  18. 18

    The best model for coding

    Jan 2026

  19. 19LI

    2018 · languagemodels.io

  20. 20WB

    Here is a production-first Keras-inspired LM framework, built with the advice of François Chollet (ex-Google, creator of Keras and ARC-AGI), our technical advisor. This system have already been deployed in production with our clients (which is why we have already every LLMOps practice implemented). It is also compatible with Jupyter and Marimo to integrate seamlessly in you Data Scientists workflows. You can try the code examples online on HF space and you can find more information in the documentation and FAQ. If you have any feedback for us don't hesitate to join our discord! More releases…

    2025 · github.com

  21. 21CT

    I had been looking to try <500M parameter language models but you wouldn't find an API to try them anywhere, so I built this cloudflare hosted static website that hosts weights and built an inference runtime for these models that uses WebGPU and runs inference from your browser. These are only so useful in a multi-turn conversation but it's still interesting to see what you can pack in a <250mb model. I tried using ONNX versions earlier, but there were too many quirks of using them with language models and the TPS wasn't too impressive. Inspired by svenflow&#x2F;webgpu-gemma, I put my codex…

    May 2026 · chonklm.com

  22. 22AG

    I’ve been building LLM tooling for a small VC fund and found myself explaining the same mental model over and over to non-technical people around me: how a stateless LLM becomes a chatbot, how tool use works, what an agent is mechanically, and why context windows shape all of it. I never found a guide that covered that full chain at the level I wanted, so I wrote one. It’s nine short chapters, each building on the last. Deliberately simplified: the goal is a useful mental model, not a textbook. Feedback, corrections, and contributions welcome: github.com&#x2F;ymyke&#x2F;aiaiai

    Apr 2026 · aiaiai.guide

  23. 23MM

    Hi HN! We (Thomas and Stéphan, hello!) recently released Model2Vec, a Python library for distilling any sentence transformer into a small set of static embeddings. This makes inference with such a model up to 500x faster, and reduces model size by a factor of 15 (7.5M params or 15&#x2F;30MB on disk, depending on whether you use float16 or float32). This allows you to embed 50-100k documents per second on a cpu on a macbook. This reduction of course comes at a cost: distilled models are worse than their parent models. Even so, they are actually a lot better than large sets of conventional…

    2024 · github.com

  24. 24HP

    Hi HN. I heard you like dev tools and AI, so we wanted to share our project that we’ve been working on. We’re working on Horizon [1] - a higher level abstraction for LLMs so that developers can spend less time trying to grapple with LLMs to make them work and more time with users. This is the starting feature set which takes an auto-ML approach to identify the optimal LLM model, hyperparameters, and prompt - instead of just giving you the tooling to figure it out yourself. You can read more about it in our documentations. Our view is that as LLMs become increasingly commoditized and prompts…

    2023 · gethorizon.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →