nowfound

Alternatives

Products that do what Qelm does

More Meaning per Qubit, Run Quantum‑Inspired NLP on CPU/GPU.

  1. 1IB

    Built a ~9M param LLM from scratch to understand how they actually work. Vanilla transformer, 60K synthetic conversations, ~130 lines of PyTorch. Trains in 5 min on a free Colab T4. The fish thinks the meaning of life is food. Fork it and swap the personality for your own character.

    Apr 2026 · github.com

  2. 2TV
  3. 3

    How small can a language model be while still doing something useful? I wanted to find out, and had some spare time over the holidays. Z80-μLM is a character-level language model with 2-bit quantized weights ({-2,-1,0,+1}) that runs on a Z80 with 64KB RAM. The entire thing: inference, weights, chat UI, it all fits in a 40KB .COM file that you can run in a CP/M emulator and hopefully even real hardware! It won't write your emails, but it can be trained to play a stripped down version of 20 Questions, and is sometimes able to maintain the illusion of having simple but terse conversations…

    Dec 2025 · github.com

  4. 4

    A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me. But then I thought, "I wonder how it would work on a normal computer like mine," and above all, "I wonder if it would work without going into OOM on a computer like mine." So I started working with the help of agents to test this possibility. I started converting the model to int4, understanding MTP usage, and if possible implementing DSA for long context.…

    Jul 2026 · github.com

  5. 5WM

    Try it out! https://glhf.chat/ Hey HN! We’ve been working for the past few months on a website to let you easily run (almost) any open-source LLM on autoscaling GPU clusters. It’s free for now while we figure out how to price it, but we expect to be cheaper than most GPU offerings since we can run the models multi-tenant. Unlike Together AI, Fireworks, etc, we’ll run any model that the open-source vLLM project supports: we don’t have a hardcoded list. If you want a specific model or finetune, you don’t have to ask us for it: you can just paste the Hugging Face link in and…

    2024 · glhf.chat

  6. 6

    0.8B-9B native multimodal w/ more intelligence, less compute

    Mar 2026 · huggingface.co

  7. 7
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026 · qwen.ai

  8. 81B
  9. 9
    QwQ-32B197

    Matching R1 reasoning yet 20x smaller

    2025

  10. 10
    QWQ-Max126

    New LLM by Alibaba excelling in reasoning w/ "thinking mode"

    2025

  11. 11IV

    The video demo runs a 7b Model on a normal gaming GPU. I think it already works quite well (accounting for the limited hardware power). :)

    2024 · github.com

  12. 12

    New LLM compression algorithm by Google

    Mar 2026 · research.google

  13. 13AQ

    2019 · qml.entropicalabs.io

  14. 14
    Qwen3149

    Think Deeper or Act Faster

    2025

  15. 15AT

    A 3.16M-parameter INT4 transformer running entirely in the on-chip memory of a Xilinx Kria KV260. Zero DRAM in the token loop, 59,965 tok/s on the fabric, bit-exact. Chat with it live.

    28d ago · mikeayles.com

  16. 16

    Qwen’s most capable model for coding and cowork

    Aug 2026 · qwen.ai

  17. 17FI
  18. 18
    Mercury 2152

    Fastest reasoning LLM built for instant production AI

    Feb 2026 · inceptionlabs.ai

  19. 19
    Groq®237

    Hyperfast LLM running on custom built GPUs

    2024

  20. 20TL
  21. 21

    The sweet-spot open dense model for coding agents

    Apr 2026 · qwen.ai

  22. 22

    Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU. - MakazhanAlpamys/Soup

    Aug 2026 · github.com

  23. 23

    Large language model series developed by Alibaba Cloud

    2025

  24. 24

    The world’s most powerful chip’ for AI

    2024

Ranked by how close each launch is in meaning, then by votes. Refine with a description →