nowfound

Alternatives

Products that do what 1-Bit Bonsai, the First Commercially Viable 1-Bit LLMs does

  1. 1BA

    Introducing Bonsai 0.5B, one of the first ternary-weight LLMs to rival full-precision models of similar size, such as Qwen 2.5 0.5B and MobileLLM 0.5B. Trained on just 3.8B tokens, using 1,000x less data than other models, Bonsai redefines what’s possible for ultra-efficient training in low-bit models. Next, we're building larger and more powerful ternary-weight models for the edge. Technical Report: https://github.com/deepgrove-ai/Bonsai/blob/main/paper/Bonsa... Model (Unpacked): https://huggingface.co/deepgrove/Bonsai Reach us:…

    2025 · github.com

  2. 2IB

    Built a ~9M param LLM from scratch to understand how they actually work. Vanilla transformer, 60K synthetic conversations, ~130 lines of PyTorch. Trains in 5 min on a free Colab T4. The fish thinks the meaning of life is food. Fork it and swap the personality for your own character.

    Apr 2026 · github.com

  3. 3RP

    The excitement surrounding PrismML’s 1-bit/ternary Bonsai models has the industry closely watching how smartphone giants, particularly Apple, will implement LLMs on edge devices. Moving AI on-device is a brilliant and necessary strategy. It ensures absolute user privacy in alignment with EU regulations, fundamentally shifts the economics away from costly cloud inference, and paves the way for a significant hardware upgrade supercycle as users seek true AI-capable silicon. To create a smart on-device "Semantic Router," models need to reach the 27B+ parameter scale. Achieving this on a…

    Jul 2026

  4. 4L3

    I spent a lot of time and money on this rather big side project of mine that attempts to replicate the mechanistic interpretability research on proprietary LLMs that was quite popular this year and produced great research papers by Anthropic [1], OpenAI [2] and Deepmind [3]. I am quite proud of this project and since I consider myself the target audience for HackerNews did I think that maybe some of you would appreciate this open research replication as well. Happy to answer any questions or face any feedback. Cheers [1]…

    2024 · github.com

  5. 5B1

    We took a recently released Bonsai 1.7B ternary model from PrismML (https://github.com/PrismML-Eng/Bonsai-demo) and ran our agentic evolution search on it for 6 hours to optimize the Metal kernels. The search was fully autonomous. Measured against unmodified upstream llama.cpp at the same Bonsai/Q2_0 commit, same M4 Max: - tg128: 309.82 → 442.42 t/s (+42.0%) - pp512: 4250.32 → 4622.63 t/s (+8.8%)

    May 2026 · agents2agents.ai

  6. 6

    0.8B-9B native multimodal w/ more intelligence, less compute

    Mar 2026

  7. 7
    Bonsai105

    AI programming platform for enterprises

    2017

  8. 8BP
  9. 9DB

    I've been doing some data cleaning for my fine tuning projects using LLMs, and decided to just build a package for it as a side project. Check it out here: https://github.com/databonsai/databonsai Some features: - categorization (labelling), transformation and decomposition (text into structured format) - validates llm outputs - batch mode batches up the inputs/outputs so you don't send the prompt (schema, fewshot examples) for every row of data, saving a significant amount of tokens There are some similarities to the Instructor repo, but this is simpler and made for…

    2024 · github.com

  10. 10TV
  11. 11

    Announcing GPT-4.1, GPT-4.1 mini, & GPT-4.1 nano in the API

    2025

  12. 12

    New LLM compression algorithm by Google

    Mar 2026

  13. 13
    QwQ-32B197

    Matching R1 reasoning yet 20x smaller

    2025

  14. 14AT

    A 3.16M-parameter INT4 transformer running entirely in the on-chip memory of a Xilinx Kria KV260. Zero DRAM in the token loop, 59,965 tok/s on the fabric, bit-exact. Chat with it live.

    27d ago · mikeayles.com

  15. 15TL
  16. 16
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  17. 17
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026

  18. 18
    Dream 7B191

    Powerful Open Diffusion LLM, Beyond Autoregressive

    2025

  19. 19

    Qwen’s most capable model for coding and cowork

    Aug 2026 · qwen.ai

  20. 20
    Bonsai88

    Curated tools and expert videos in one beautiful box

    2014

  21. 21KR

    I discovered that in LLM inference, keys and values in the KV cache have very different quantization sensitivities. Keys need higher precision than values to maintain quality. I patched llama.cpp to enable different bit-widths for keys vs. values on Apple Silicon. The results are surprising: - K8V4 (8-bit keys, 4-bit values): 59% memory reduction with only 0.86% perplexity loss - K4V8 (4-bit keys, 8-bit values): 59% memory reduction but 6.06% perplexity loss - The configurations use the same number of bits, but K8V4 is 7× better for quality This means you can run LLMs with 2-3× longer…

    2025 · github.com

  22. 22

    Get your unstructured data AI-ready in minutes

    2024

  23. 23

    The sweet-spot open dense model for coding agents

    Apr 2026

  24. 24

    The first open model to beat Sonnet made for productivity

    Feb 2026

Ranked by how close each launch is in meaning, then by votes. Refine with a description →