nowfound

Alternatives

Products that do what Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules does

The excitement surrounding PrismML’s 1-bit/ternary Bonsai models has the industry closely watching how smartphone giants, particularly Apple, will implement LLMs on edge devices. Moving AI on-device is a brilliant and necessary strategy. It ensures absolute user privacy in alignment with EU regulations, fundamentally shifts the economics away from costly cloud inference, and paves the way for a significant hardware upgrade supercycle as users seek true AI-capable silicon. To create a smart on-device "Semantic Router," models need to reach the 27B+ parameter scale. Achieving this on a…

  1. 11B
  2. 2B1

    We took a recently released Bonsai 1.7B ternary model from PrismML (https://github.com/PrismML-Eng/Bonsai-demo) and ran our agentic evolution search on it for 6 hours to optimize the Metal kernels. The search was fully autonomous. Measured against unmodified upstream llama.cpp at the same Bonsai/Q2_0 commit, same M4 Max: - tg128: 309.82 → 442.42 t/s (+42.0%) - pp512: 4250.32 → 4622.63 t/s (+8.8%)

    May 2026 · agents2agents.ai

  3. 3CO

    Hey HN, Henry and Roman here - we've been building a cross-platform framework for deploying LLMs, VLMs, Embedding Models and TTS models locally on smartphones. Ollama enables deploying LLMs models locally on laptops and edge severs, Cactus enables deploying on phones. Deploying directly on phones facilitates building AI apps and agents capable of phone use without breaking privacy, supports real-time inference with no latency, we have seen personalised RAG pipelines for users and more. Apple and Google actively went into local AI models recently with the launch of Apple Foundation Frameworks…

    2025 · github.com

  4. 4

    Massive local model speedup on Apple Silicon with MLX

    Apr 2026

  5. 5
    Bonsai105

    AI programming platform for enterprises

    2017

  6. 6RA
  7. 7

    Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    27d ago · cactuscompute.com

  8. 8OS

    Hi HN, I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal. I have always adored on-device AI. It feels like magic that you can run a powerful NN on your Mac or iPhone. So I wanted to push the limits a bit and run a model whose weights don’t fit in memory. The model’s 4-bit quantized weights occupy roughly 14 GB, which makes running it with conventional inference tools almost impossible on an 8 GB or even 16 GB Mac once the OS, applications, and KV cache are…

    Jul 2026 · github.com

  9. 9

    Ollama but for mobile, with a cloud fallback

    2025

  10. 108F

    Hi HN! I'm just sharing a project I've been working on during the LLM Efficiency Challenge - you can now finetune Llama with QLoRA 5x faster than Huggingface's original implementation on your own local GPU. Some highlights: 1. Manual autograd engine - hand derived backprop steps. 2. QLoRA / LoRA 80% faster, 50% less memory. 3. All kernels written in OpenAI's Triton language. 4. 0% loss in accuracy - no approximation methods - all exact. 5. No change of hardware necessary. Supports NVIDIA GPUs since 2018+. CUDA 7.5+. 6. Flash Attention support via Xformers. 7. Supports 4bit and 16bit…

    2023 · github.com

  11. 11BA

    Introducing Bonsai 0.5B, one of the first ternary-weight LLMs to rival full-precision models of similar size, such as Qwen 2.5 0.5B and MobileLLM 0.5B. Trained on just 3.8B tokens, using 1,000x less data than other models, Bonsai redefines what’s possible for ultra-efficient training in low-bit models. Next, we're building larger and more powerful ternary-weight models for the edge. Technical Report: https://github.com/deepgrove-ai/Bonsai/blob/main/paper/Bonsa... Model (Unpacked): https://huggingface.co/deepgrove/Bonsai Reach us:…

    2025 · github.com

  12. 12

    The on-device model for your personal data

    Sep 2025

  13. 13WM

    We wrote our inference engine on Rust, it is faster than llama cpp in all of the use cases. Your feedback is very welcomed. Written from scratch with idea that you can add support of any kernel and platform.

    2025 · github.com

  14. 14AT

    A 3.16M-parameter INT4 transformer running entirely in the on-chip memory of a Xilinx Kria KV260. Zero DRAM in the token loop, 59,965 tok/s on the fabric, bit-exact. Chat with it live.

    27d ago · mikeayles.com

  15. 15
    Apollo AI280

    Run local models like Llama on iOS

    2025

  16. 16TV
  17. 17WB

    Hey HN! Alex and Zack from Nexa AI here. We are excited to share a project our team has been passionately working on recently, in collaboration with Jiajun from Meta, Qun from San Francisco State University, and Xin and Qi from the University of North Texas. Running AI models on edge devices is becoming increasingly important. It's cost-effective, ensures privacy, offers low-latency responses, and allows for customization. Plus, it's always available, even offline. What's really exciting is that smaller-scale models are now approaching the performance of large-scale closed-source models for…

    2024 · github.com

  18. 18MP
  19. 19

    Calculate the GPU memory you need for LLM inference

    2025

  20. 20

    Ultra-efficient on-device AI, now even faster

    2025

  21. 21

    Ultra-Fast, Light-as-Air AI Generation

    Jun 2026 · bonsaiimage.com

  22. 22
    NobodyWho106

    Run AI models on any device

    17d ago · github.com

  23. 23

    Run a 4B image model on your phone

    May 2026 · prismml.com

  24. 24
    LFM2.5134

    The next generation of on-device AI

    Jan 2026

Ranked by how close each launch is in meaning, then by votes. Refine with a description →