nowfound

Alternatives

Products that do what I benchmarked Gemma 4 E2B – the 2B model beat the 12B on multi-turn does

  1. 1

    Run multimodal AI locally with an encoder-free architecture

    Jun 2026 · blog.google

  2. 2

    Google's most intelligent open models to date

    Apr 2026

  3. 3
    Gemma 2279

    Lightweight, state-of-the-art open models from Google

    2024

  4. 4OS

    Hi HN, I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal. I have always adored on-device AI. It feels like magic that you can run a powerful NN on your Mac or iPhone. So I wanted to push the limits a bit and run a model whose weights don’t fit in memory. The model’s 4-bit quantized weights occupy roughly 14 GB, which makes running it with conventional inference tools almost impossible on an 8 GB or even 16 GB Mac once the OS, applications, and KV cache are…

    Jul 2026 · github.com

  5. 5CH

    Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score between 0 and 1. Developers can accept the on-device when it's high, hand off to a bigger cloud model when it's low. By routing only 15-35% of queries to Gemini 3.1 Flash-Lite, Gemma-4-E2B matches Gemini 3.1 Flash-Lite on most benchmarks. - ChartQA: 15-20% - LibriSpeech: 25-30% - MMBench, GigaSpeech, MMAU: 30-35% -…

    Jul 2026 · github.com

  6. 6
    Gemma 3n199

    Run powerful multimodal AI right on your phone

    2025

  7. 7G4

    About six months ago, I started working on a project to fine-tune Whisper locally on my M2 Ultra Mac Studio with a limited compute budget. I got into it. The problem I had at the time was I had 15,000 hours of audio data in Google Cloud Storage, and there was no way I could fit all the audio onto my local machine, so I built a system to stream data from my GCS to my machine during training. Gemma 3n came out, so I added that. Kinda went nuts, tbh. Then I put it on the shelf. When Gemma 4 came out a few days ago, I dusted it off, cleaned it up, broke out the Gemma part from the Whisper…

    Apr 2026 · github.com

  8. 8
    Gemma259

    Google’s new state-of-the-art open source LLMs

    2024

  9. 9
    Gemma 3200

    Build with multimodal AI from Google

    2025

  10. 10GG

    Gemma Gem is a Chrome extension that loads Google's Gemma 4 (2B) through WebGPU in an offscreen document and gives it tools to interact with any webpage: read content, take screenshots, click elements, type text, scroll, and run JavaScript. You get a small chat overlay on every page. Ask it about the page and it (usually) figures out which tools to call. It has a thinking mode that shows chain-of-thought reasoning as it works. It's a 2B model in a browser. It works for simple page questions and running JavaScript, but multi-step tool chains are unreliable and it sometimes ignores its tools…

    Apr 2026 · github.com

  11. 11M4

    2011 · snowday2011.com

  12. 12FA

    Sharing instructions on finetuning Gemma2b for a codegen usecase!

    2024 · github.com

  13. 13

    Fine-tuned Gemma 2: 2B model for Kazakh Instructions (SLLM)

    2025

  14. 14IM

    I made my first macOS utility app that ships with a bundled Gemma 4 model, specifically the Gemma E4B one. It made my app DMG have 5.3 GB in size, but I think it is a small size for the power that this free local model can provide. It runs fine on CPU, but can also run on Apple Silicon GPU, although I did not notice any performance improvements with GPU (tested on a M5 chip). I think these local lightweight and multimodal models will open multiple possibilities for new software tools where privacy is essential.

    May 2026 · snapname.app

  15. 15

    Two wheel smart self balancing scooter hoverboard thingy

    2015

  16. 16PT
  17. 17RG

    I wanted to know how fast a 26B mixture-of-experts model could run on a desktop CPU with no GPU. Got ~40 tok/s single-stream (lossless) and ~124 batched. The surprising part was the byte budget: for this model you compress the output head (32% of per-token bytes), not the experts (16%). The writeup has the bandwidth roofline and the dead-ends; the repo has the reproducible recipe. Happy to answer questions. Repo: https://github.com/arun-prasath2005/gemma4-cpu-moe

    Jun 2026 · apeg.dev

  18. 18IB
  19. 19BM
  20. 20IJ

    2010 · 4sqvision.jazzychad.com

  21. 21MS
  22. 22

    Open-source AI playground for image & video generation

    Apr 2026

  23. 23LL
  24. 24IB

Ranked by how close each launch is in meaning, then by votes. Refine with a description →