nowfound

Alternatives

Products that do what Gemma4 does

Playground, mobile guide & tools for Google's Gemma 4 models

  1. 1

    Google's most intelligent open models to date

    Apr 2026 · blog.google

  2. 2
    Gemma259

    Google’s new state-of-the-art open source LLMs

    2024

  3. 3
    Gemma 2279

    Lightweight, state-of-the-art open models from Google

    2024

  4. 4

    Run multimodal AI locally with an encoder-free architecture

    Jun 2026 · blog.google

  5. 5GG

    Gemma Gem is a Chrome extension that loads Google's Gemma 4 (2B) through WebGPU in an offscreen document and gives it tools to interact with any webpage: read content, take screenshots, click elements, type text, scroll, and run JavaScript. You get a small chat overlay on every page. Ask it about the page and it (usually) figures out which tools to call. It has a thinking mode that shows chain-of-thought reasoning as it works. It's a 2B model in a browser. It works for simple page questions and running JavaScript, but multi-step tool chains are unreliable and it sometimes ignores its tools…

    Apr 2026 · github.com

  6. 6
    Gemma 3n199

    Run powerful multimodal AI right on your phone

    2025

  7. 7
    Gemma 3200

    Build with multimodal AI from Google

    2025

  8. 8OS

    Hi HN, I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal. I have always adored on-device AI. It feels like magic that you can run a powerful NN on your Mac or iPhone. So I wanted to push the limits a bit and run a model whose weights don’t fit in memory. The model’s 4-bit quantized weights occupy roughly 14 GB, which makes running it with conventional inference tools almost impossible on an 8 GB or even 16 GB Mac once the OS, applications, and KV cache are…

    Jul 2026 · github.com

  9. 9G4

    About six months ago, I started working on a project to fine-tune Whisper locally on my M2 Ultra Mac Studio with a limited compute budget. I got into it. The problem I had at the time was I had 15,000 hours of audio data in Google Cloud Storage, and there was no way I could fit all the audio onto my local machine, so I built a system to stream data from my GCS to my machine during training. Gemma 3n came out, so I added that. Kinda went nuts, tbh. Then I put it on the shelf. When Gemma 4 came out a few days ago, I dusted it off, cleaned it up, broke out the Gemma part from the Whisper…

    Apr 2026 · github.com

  10. 10CH

    Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score between 0 and 1. Developers can accept the on-device when it's high, hand off to a bigger cloud model when it's low. By routing only 15-35% of queries to Gemini 3.1 Flash-Lite, Gemma-4-E2B matches Gemini 3.1 Flash-Lite on most benchmarks. - ChartQA: 15-20% - LibriSpeech: 25-30% - MMBench, GigaSpeech, MMAU: 30-35% -…

    Jul 2026 · github.com

  11. 11

    Run leading vision models locally with the new engine

    2025

  12. 12PT
  13. 13OR

    Hi HN A few folks and I have been working on this project for a couple weeks now. After previously working on the Docker project for a number of years (both on the container runtime and image registry side), the recent rise in open source language models made us think something similar needed to exist for large language models too. While not exactly the same as running linux containers, running LLMs shares quite a few of the same challenges. There are "base layers" (e.g. models like Llama 2), specific configuration to run correctly (parameters, temperature, context window sizes etc). There's…

    2023 · github.com

  14. 14CO

    Hey HN, Henry and Roman here - we've been building a cross-platform framework for deploying LLMs, VLMs, Embedding Models and TTS models locally on smartphones. Ollama enables deploying LLMs models locally on laptops and edge severs, Cactus enables deploying on phones. Deploying directly on phones facilitates building AI apps and agents capable of phone use without breaking privacy, supports real-time inference with no latency, we have seen personalised RAG pipelines for users and more. Apple and Google actively went into local AI models recently with the launch of Apple Foundation Frameworks…

    2025 · github.com

  15. 15
    Llama312

    3.1-405B: an open source model to rival GPT-4o / Claude-3.5

    2024

  16. 16
    Llama 4423

    A new era of natively multimodal AI innovation

    2025

  17. 17IM

    I made my first macOS utility app that ships with a bundled Gemma 4 model, specifically the Gemma E4B one. It made my app DMG have 5.3 GB in size, but I think it is a small size for the power that this free local model can provide. It runs fine on CPU, but can also run on Apple Silicon GPU, although I did not notice any performance improvements with GPU (tested on a M5 chip). I think these local lightweight and multimodal models will open multiple possibilities for new software tools where privacy is essential.

    May 2026 · snapname.app

  18. 18
    ChattyUI149

    Run open-source LLMs locally in the browser using WebGPU

    2024

  19. 19
    Ollama235

    The easiest way to run large language models locally

    2023

  20. 20

    Gemma 4 on your phone. 46 AI doctors. Zero data leaks.

    May 2026 · github.com

  21. 21
    Apollo AI280

    Run local models like Llama on iOS

    2025

  22. 22
    Llama 2263

    The next generation of Meta's open source LLM

    2023

  23. 23

    Massive local model speedup on Apple Silicon with MLX

    Apr 2026

  24. 24LL

Ranked by how close each launch is in meaning, then by votes. Refine with a description →