nowfound

Alternatives

Products that do what Gemma 4 Multimodal Fine-Tuner for Apple Silicon does

About six months ago, I started working on a project to fine-tune Whisper locally on my M2 Ultra Mac Studio with a limited compute budget. I got into it. The problem I had at the time was I had 15,000 hours of audio data in Google Cloud Storage, and there was no way I could fit all the audio onto my local machine, so I built a system to stream data from my GCS to my machine during training. Gemma 3n came out, so I added that. Kinda went nuts, tbh. Then I put it on the shelf. When Gemma 4 came out a few days ago, I dusted it off, cleaned it up, broke out the Gemma part from the Whisper…

  1. 1

    Run multimodal AI locally with an encoder-free architecture

    Jun 2026 · blog.google

  2. 2OS

    Hi HN, I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal. I have always adored on-device AI. It feels like magic that you can run a powerful NN on your Mac or iPhone. So I wanted to push the limits a bit and run a model whose weights don’t fit in memory. The model’s 4-bit quantized weights occupy roughly 14 GB, which makes running it with conventional inference tools almost impossible on an 8 GB or even 16 GB Mac once the OS, applications, and KV cache are…

    Jul 2026 · github.com

  3. 3
    Gemma 3n199

    Run powerful multimodal AI right on your phone

    2025

  4. 4CH

    Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score between 0 and 1. Developers can accept the on-device when it's high, hand off to a bigger cloud model when it's low. By routing only 15-35% of queries to Gemini 3.1 Flash-Lite, Gemma-4-E2B matches Gemini 3.1 Flash-Lite on most benchmarks. - ChartQA: 15-20% - LibriSpeech: 25-30% - MMBench, GigaSpeech, MMAU: 30-35% -…

    Jul 2026 · github.com

  5. 5TN

    Kitten TTS (https:&#x2F;&#x2F;github.com&#x2F;KittenML&#x2F;KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=44807868. Today we're releasing three new models with 80M, 40M and 14M parameters. The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in…

    Mar 2026 · github.com

  6. 6
    Gemma 3200

    Build with multimodal AI from Google

    2025

  7. 7

    Google's most intelligent open models to date

    Apr 2026

  8. 8

    Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens&#x2F;sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens&#x2F;sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    27d ago · cactuscompute.com

  9. 9

    High quality text transcription with OpenAI's whisper on Mac

    2023

  10. 10FT

    Aug 2026 · github.com

  11. 11GG

    Gemma Gem is a Chrome extension that loads Google's Gemma 4 (2B) through WebGPU in an offscreen document and gives it tools to interact with any webpage: read content, take screenshots, click elements, type text, scroll, and run JavaScript. You get a small chat overlay on every page. Ask it about the page and it (usually) figures out which tools to call. It has a thinking mode that shows chain-of-thought reasoning as it works. It's a 2B model in a browser. It works for simple page questions and running JavaScript, but multi-step tool chains are unreliable and it sometimes ignores its tools…

    Apr 2026 · github.com

  12. 12
    Blankie181

    Open-source ambient sound mixer for macOS

    2025

  13. 13

    Record your microphone, system audio simultaneously

    Apr 2026

  14. 14

    Open-source, local-first dictation you can trust

    2025

  15. 15
    FnKey104

    macOS dictation with Deepgram stream

    Mar 2026

  16. 16IM

    I made my first macOS utility app that ships with a bundled Gemma 4 model, specifically the Gemma E4B one. It made my app DMG have 5.3 GB in size, but I think it is a small size for the power that this free local model can provide. It runs fine on CPU, but can also run on Apple Silicon GPU, although I did not notice any performance improvements with GPU (tested on a M5 chip). I think these local lightweight and multimodal models will open multiple possibilities for new software tools where privacy is essential.

    May 2026 · snapname.app

  17. 17

    Voice transcription lives in the Mac notch

    May 2026 · coddo.ai

  18. 18

    Offline AI Speech to Text Transcription for iOS & macOS

    2025

  19. 19

    Not just dictation and private AI voice toolkit

    Apr 2026

  20. 20
    Thoth 64

    Private, local AI transcription for your Mac

    Apr 2026

  21. 21IM

    Hello all, I made a small transcription app for your Mac based on OpenAI’s Whisper. Would love some feedback. My plan is to make it easy to load weights from any fine-tuned whisper model to enable specialized dictation for any subfield. It’s still early in development. Thanks!

    2023 · twitter.com

  22. 22
    SAND80

    Sequencer and host for audio plugins on iOS

    2022

  23. 23RG

    I wanted to know how fast a 26B mixture-of-experts model could run on a desktop CPU with no GPU. Got ~40 tok&#x2F;s single-stream (lossless) and ~124 batched. The surprising part was the byte budget: for this model you compress the output head (32% of per-token bytes), not the experts (16%). The writeup has the bandwidth roofline and the dead-ends; the repo has the reproducible recipe. Happy to answer questions. Repo: https:&#x2F;&#x2F;github.com&#x2F;arun-prasath2005&#x2F;gemma4-cpu-moe

    Jun 2026 · apeg.dev

  24. 24
    Mimic 377

    Privacy-focused neural text-to-speech (TTS) engine

    2022

Ranked by how close each launch is in meaning, then by votes. Refine with a description →