Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
In plain words
Maple-Preview is a large language model that runs directly on iPhones, delivering inference speeds of 120 tokens per second. Built on a ternary 20 billion parameter mixture-of-experts architecture, it enables users to run advanced AI locally without relying on cloud services. The software is designed for iPhone users seeking fast, on-device language processing capabilities.
written from the facts on this page · September 2026
Does a similar job
all alternatives →
Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhoneAug 2026 · github.com · ▲312Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone. - leonickson1/Swiftlet
Gan.AI TTS Model & API Playground2024 · ▲344First TTS model to support all 22 Indic languages + English



More life & fun this month
the category →- TL
Life & fun · 11d ago · louisabraham.github.io



Photosynthesis fires two of your iPhone
Life & fun · 29d ago · photosynthesis.camera
- CCCreatium Coach▲319
Your multimedia mentor that takes you from mid to great
Life & fun · 11d ago · producthunt.creatium.info
SoloUno▲310Take control of hair pulling, nail biting & skin picking
Life & fun · 29d ago · solouno.io
Launched alongside, August 2026
the whole month →- TL
Life & fun · 11d ago · louisabraham.github.io



Hello HN! I found that picking out plausible but diverse skin tones for my digital art and game development projects was kind of difficult, and I got curious about if there was a way to define a color space that made it easy. I've built a color picker and procedural generation algorithm based on the space as well as a bunch of other fun js features and demos throughout the page that use the equations. If you find it interesting, I have lots of explanations of how I built it and what properties the space has. The methodology might be a bit shaky, but hopefully the result is as helpful for…
Life & fun · Aug 2026 · toneyalexander.github.io


I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 18d ago · simedw.com