Alternatives
Products that do what Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone does
- 1

- 2

Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone. - leonickson1/Swiftlet
Aug 2026 · github.com
- 3KT
Kitten TTS is an open-source series of tiny and expressive text-to-speech models for on-device applications. We are excited to launch a preview of our smallest model, which is less than 25 MB. This model has 15M parameters. This release supports English text-to-speech applications in eight voices: four male and four female. The model is quantized to int8 + fp16, and it uses onnx for runtime. The model is designed to run literally anywhere eg. raspberry pi, low-end smartphones, wearables, browsers etc. No GPU required! We're releasing this to give early users a sense of the latency and voices…
2025 · github.com
- 4

First TTS model to support all 22 Indic languages + English
2024
- 5

- 6

- 7

- 8ZΜ
How small can a language model be while still doing something useful? I wanted to find out, and had some spare time over the holidays. Z80-μLM is a character-level language model with 2-bit quantized weights ({-2,-1,0,+1}) that runs on a Z80 with 64KB RAM. The entire thing: inference, weights, chat UI, it all fits in a 40KB .COM file that you can run in a CP/M emulator and hopefully even real hardware! It won't write your emails, but it can be trained to play a stripped down version of 20 Questions, and is sometimes able to maintain the illusion of having simple but terse conversations…
Dec 2025 · github.com
- 9

- 10

- 11

- 12

- 13TN
Kitten TTS (https://github.com/KittenML/KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https://news.ycombinator.com/item?id=44807868. Today we're releasing three new models with 80M, 40M and 14M parameters. The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in…
Mar 2026 · github.com
- 14RQ
Sep 2025 · github.com
- 15

- 16

- 17
Qwen3.5 Small▲3590.8B-9B native multimodal w/ more intelligence, less compute
Mar 2026 · huggingface.co
- 18WM
We wrote our inference engine on Rust, it is faster than llama cpp in all of the use cases. Your feedback is very welcomed. Written from scratch with idea that you can add support of any kernel and platform.
2025 · github.com
- 19

- 20

- 21

- 22

Metal-first MoE inference for Apple Silicon with bounded SSD expert streaming and local OpenAI Chat, Responses, and Anthropic Messages endpoints. - hebrus-labs/hebrus
Jul 2026 · github.com
- 23

- 24EL
2023 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →