nowfound

Alternatives

Products that do what Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone does

  1. 1

    Massive local model speedup on Apple Silicon with MLX

    Apr 2026 · ollama.com

  2. 2

    Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone. - leonickson1/Swiftlet

    Aug 2026 · github.com

  3. 3KT

    Kitten TTS is an open-source series of tiny and expressive text-to-speech models for on-device applications. We are excited to launch a preview of our smallest model, which is less than 25 MB. This model has 15M parameters. This release supports English text-to-speech applications in eight voices: four male and four female. The model is quantized to int8 + fp16, and it uses onnx for runtime. The model is designed to run literally anywhere eg. raspberry pi, low-end smartphones, wearables, browsers etc. No GPU required! We're releasing this to give early users a sense of the latency and voices…

    2025 · github.com

  4. 4

    First TTS model to support all 22 Indic languages + English

    2024

  5. 5

    Ultra-fast 309B MoE model for coding & agents

    Dec 2025 · mimo.xiaomi.com

  6. 6

    Voice-first writing—now on iPhone

    2025

  7. 7

    The low-code platform for testing AI apps

    2024

  8. 8

    How small can a language model be while still doing something useful? I wanted to find out, and had some spare time over the holidays. Z80-μLM is a character-level language model with 2-bit quantized weights ({-2,-1,0,+1}) that runs on a Z80 with 64KB RAM. The entire thing: inference, weights, chat UI, it all fits in a 40KB .COM file that you can run in a CP/M emulator and hopefully even real hardware! It won't write your emails, but it can be trained to play a stripped down version of 20 Questions, and is sometimes able to maintain the illusion of having simple but terse conversations…

    Dec 2025 · github.com

  9. 9

    The open sparse MoE model for agentic coding

    Apr 2026 · qwen.ai

  10. 10

    The open-source era of 1M context intelligence

    Apr 2026 · huggingface.co

  11. 11

    Write 5x faster with voice dictation now Available on iOS

    Nov 2025

  12. 12

    Text-to-Speech built for Voice Agents

    Apr 2026 · smallest.ai

  13. 13TN

    Kitten TTS (https:&#x2F;&#x2F;github.com&#x2F;KittenML&#x2F;KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=44807868. Today we're releasing three new models with 80M, 40M and 14M parameters. The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in…

    Mar 2026 · github.com

  14. 14RQ
  15. 15
    Wan 2.2208

    The first open MoE model for AI video generation

    2025

  16. 16
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026 · qwen.ai

  17. 17

    0.8B-9B native multimodal w/ more intelligence, less compute

    Mar 2026 · huggingface.co

  18. 18WM

    We wrote our inference engine on Rust, it is faster than llama cpp in all of the use cases. Your feedback is very welcomed. Written from scratch with idea that you can add support of any kernel and platform.

    2025 · github.com

  19. 19

    Blazing fast keyboard for iOS now with GIFs and more

    2016

  20. 20

    A native omni model for voice, video, and tools

    Mar 2026 · qwen.ai

  21. 21

    The most expressive Text to Speech model ever

    2025

  22. 22

    Metal-first MoE inference for Apple Silicon with bounded SSD expert streaming and local OpenAI Chat, Responses, and Anthropic Messages endpoints. - hebrus-labs/hebrus

    Jul 2026 · github.com

  23. 23

    The open-weight preview of Qwen4

    11d ago · qwen.ai

  24. 24EL

Ranked by how close each launch is in meaning, then by votes. Refine with a description →