nowfound

Alternatives

Products that do what Samosa Chat - Run Qwen3.6-35B-A3B Locally on a 16 GB Mac does

  1. 1RA
  2. 2

    Qwen Chat are now available for macOS

    2025

  3. 3Q6
  4. 4RQ
  5. 5VV
  6. 6

    Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API. - carloslfu/slotstream

    5d ago · github.com

  7. 7PM

    2021 · apps.apple.com

  8. 8MA
  9. 9IM
  10. 10AM

    2024 · github.com

  11. 11

    An open-source macOS client for Codex

    Mar 2026

  12. 12AQ
  13. 13PG

    2015 · github.com

  14. 14C2
  15. 15LM

    2021 · jonathanalland.com

  16. 16TF

    I’d originally launched my app: Private LLM[1][2] on HN around 10 months ago, with a single RedPajama Chat 3B model. The app has come a long way since then. About a month ago, I added support for 4-bit OmniQuant quantized Mixtral 8x7B Instruct model, and it seems to outperform Q4 models at inference speed and Q8 models at text generation quality, while consuming only about 24GB of RAM[3] at 8k context length. The trick is: a) to use a better quantization algorithm and b) to use unquantized embeddings and the MoE gates (the overhead is quite small). Other notable features include many more…

    2024

  17. 17IF
  18. 18MW

    I always find myself frustrated by how many steps I have to take to video chat with someone online. There's always too much software to download and install, and too many accounts to remember. The high-end, high-price Cisco conferencing systems I've used are especially fragile. So I built Vidless in a weekend for myself, and I hope you find it useful too. Just create a room and share the link, and you'll be chatting in seconds. You can invite several people (I haven't really tested a max yet), and chats on Vidless are always private. Enjoy!

    2012 · vidless.com

  19. 19RM
  20. 20FP
  21. 21IM

    I made my first macOS utility app that ships with a bundled Gemma 4 model, specifically the Gemma E4B one. It made my app DMG have 5.3 GB in size, but I think it is a small size for the power that this free local model can provide. It runs fine on CPU, but can also run on Apple Silicon GPU, although I did not notice any performance improvements with GPU (tested on a M5 chip). I think these local lightweight and multimodal models will open multiple possibilities for new software tools where privacy is essential.

    May 2026 · snapname.app

  22. 22IB

    hey hn, I built an open-source Perplexity clone that can run local LLMs and cloud LLMs. It's fully self-hostable through Docker and uses ollama to support local LLMs. The demo video in the repository shows me running it locally with llama3 on my M1 Macbook Pro. I'm open to any suggestions or feedback, thanks!

    2024 · github.com

  23. 23NI

    2022 · github.com

  24. 24AC

    2016 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →