Alternatives
Products that do what Samosa Chat - Run Qwen3.6-35B-A3B Locally on a 16 GB Mac does
- 1RA
Aug 2026 · github.com
- 2

- 3Q6
Jul 2026 · github.com
- 4RQ
Sep 2025 · github.com
- 5VV
2021 · github.com
- 6

Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API. - carloslfu/slotstream
5d ago · github.com
- 7PM
2021 · apps.apple.com
- 8MA
2016 · github.com
- 9IM
2015 · github.com
- 10AM
2024 · github.com
- 11

- 12AQ
2020 · getutm.app
- 13PG
2015 · github.com
- 14C2
2021 · github.com
- 15LM
2021 · jonathanalland.com
- 16TF
I’d originally launched my app: Private LLM[1][2] on HN around 10 months ago, with a single RedPajama Chat 3B model. The app has come a long way since then. About a month ago, I added support for 4-bit OmniQuant quantized Mixtral 8x7B Instruct model, and it seems to outperform Q4 models at inference speed and Q8 models at text generation quality, while consuming only about 24GB of RAM[3] at 8k context length. The trick is: a) to use a better quantization algorithm and b) to use unquantized embeddings and the MoE gates (the overhead is quite small). Other notable features include many more…
2024
- 17IF
Mar 2026 · github.com
- 18MW
I always find myself frustrated by how many steps I have to take to video chat with someone online. There's always too much software to download and install, and too many accounts to remember. The high-end, high-price Cisco conferencing systems I've used are especially fragile. So I built Vidless in a weekend for myself, and I hope you find it useful too. Just create a room and share the link, and you'll be chatting in seconds. You can invite several people (I haven't really tested a max yet), and chats on Vidless are always private. Enjoy!
2012 · vidless.com
- 19RM
Apr 2026 · github.com
- 20FP
2015 · github.com
- 21IM
I made my first macOS utility app that ships with a bundled Gemma 4 model, specifically the Gemma E4B one. It made my app DMG have 5.3 GB in size, but I think it is a small size for the power that this free local model can provide. It runs fine on CPU, but can also run on Apple Silicon GPU, although I did not notice any performance improvements with GPU (tested on a M5 chip). I think these local lightweight and multimodal models will open multiple possibilities for new software tools where privacy is essential.
May 2026 · snapname.app
- 22IB
hey hn, I built an open-source Perplexity clone that can run local LLMs and cloud LLMs. It's fully self-hostable through Docker and uses ollama to support local LLMs. The demo video in the repository shows me running it locally with llama3 on my M1 Macbook Pro. I'm open to any suggestions or feedback, thanks!
2024 · github.com
- 23NI
2022 · github.com
- 24AC
2016 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →