Alternatives
Products that do what Rapid-MLX – Run local LLMs on Mac, 2-3x faster than alternatives does
- 1VV
2021 · github.com
- 2

- 3MO
2018 · micromdm.io
- 4PA
2020 · github.com
- 5MA
2017 · marta.yanex.org
- 6TF
I’d originally launched my app: Private LLM[1][2] on HN around 10 months ago, with a single RedPajama Chat 3B model. The app has come a long way since then. About a month ago, I added support for 4-bit OmniQuant quantized Mixtral 8x7B Instruct model, and it seems to outperform Q4 models at inference speed and Q8 models at text generation quality, while consuming only about 24GB of RAM[3] at 8k context length. The trick is: a) to use a better quantization algorithm and b) to use unquantized embeddings and the MoE gates (the overhead is quite small). Other notable features include many more…
2024
- 7LS
2020 · lacona.app
- 8MG
Hello HN, I've been working on this project for a while, and it has been in an "open" beta for some time. I finally believe it's ready for its first release. I hope you like it. Here are some potential questions that may arise: 1. How does it compare to LM Studio? It's likely that if you're already using LM Studio, you'll continue to do so. This project is designed to be more user-friendly. 2. Is it open-source? No, it is not. 3. Does it use any open-source libraries? Yes, it uses llama.cpp and a few others, as indicated in the license information included with the application. 4. Why is not…
2023 · avapls.com
- 9PF
2021 · github.com
- 10MA
2014 · mjolnir.io
- 11LL
2023 · github.com
- 12LA
Feb 2026 · github.com
- 13IB
hey hn, I built an open-source Perplexity clone that can run local LLMs and cloud LLMs. It's fully self-hostable through Docker and uses ollama to support local LLMs. The demo video in the repository shows me running it locally with llama3 on my M1 Macbook Pro. I'm open to any suggestions or feedback, thanks!
2024 · github.com
- 14

- 15CO
Apr 2026 · npmjs.com
- 16LA
Jun 2026 · rorlikowski.github.io
- 17IB
Built a simple web app that tells you which open-source LLMs will work on your hardware. It auto-detects your specs, shows compatible models from Hugging Face, gives realistic performance estimates (tokens/sec), and recommends quantization settings. You can also manually input specs to see "what if I upgraded my RAM?" Made this after wasting time downloading giant models only to find they crawled on my hardware. Hope it saves you some frustration!
2025 · caniusellm.com
- 18LA
2017 · github.com
- 19QA
2025 · qspeak.app
- 20JA
Mar 2026 · github.com
- 21IL
2023 · github.com
- 22TT
2020 · typer.tiangolo.com
- 23VA
May 2026 · voxxy.io
- 24SH
Jun 2026 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →