nowfound

Alternatives

Products that do what Rapid-MLX – Run local LLMs on Mac, 2-3x faster than alternatives does

  1. 1VV
  2. 2
    BaseRT222

    6.4x faster than llama.cpp, 3.9x faster than MLX

    Jul 2026 · basecompute.co

  3. 3MO
  4. 4PA
  5. 5MA
  6. 6TF

    I’d originally launched my app: Private LLM[1][2] on HN around 10 months ago, with a single RedPajama Chat 3B model. The app has come a long way since then. About a month ago, I added support for 4-bit OmniQuant quantized Mixtral 8x7B Instruct model, and it seems to outperform Q4 models at inference speed and Q8 models at text generation quality, while consuming only about 24GB of RAM[3] at 8k context length. The trick is: a) to use a better quantization algorithm and b) to use unquantized embeddings and the MoE gates (the overhead is quite small). Other notable features include many more…

    2024

  7. 7LS
  8. 8MG

    Hello HN, I've been working on this project for a while, and it has been in an "open" beta for some time. I finally believe it's ready for its first release. I hope you like it. Here are some potential questions that may arise: 1. How does it compare to LM Studio? It's likely that if you're already using LM Studio, you'll continue to do so. This project is designed to be more user-friendly. 2. Is it open-source? No, it is not. 3. Does it use any open-source libraries? Yes, it uses llama.cpp and a few others, as indicated in the license information included with the application. 4. Why is not…

    2023 · avapls.com

  9. 9PF
  10. 10MA
  11. 11LL
  12. 12LA
  13. 13IB

    hey hn, I built an open-source Perplexity clone that can run local LLMs and cloud LLMs. It's fully self-hostable through Docker and uses ollama to support local LLMs. The demo video in the repository shows me running it locally with llama3 on my M1 Macbook Pro. I'm open to any suggestions or feedback, thanks!

    2024 · github.com

  14. 14

    Run local LLMs faster and smoother on your device

    May 2026 · autotunellm.com

  15. 15CO
  16. 16LA
  17. 17IB

    Built a simple web app that tells you which open-source LLMs will work on your hardware. It auto-detects your specs, shows compatible models from Hugging Face, gives realistic performance estimates (tokens/sec), and recommends quantization settings. You can also manually input specs to see "what if I upgraded my RAM?" Made this after wasting time downloading giant models only to find they crawled on my hardware. Hope it saves you some frustration!

    2025 · caniusellm.com

  18. 18LA
  19. 19QA
  20. 20JA

    Mar 2026 · github.com

  21. 21IL
  22. 22TT

    2020 · typer.tiangolo.com

  23. 23VA
  24. 24SH

    Jun 2026 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →