nowfound

Alternatives

Products that do what BaseRT does

6.4x faster than llama.cpp, 3.9x faster than MLX

  1. 1

    Spin up a Mac Mini M1 in the cloud.

    2021

  2. 2

    Flat rate to the best LLMs for OpenClaw, Hermes Agent, etc.

    Apr 2026

  3. 3
    Cai179

    Press ⌥C on anything to run smart actions, locally

    Apr 2026

  4. 4RM
  5. 5
    Ollama235

    The easiest way to run large language models locally

    2023

  6. 6

    Vibe-check many open-source and proprietary LLMs at once

    2024

  7. 7

    The fastest ipsum generator on the 🌎⚡

    2019

  8. 8
    ChattyUI149

    Run open-source LLMs locally in the browser using WebGPU

    2024

  9. 9
    Mammouth190

    Get access to the best LLMs in one place for 10€

    2024

  10. 10

    LLM reinforcement fine-tuning platform to improve LLM output

    2025

  11. 11
    Radio LLM141

    Off-grid, disaster-proof LLM platform using Meshtastic

    2024

  12. 12NT

    With the latest launch from Google I've added support for Gemma 3 270M, the speed for local LLM to TTS token time is incredible! This is an heavy obvious work in progress - any contributions or tips would be welcome. The idea is to have a fast moving edge model playground, and maybe have some utility (like the e reader) on the side.

    2025 · github.com

  13. 13LA

    We have released a high-performance CMS (2-10ms in our tests) CMS, that needs to be set-up by a programmer. Enjoy! EDIT: Since links in a text post do not become clickable, I've moved them to the comments.

    2015

  14. 14IB

    hey hn, I built an open-source Perplexity clone that can run local LLMs and cloud LLMs. It's fully self-hostable through Docker and uses ollama to support local LLMs. The demo video in the repository shows me running it locally with llama3 on my M1 Macbook Pro. I'm open to any suggestions or feedback, thanks!

    2024 · github.com

  15. 15IB

    After fine-tuning GPT for a personal project, I realized how tedious it is to write plain text in a massive JSON file. That's why I built this app for my own use, and I want to see if others could benefit from a tool like this as well ;)

    2024 · finetuna-ui.com

  16. 16LT

    I wanted to share a project I've been working on for the past few weeks: llgtrt. It's a Rust implementation of a HTTP REST server for hosting Large Language Models using llguidance library for constrained output with NVIDIA TensorRT-LLM. The server is compatible with the OpenAI REST API and supports structured JSON schema enforcement as well as full context-free grammars (via Guidance). It's similar in spirit to the Python-based TensorRT-LLM OpenAI server example but written entirely in Rust and built with constraints in mind. No Triton Inference Server involved. This also serves as a demo…

    2024 · github.com

  17. 17LF

    I submitted an earlier version of this a few months ago (as llama2.f90). At that time it had a lot of steps to run and was just a toy, now it's easy to run and is a competitive option for llm inference. See the motivation section for discussion and the `Performance` issue for an ongoing discussion about performance.

    2023 · github.com

  18. 18AO

    I've built an airgapped Retrieval-Augmented Generation (RAG) system for question-answering on documents, running entirely offline with local inference. Using Llama 3, Mistral, and Gemini, this setup allows secure, private NLP on your own machine. Perfect for researchers, data scientists, and developers who need to process sensitive data without cloud dependencies. Built with Llama C++, LangChain, and Streamlit, it supports quantized models and provides a sleek UI for document processing. Check it out, contribute, or suggest new features!

    2024 · github.com

  19. 19CG

    Just added support for Llama-3 models to our AI app platform Promptly. We decided to try Groq cloud for powering these models and the results have so far been pretty good comparing Llama-3-70B with GPT-4 Turbo. Put an app together to compare these models. Check it out at https://trypromptly.com/a/groq-llama-3-70b-vs-gpt-4-turbo. https://trypromptly.com/s/iQG7EoJ4Pm is a sample output comparison between Llama-3-70B and GPT-4 turbo. https://youtu.be/1UChY6EDwFA shows the inference speed of Groq compared to GPT-4.

    2024 · trypromptly.com

  20. 20ML

    Time to first token is 39% faster Agent wall times decrease by 46% No swaps Tracks your resource usage in real-time and adjusts how the model runs so that it works perfectly on your device. Implements KV cache sizing, prefix caching, live RAM pressure management, context trimming, KV quantization, and more. Built a ton of features

    Jun 2026 · autotunellm.com

  21. 21NL

    Refuel LLM (84.2%) outperforms trained human annotators (80.4%), GPT-3-5-turbo (81.3%), PaLM-2 (82.3%) and Claude (79.3%) across a benchmark of 15 text labeling datasets. It is a Llama-v2-13b base model, trained on over 2500 unique datasets (5.24B tokens) spanning categories such as classification, entity resolution, matching, reading comprehension and information extraction. Here is the interactive demo: https://labs.refuel.ai/playground. Pretty fun to play with!

    2023

  22. 22MI

    Hi HN! I lead product at Vectara and we've just released a new LLM in our platform that outperforms GPT4 and Gemini 1.5 Pro on RAG tasks. Vectara is a Retrieval Augmented Generation (RAG) platform primarily deployed as a SaaS service which includes a generous free tier so you can try it for free. The way we've been able to offer a "better but cheaper" is that we focus a lot of our attention on taking smaller models (which can be hosted in a cost efficient way) and fine tuning them to specific tasks: in this case RAG. This ends up with a model that is less capable of arbitrary tasks like…

    2024 · vectara.com

  23. 23AL

    Try it out here: https://labs.refuel.ai/playground Refuel LLM (84.2%) outperforms trained human annotators (80.4%), GPT-3-5-turbo (81.3%), PaLM-2 (82.3%) and Claude (79.3%) across a benchmark of 15 text labeling datasets. It is a Llama-v2-13b base model, trained on over 2500 unique datasets (5.24B tokens) spanning categories such as classification, entity resolution, matching, reading comprehension and information extraction.

    2023

  24. 24AO

Ranked by how close each launch is in meaning, then by votes. Refine with a description →