nowfound

Alternatives

Products that do what Xybrid – run LLM and speech locally in your app (no back end, Rust) does

Hi HN, We built Xybrid, a Rust library for running LLM + speech pipelines directly inside your app, no server, no daemon, just one binary. We started building it while working on a privacy-focused LLM app with Tauri and realized there wasn’t a straightforward way to embed models directly into shipped applications without relying on a separate server process. Xybrid links into your process like any other library. It supports GGUF / ONNX / CoreML and integrates with Flutter, Swift, Kotlin, Unity, and Tauri, letting you run pipelines like speech → LLM → speech in a single call. On…

  1. 1LA

    G'day, HN! I'm one of the maintainers of `llm`. I've been working alongside a trusty group of contributors to bring this project to life, and we're now at a point where we're ready to share it with the world. Large language models (LLMs) are taking the computing world by storm due to their emergent abilities that allow them to perform a wide variety of tasks, including translation, summarization, code generation, and even some degree of reasoning. However, the ecosystem around LLMs is still in its infancy, and it can be difficult to get started with these models. `llm` is a one-stop shop for…

    2023 · github.com

  2. 2
    ChattyUI149

    Run open-source LLMs locally in the browser using WebGPU

    2024

  3. 3
    Aqueduct107

    The easiest way to run open source LLMs

    2023

  4. 4
    OpenWispr190

    100% local open source AI speech-to-text model

    2025

  5. 5WM

    Try it out! https://glhf.chat/ Hey HN! We’ve been working for the past few months on a website to let you easily run (almost) any open-source LLM on autoscaling GPU clusters. It’s free for now while we figure out how to price it, but we expect to be cheaper than most GPU offerings since we can run the models multi-tenant. Unlike Together AI, Fireworks, etc, we’ll run any model that the open-source vLLM project supports: we don’t have a hardcoded list. If you want a specific model or finetune, you don’t have to ask us for it: you can just paste the Hugging Face link in and…

    2024 · glhf.chat

  6. 6SL

    For speech-to-text, large-language-model inference and text-to-speech I created three wrapper libraries in C/C++ (using Whisper.cpp, Llama.cpp and Piper). Follow the URL to see an example that shows how to use these libraries for a speech-to-text, LLM inference, text-to-speech pipeline. Windows and Linux are supported.

    Sep 2025 · github.com

  7. 7
    NobodyWho106

    Hey, Hello, we are the team behind NobodyWho; an open-source plugin that embeds large language models functionality inside Unity and Godot games - completely locally and offline. We try to make it as easy and user friendly as possible to use our plugin, so feel free to create an issue on our GitHub if you have any inconveniences - and we will try to improve it! So why might this interest you: - Written in Rust with llama.cpp bindings and optional GPU back-ends (Vulkan / Metal) for fast inference. - Drop-in nodes / components: add a model asset and a chat or embedding node and get a…

    18d ago · github.com

  8. 8OL

    OpenBrief is basically a GUI for yt-dlp with some AI on top — paste a link, it downloads locally, and transcription and voice generation run with local AI on your machine. Summaries and chat over the transcript use an LLM, which is bring-your-own-key for now. It's open source and free.

    May 2026 · github.com

  9. 9WR

    A WebAssembly runtime embedded in Godot game engine projects. Interact with Wasm modules from GDScript. Accessing Wasm modules via GDScript provides the following benefits. - Sandboxed environment allows safely loading zero-trust mods/extensions to a project. - Single target means that Wasm modules can be built from any language e.g. Rust, Go, AssemblyScript and the single binary can be run on all platforms. - Fast execution compared to GDScript allows for offloading compute-heavy operations or running bots/mods/etc. at high FPS. This might be useful for loading mods, bot AI,…

    2023 · github.com

  10. 10CI

    One of the most frequent questions one faces while running LLMs locally is: I have xx RAM and yy GPU, Can I run zz LLM model ? I have vibe coded a simple application to help you with just that. Update: A lot of great feedback for me to improve the app. Thank you all.

    2025 · can-i-run-this-llm-blue.vercel.app

  11. 11

    Harness local AI for notes

    Jul 2026 · voice-to-md.xajik0.workers.dev

  12. 12LT

    Current AI-assisted CLI tools are often part of larger systems and work better on Linux. I built llm-term to address these. It's a Rust-based tool that compiles into a single binary file. You only need to download the binary, add it to your PATH, and configure your OpenAI key to get started. While llm-term offers an option for gpt-4o, it works great with gpt-4o-mini. So it's not costly. I appreciate any feedback or suggestions.

    2024 · github.com

  13. 13MG

    Hello HN, I've been working on this project for a while, and it has been in an "open" beta for some time. I finally believe it's ready for its first release. I hope you like it. Here are some potential questions that may arise: 1. How does it compare to LM Studio? It's likely that if you're already using LM Studio, you'll continue to do so. This project is designed to be more user-friendly. 2. Is it open-source? No, it is not. 3. Does it use any open-source libraries? Yes, it uses llama.cpp and a few others, as indicated in the license information included with the application. 4. Why is not…

    2023 · avapls.com

  14. 14

    High performance secure & portable Rust functions in Node.js

    2020

  15. 15
    Verby72

    Free text to speech converter with SSML editor

    2019

  16. 16UL

    Recently featured in a LangChain blog https://blog.langchain.dev/empowering-development-with-flowt... , use LLMs to construct an API first runnable workflow with an IDE experience.

    2024 · github.com

  17. 17CE

    Hi all, I open sourced my toy project that runs Generative AI models LOCALLY in the side panel of a Chrome extension. The Chrome extension uses Transformers.js to run models in browser under the hood. I've integrated and tested these models so far. \1. LLM: Llama 3, Phi 3.5, Qwen 2.5, SmolLM2 \2. Reasoning: DeepSeek R1 \3. Multimodal LLM: Janus \4. Speech-to-Text: Whisper On an M1 MacBook, DeepSeek R1 1.5B runs at ~30 tokens/sec If you're interested in, you can download the extension from chrome web store or clone my github repository. \1. chrome web store:…

    2025 · github.com

  18. 18GA

    Gerbil is an open source app that I've been working on for the last couple of months. The development now is largely done and I'm unlikely to add anymore major features. Instead I'm focusing on any bug fixes, small QoL features and dependency upgrades. Under the hood it runs llama.cpp (via koboldcpp) backends and allows easy integration with the popular modern frontends like Open WebUI, SillyTavern, ComfyUI, StableUI (built-in) and KoboldAI Lite (built-in). Why did I create this? I wanted an all-in-one solution for simple text and image-gen local LLMs. I got fed up with needing to manage…

    Nov 2025 · github.com

  19. 19AR
  20. 20LL

    What it is A single 45 MB Windows .exe that embeds llama.cpp and a minimal Tk UI. Copy it (plus any .gguf model) to a flash drive, double-click on any Windows PC, and you’re chatting with an LLM—no admin rights, Cloud, or network. Why I built it Existing “local LLM” GUIs assume you can pip install, pass long CLI flags, or download GBs of extras. I wanted something my less-technical colleagues could run during a client visit by literally plugging in a USB drive. How it works PyInstaller one-file build → bundles Python runtime, llama_cpp_python, and the UI into a single PE. On first launch, it…

    2025 · github.com

  21. 21AE
  22. 22

    Make your coding agent talk to you! Offline!

    14d ago · github.com

  23. 23ML

    Motörhead is an open-source project aimed at helping developers build and manage LLM-based chat applications. We built Motörhead in Rust and it currently handles user sessions and context windows with the following features: -Short-term memory management: Motörhead maintains chat session histories for n users, keeping track of messages sent and received. -Incremental summarization: When the message buffer reaches a maximum window size, Motörhead incrementally summarizes the content into a context key. This process involves taking an existing summary and the messages removed from the list,…

    2023 · github.com

  24. 24HA

    Demo starts at 50m into the video. This was a bit terrifying to record because 2am the previous night everything was totally broken after a major refactor (so that we could add external LLM support as well as local GPUs). But pressure can be a useful force :-D We start with a stack deployed on my laptop without a GPU, pointing to together.ai so we can run open source LLMs easily without having to have access to a GPU. We show simple inference through the ChatGPT-like web interface (with users, sessions etc) and then simple drag'n'drop RAG. Then we show some helix apps defined as yaml: Marvin…

    2024 · youtube.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →