nowfound

Alternatives

Products that do what Llmpm – NPM for LLMs does

npm for LLMs — install, run, and share AI models. We’ve built llmpm, a CLI tool that makes open-source LLMs installable like packages. llmpm install llama3 llmpm run llama3 You can also package models with your projects so others can reproduce the same setup easily. Website: https://llmpm.co GitHub:https://github.com/llmpm/llmpm-dev

  1. 1
    AskCodi230

    Custom LLMs, without training. Use via openai compatible api

    Nov 2025

  2. 2
    LM Studio209

    Discover, download, and run local LLMs (incl. DeepSeek R1)

    2025

  3. 3

    Vibe-check many open-source and proprietary LLMs at once

    2024

  4. 4
    Aqueduct107

    The easiest way to run open source LLMs

    2023

  5. 5
    Gradient153

    Developer API for building private LLMs that you own

    2023

  6. 6
    LLMTest125

    Use the right LLMs in your apps. Setup fallbacks. Be happy.

    May 2026

  7. 7
    Taylor AI118

    Fine-tune open source LLMs in minutes

    2023

  8. 8

    Build local LLMs using top data science libraries

    2023

  9. 9
    ChattyUI149

    Run open-source LLMs locally in the browser using WebGPU

    2024

  10. 10

    Aggregate uptime monitoring across OpenAI, Claude, and more

    Apr 2026

  11. 11
    Llama91

    A terminal file manager

    2021

  12. 12
    Colossal135

    Effortlessly integrate tool-using agents with a single fetch

    2025

  13. 13

    Test-driven development for LLMs

    2023

  14. 14

    Bring reliable AI virtual assistants to your app

    2024

  15. 15
    Harbor75

    CLI + companion App to spin up complete local LLM stacks

    May 2026

  16. 16AC

    Hi HN, we're Ashpreet, Eli and Yash and we're excited to share Phidata: a collection of AI Apps built with open-source tools. While helping teams build AI products, we built templates for spinning up LLM Apps quickly. Today we're open-sourcing our templates for building: - RAG LLM Apps - Autonomous LLM Apps - Multimodal LLM Apps - Data Engineering LLM Apps Templates are built with FastApi for serving, Streamlit for prototyping, PgVector for vectors and PosgreSQL for storage. Run them locally using docker and in production on AWS - with 1 command. - Github:…

    2023 · github.com

  17. 17DM

    Hi HN, I’m one of the authors of this post. We’ve updated Docker Model Runner to support vLLM alongside the existing llama.cpp backend. The goal is to bridge the gap between local prototyping (often done with GGUF/llama.cpp) and high-throughput production (often done with Safetensors/vLLM) using a consistent Docker workflow. Key technical details: Auto-routing: The tool detects the model format. If you pull a GGUF model, it routes to llama.cpp. If you pull a Safetensors model, it routes to vLLM. API: It exposes an OpenAI-compatible API (/v1/chat/completions), so the…

    Nov 2025 · github.com

  18. 18LA

    You build LLM applications with YAML files, that define an execution graph. Nodes can be either LLM API calls, regular function executions or other graphs themselves. Because you can nest graphs easily, building complex applications is not an issue, but at the same time you don't lose control. The YAML basically states what are the tasks that need to be done and how they connect. Other than that, you only write individual python functions to be called during the execution. No new classes and abstractions to learn.

    2024 · github.com

  19. 19IB

    hey hn, I built an open-source Perplexity clone that can run local LLMs and cloud LLMs. It's fully self-hostable through Docker and uses ollama to support local LLMs. The demo video in the repository shows me running it locally with llama3 on my M1 Macbook Pro. I'm open to any suggestions or feedback, thanks!

    2024 · github.com

  20. 20RA

    Hi there, looking for feedback on my new project "Featherless.AI" The idea is to allow users to run all the models on hugging face instantly. Via the OpenAI API compatible endpoint. Why? Because its a real chore to download models and spin up GPUs, especially if you want to test multiple models. Not to mention GPUs cost multiple dollars an hour to rent. And if we want more people to use open source AI, we got to make it easier for them to try and play with all of them. So what if instead of spinning up dedicated GPUs per model (which is what every provider is doing) We can startup a LLM…

    2024 · featherless.ai

  21. 21AP

    Hey HackerNews, I'm building an open-source library aiming to make it very easy for anyone to use plugins with any LLM (plugins as defined by OpenAI). I just finished this very simple tutorial: https://github.com/edreisMD/plugnplai/blob/master/examples/a.... And would love to get some feedback, and suggestions on how to improve it / make it useful for you. Steps: 1. Load plugins from https://plugnplai.com directory (now with ~150 plugins) 2. Install and activate: Load specifications and build a default prompt describing the plugins to…

    2023 · twitter.com

  22. 22LS

    LLMStack is a low-code platform that can be used to build LLM apps, chatbots and integrate AI experiences into existing products/workflows. It comes with everything out of the box that one needs to build LLM apps locally. It can also be used in a multi-tenant setting, making it available for everyone to use in an enterprise. Some highlights of the platform: - Chain multiple LLM models allowing for complex pipelines - Includes a vector database and necessary connectors to help enrich LLM responses with private data - App templates tailored to specific use cases to quickly build LLM apps…

    2023 · github.com

  23. 23LK

    Hi HN! I built LLMKube, a Kubernetes operator for deploying GPU-accelerated LLMs in production. One command gets you from zero to inference with full observability. Why this exists: Regulated industries (healthcare, defense, finance) need air-gapped LLM deployments, but existing tools are either single-node only (Ollama) or lack GPU optimization and SLO enforcement. LLMKube bridges the gap. What's working: - 17x speedup with NVIDIA GPUs (64 tok/s on Llama 3.2 3B vs 4.6 tok/s CPU) - One command: llmkube deploy llama-3b --gpu (auto CUDA setup, scheduling, layer offloading) -…

    Nov 2025 · github.com

  24. 24HF

    We have a massive GPU cluster and developed our own infrastructure to manage the cluster and train massive models. There's how it works: 1. You upload the dataset with preconfigured format into HuggingFaсe [1]. 2. Choose your LLM (e.g. LLaMa 70B, Mistral 7B) 3. Place your submission into the queue 4. Wait for it to get trained. 5. Then you get your trained model there on HuggingFace. Essentially, why would we want to do it? 1. We already have an experience with training big LLMs. 2. We could achieve near-perfect infrastructure performance for training. 3. Sometimes GPUs have just nothing to…

    2023 · higgsfield.xyz

Ranked by how close each launch is in meaning, then by votes. Refine with a description →