nowfound

Alternatives

Products that do what GPT Router does

Avoid OpenAI downtimes - one API for 30+ LLMs

  1. 1

    Ollama but for mobile, with a cloud fallback

    2025

  2. 2
    liteLLM120

    One library to standardize all LLM APIs

    2023

  3. 3
    AskCodi230

    Custom LLMs, without training. Use via openai compatible api

    Nov 2025

  4. 4

    Access 1 billion tokens per month for free

    Apr 2026

  5. 5

    Trace LLM requests + costs with OpenTelemetry monitoring

    Oct 2025

  6. 6
    Taylor AI118

    Fine-tune open source LLMs in minutes

    2023

  7. 7

    LLM Provider arbitrage to get the best performance for the $

    2025

  8. 8

    Calculate the GPU memory you need for LLM inference

    2025

  9. 9
    Dolly113

    Democratizing the magic of ChatGPT with open models

    2023

  10. 10
    ChattyUI149

    Run open-source LLMs locally in the browser using WebGPU

    2024

  11. 11
    Openlit152

    One click observability & evals for LLMs & GPUs

    2024

  12. 12
    LLMTest125

    Use the right LLMs in your apps. Setup fallbacks. Be happy.

    May 2026

  13. 13
    Tokenwise143

    A smart LLM proxy that shows where you're overpaying

    Jun 2026

  14. 14

    Aggregate uptime monitoring across OpenAI, Claude, and more

    Apr 2026

  15. 15
    Aqueduct107

    The easiest way to run open source LLMs

    2023

  16. 16
    Perssua61

    Real-time guidance from any LLM (including local ones)

    Nov 2025

  17. 17

    Your AI, fully offline with Zero data collection & 100% free

    May 2026

  18. 18

    One balance. Every model. Chat, image, video & audio.

    Jun 2026

  19. 19

    One AI API for production - streaming, failover, logs

    Jan 2026

  20. 20

    Version, test, and collaborate on LLM prompts— like code

    2025

  21. 21

    Free multi-provider LLM proxy with automatic failover

    11d ago · github.com

  22. 22SA

    Hi HN, We’re building https://www.switchpoint.dev – a drop-in replacement for OpenAI’s API that reduces LLM cost by smartly routing across models (e.g., Claude, Gemini, GPT-4) depending on subject and difficulty of the task. Why we built this: LLM costs are spiraling—especially for products doing retrieval, agentic reasoning, or even just high-volume chat. We were frustrated with paying GPT-4 rates when most queries didn’t need it. So we built a router that: - Starts with cheaper/free models (like Llama 8B, 4o-mini, 2.0 flash) - Streams responses and upgrades on failure - Acts…

    2025 · switchpoint.dev

  23. 23LT

    I wanted to share a project I've been working on for the past few weeks: llgtrt. It's a Rust implementation of a HTTP REST server for hosting Large Language Models using llguidance library for constrained output with NVIDIA TensorRT-LLM. The server is compatible with the OpenAI REST API and supports structured JSON schema enforcement as well as full context-free grammars (via Guidance). It's similar in spirit to the Python-based TensorRT-LLM OpenAI server example but written entirely in Rust and built with constraints in mind. No Triton Inference Server involved. This also serves as a demo…

    2024 · github.com

  24. 24AC

    There's LLM Council and similar tools, but they use predefined model lineups. This one is different in a few ways that mattered to me: *Bring your own models.* Mix Ollama (local), OpenAI, Anthropic, Groq, Google — or any OpenAI-compatible endpoint — in whatever combination you want. A council of DeepSeek-R1 + llama2-uncensored + mistral-nemo is a very different deliberation than GPT-4o + Claude + Gemini. *Zero server, zero account, zero storage.* The app is purely static. API calls go directly from your browser to providers. Nothing touches a backend. No tokens, no sessions, no analytics.…

    Feb 2026 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →