Alternatives
Products that do what Selfhostllm.org – Plan GPU capacity for self-hosting LLMs does
A simple calculator that estimates how many concurrent requests your GPU can handle for a given LLM, with shareable results.
- 1

- 2

- 3

- 4

- 5

- 6

- 7SH
Jun 2026 · github.com
- 8

- 9

- 10

- 11

- 12AS
2019 · github.com
- 13IB
Built a simple web app that tells you which open-source LLMs will work on your hardware. It auto-detects your specs, shows compatible models from Hugging Face, gives realistic performance estimates (tokens/sec), and recommends quantization settings. You can also manually input specs to see "what if I upgraded my RAM?" Made this after wasting time downloading giant models only to find they crawled on my hardware. Hope it saves you some frustration!
2025 · caniusellm.com
- 14

- 155L
We've built InferX, a specialized runtime environment that fundamentally changes how LLMs are served. The core problem we solve is the latency bottleneck in AI inference, especially with large models. Current systems waste resources or suffer from painfully slow cold starts. InferX's AI-native architecture, with its "snapshot" technology, enables: * *Sub-2s cold starts:* Spin up models instantly. * *High density:* Serve more LLMs on the same GPUs. * *Optimal efficiency:* Maximize GPU utilization. This isn't just another API; it's a new execution layer designed from the ground up for the…
2025 · github.com
- 16IB
I was overspending on GPT-4o. It was really hard to compare different models I could switch to, so I built this LLM comparison tool. It shows leaderboards, pricing, and performance data across 100+ LLMs (including all major providers and open-source models). Key features: - Live pricing comparisons - Benchmark Scores (MMLU, HumanEval, GPQA, etc.) - Context length vs cost analysis - Speed/throughput tests across providers - Quality vs price visualizations - Open source (all data verifiable) Try it out: https://llmstats.com I'd like to know your opinion :) Tech stack: Next.js,…
2025 · llm-stats.com
- 17

- 18MI
I've been working on a platform that uses LLMs to build maintain and manage k8s clusters on any cloud. The system writes infra as code to your Github repos and automatically containerizes and scales any services (public or private). The goal is to give your average engineer a vercel-like deployment experience for any service in any language at minimal cost. We have humans involved at the moment auditing LLM outputs and keeping an eye on clusters. We are looking for folks who may be thinking about their first infra/devops hire. Just connect your github and your cloud provider. The system…
2024 · milkinfrastructure.com
- 19CT
2012 · kloudcalc.com
- 20SA
Hi guys, we're Wilhem from Paris and Jean-Daniel from Tokyo, software engineers with a passion for all things cloud (IaaS, PaaS, SaaS). We recently decided to tackle the problem of Capacity Planning with Stacktical, a Scalability Prediction service (https://stacktical.com). For a decade, we've been observing our clients and colleagues trying to nail down their strategy using repeated cycles of defining, collecting and interpreting load testing campaigns. It's funny how most people don't realize how demanding the work of infrastructure managers and their teams really is... While…
2016
- 21HF
We have a massive GPU cluster and developed our own infrastructure to manage the cluster and train massive models. There's how it works: 1. You upload the dataset with preconfigured format into HuggingFaсe [1]. 2. Choose your LLM (e.g. LLaMa 70B, Mistral 7B) 3. Place your submission into the queue 4. Wait for it to get trained. 5. Then you get your trained model there on HuggingFace. Essentially, why would we want to do it? 1. We already have an experience with training big LLMs. 2. We could achieve near-perfect infrastructure performance for training. 3. Sometimes GPUs have just nothing to…
2023 · higgsfield.xyz
- 22LB
2025 · github.com
- 23LA
Feb 2026 · github.com
- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →