Alternatives
Products that do what High-end desktop supercomputers for AI does
Run and tune the biggest large language models locally
- 1

- 2SF
Hey folks! We're Alex and Evan, and we're working on putting together a 512 H100 compute cluster for startups and researchers to train large generative models on. - it runs at the lowest possible margins (<$2.00/hr per H100) - designed for bursty training runs, so you can take say 128 H100s for a week - you don’t need to commit to multiple years of compute or pay for a year upfront Big labs like OpenAI and Deepmind have big clusters that support this kind of bursty allocation for their researchers, but startups so far have had to get very small clusters on very long term contracts, wait…
2023 · sfcompute.org
- 3

A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me. But then I thought, "I wonder how it would work on a normal computer like mine," and above all, "I wonder if it would work without going into OOM on a computer like mine." So I started working with the help of agents to test this possibility. I started converting the model to int4, understanding MTP usage, and if possible implementing DSA for long context.…
Jul 2026 · github.com
- 4
General Compute▲315AI models that run on an inference cloud optimized for speed
May 2026 · generalcompute.com
- 5

- 6

- 7

- 8

- 9MO
I wanted to share our new speech to text model, and the library to use them effectively. We're a small startup (six people, sub-$100k monthly GPU budget) so I'm proud of the work the team has done to create streaming STT models with lower word-error rates than OpenAI's largest Whisper model. Admittedly Large v3 is a couple of years old, but we're near the top the HF OpenASR leaderboard, even up against Nvidia's Parakeet family. Anyway, I'd love to get feedback on the models and software, and hear about what people might build with it.
Feb 2026 · github.com
- 10

Powers faster, efficient reasoning for long-running agents
Jun 2026 · developer.nvidia.com
- 11

- 12WM
Try it out! https://glhf.chat/ Hey HN! We’ve been working for the past few months on a website to let you easily run (almost) any open-source LLM on autoscaling GPU clusters. It’s free for now while we figure out how to price it, but we expect to be cheaper than most GPU offerings since we can run the models multi-tenant. Unlike Together AI, Fireworks, etc, we’ll run any model that the open-source vLLM project supports: we don’t have a hardcoded list. If you want a specific model or finetune, you don’t have to ask us for it: you can just paste the Hugging Face link in and…
2024 · glhf.chat
- 13

- 14

- 15

- 16

Affordable H100, H200, GB300, and B200 GPU compute for training, inference, and everything in between.
4d ago · compute.cheap
- 17

- 18

- 19

- 20RQ
Sep 2025 · github.com
- 21

- 22EL
2023 · github.com
- 23

Your private, planet-friendly AI assistant from the UK.
Mar 2026 · gb1.ai
- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →