We made glhf.chat – run almost any open-source LLM, including 405B
Try it out! https://glhf.chat/ Hey HN! We’ve been working for the past few months on a website to let you easily run (almost) any open-source LLM on autoscaling GPU clusters. It’s free for now while we figure out how to price it, but we expect to be cheaper than most GPU offerings since we can run the models multi-tenant. Unlike Together AI, Fireworks, etc, we’ll run any model that the open-source vLLM project supports: we don’t have a hardcoded list. If you want a specific model or finetune, you don’t have to ask us for it: you can just paste the Hugging Face link in and…
In plain words
glhf.chat lets users run open-source large language models on autoscaling GPU clusters through a web interface. It supports any model compatible with the vLLM project, including models up to 405B parameters, by accepting Hugging Face links directly without requiring pre-approval. The platform runs models multi-tenant for cost efficiency and is currently free. It's designed for developers and researchers who want flexible access to diverse open-source language models without vendor constraints.
written from the facts on this page · September 2026
From the sources
In the maker’s words, at launch
Try it out! https://glhf.chat/ Hey HN! We’ve been working for the past few months on a website to let you easily run (almost) any open-source LLM on autoscaling GPU clusters. It’s free for now while we figure out how to price it, but we expect to be cheaper than most GPU offerings since we can run the models multi-tenant. Unlike Together AI, Fireworks, etc, we’ll run any model that the open-source vLLM project supports: we don’t have a hardcoded list. If you want a specific model or finetune, you don’t have to ask us for it: you can just paste the Hugging Face link in and it’ll work (as long as vLLM supports the base model architecture, we’ll run anything up to ~640GB of VRAM, give or take a little for some overhead buffer). Large models will take a few minutes to boot, but if a bunch of people are trying to use the same model, it might already be loaded and not need boot time at all. The Llama-3-70b finetunes are especially nice, since they’re basically souped-up versions of the 8b finetunes a lot of people like to run locally but don’t have the VRAM for. We’re expecting the Llama-3.1 finetunes to be pretty great too once they start getting released. There are some caveats for now — for example, while we support the Deepseek V2 architecture, we actually can only run their smaller “Lite” models due to some underlying NVLink limitations (though we’re working on it). But for the most part if vLLM supports it, we should too! We figured Llama-3.1-405B Launch Day was a good day to launch ourselves too — let us know in the comments if there’s anything you want us to support, or if you run into any issues. I know it’s not “local” Llama, but, well, that’s a lot of GPUs…
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 26d ago · cactuscompute.com


Launched alongside, July 2024
the whole month →



- IC
Many years ago, I made VJ softwares (to mix live visuals in clubs) for unexpected platforms like the Game Boy Advance, the Playstation 2 and the Raspberry Pi. This year, I’m back with a new web-app: Pikimov. Inspired by Photopea (a free Photoshop clone), I created this web-based motion design & video editor as an alternative to After Effects, to fill empty void. It's free, without signup, without cloud uploads (your files stay on your machine), and your projects are not used for AI models training.
AI · 2024 · pikimov.com
