nowfound

AI · October 9, 2024

SY

Slash your LLM Inference Costs with Overnight Processing

Hey HN, If you tried running open-source models like Llama 3.1 70B or 405B, you might have noticed that it gets very expensive. The reason looks obvious enough that you might have stopped even before trying it! - GPUs are very expensive to buy or rent - Running the most performing LLMs need 4, 8 or even 16 top of the line Nvidia GPUs - And that won’t get you anywhere near the level of VRAM needed to batch enough to get a decent throughput and efficiency Some have even questioned if open-source LLM providers are not doing some shenanigans to provide the prices they offer. VC funded…

What it does

In the maker’s words, at launch

Hey HN, If you tried running open-source models like Llama 3.1 70B or 405B, you might have noticed that it gets very expensive. The reason looks obvious enough that you might have stopped even before trying it! - GPUs are very expensive to buy or rent - Running the most performing LLMs need 4, 8 or even 16 top of the line Nvidia GPUs - And that won’t get you anywhere near the level of VRAM needed to batch enough to get a decent throughput and efficiency Some have even questioned if open-source LLM providers are not doing some shenanigans to provide the prices they offer. VC funded bait-and-switch? Unclear quantization? Even the most well funded LLM inference startups, with the best inference optimization teams in the world have got into controversy about this. At EXXA, we wanted to make affordable the best open-source LLMs in all their FP16 glory. And I don’t know for others, but we’re a bootstrapped team of 3, so the subsidizing part isn’t an option :D I won’t tell you we found the magical solution for all use cases… But we found one for batch overnight jobs to generate synthetic data or things like: - Data pre-processing (e.g. contextual retrieval for RAG enhancement, knowledge graph) - LLM-as-a-judge evaluation Why overnight? Because it gives us time to: - Get GPU for a high discount (30-70%) as they would otherwise sit idle in cloud providers data centers - Heavily optimize inference for maximum throughput instead of minimum latency Today, our batch inference API is live for Llama 3.1 8B & 70B FP16 with output under 24h. We offer the lowest price per token in the market! 60% cheaper than fireworks, 40% cheaper than deepinfra. Without any hard rate limits and with prompt caching available. Try it now at withexxa.com —-------- If you have any questions: you can contact us at [email protected] If you want to generate a large amount of tokens with custom LLM models, we can host them and offer the same price ranges as Llama 3.1 8B & 70B. What do you think of our approach? Are you willing to wait for super cheap prices for AI inference?

Does the same job

all alternatives →
  • SelfHostLLM2025 · ▲134

    Calculate the GPU memory you need for LLM inference

  • Soup CLI28d ago · trysoup.dev · ▲107

    Fine-tune an 8B LLM on a 4 GB laptop GPU

  • IB
    I built a tool to check if your computer can run LLMs locally2025 · caniusellm.com · ▲8

    Built a simple web app that tells you which open-source LLMs will work on your hardware. It auto-detects your specs, shows compatible models from Hugging Face, gives realistic performance estimates (tokens/sec), and recommends quantization settings. You can also manually input specs to see "what if I upgraded my RAM?" Made this after wasting time downloading giant models only to find they crawled on my hardware. Hope it saves you some frustration!

  • 5L
    50+ LLMs on 2 GPUs with 2-Second Swapping? We built AI-Native Runtime2025 · github.com · ▲5

    We've built InferX, a specialized runtime environment that fundamentally changes how LLMs are served. The core problem we solve is the latency bottleneck in AI inference, especially with large models. Current systems waste resources or suffer from painfully slow cold starts. InferX's AI-native architecture, with its "snapshot" technology, enables: * *Sub-2s cold starts:* Spin up models instantly. * *High density:* Serve more LLMs on the same GPUs. * *Optimal efficiency:* Maximize GPU utilization. This isn't just another API; it's a new execution layer designed from the ground up for the…

  • RA
    Run any Llama model finetune and more, instantly2024 · featherless.ai · ▲7

    Hi there, looking for feedback on my new project "Featherless.AI" The idea is to allow users to run all the models on hugging face instantly. Via the OpenAI API compatible endpoint. Why? Because its a real chore to download models and spin up GPUs, especially if you want to test multiple models. Not to mention GPUs cost multiple dollars an hour to rent. And if we want more people to use open source AI, we got to make it easier for them to try and play with all of them. So what if instead of spinning up dedicated GPUs per model (which is what every provider is doing) We can startup a LLM…

  • NeuroGridNov 2025 · ▲25

    Turn idle GPUs into cash. Get affordable AI for everyone.

More ai this month

the category →
  • I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.

    AI · 17d ago · simedw.com

  • Astute585

    Automate your B2B brand going viral, with new media creators

    AI · 18d ago · company-app.joinastute.com

  • Grok Bot547

    AI teammates that you can give real work to

    AI · 25d ago · x.ai

  • Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    AI · 26d ago · cactuscompute.com

  • Make your software self-driving

    AI · 30d ago · coldtea.ai

  • Soloop472

    Approval-first Agent OS for solo founders

    AI · 30d ago · soloop.io

Launched alongside, October 2024

the whole month →
  • buzzabout1,267

    Audience insights from 1B+ online discussions in 2 mins

    AI · 2024 · buzzabout.ai

  • bolt.new1,233

    Prompt, run, edit & deploy full-stack web apps

    AI · 2024 · bolt.new

  • Feta1,217

    Run smarter stand-ups, build better products

    AI · 2024 · feta.io

  • Trag955

    AI code review companion

    AI · 2024 · usetrag.com

  • Turn Notion databases into portals & apps with no code

    Dev tools · 2024 · softr.io

  • One inbox for all your work discussions

    Work · 2024 · generalcollaboration.com