nowfound

Alternatives

Products that do what Groq® does

Hyperfast LLM running on custom built GPUs

  1. 1

    The world’s most powerful chip’ for AI

    2024

  2. 2

    Fast multimodal-native inference at scale

    Dec 2025

  3. 3

    Calculate the GPU memory you need for LLM inference

    2025

  4. 4
    Banana235

    Serverless GPUs for Machine Learning inference

    2022

  5. 5
    Mercury 2152

    Fastest reasoning LLM built for instant production AI

    Feb 2026

  6. 6

    Powers faster, efficient reasoning for long-running agents

    Jun 2026

  7. 7
    Mu126

    Fast, local AI comes to Windows Copilot+ PCs

    2025

  8. 8

    Fast LLMs for low-latency and high-performance workflows

    Jun 2026

  9. 9
    MTIA v2116

    Meta training and inference accelerator

    2024

  10. 10

    Pool compute to run powerful open models

    Apr 2026

  11. 11
    RunInfra156

    Describe the AI model you need and get an optimized AI

    Jul 2026

  12. 12
    GPU.LAND126

    Affordable cloud GPUs for deep learning

    2021

  13. 13

    The fastest generative AI Text-to-Speech API

    2023

  14. 14GA

    2021 · inferrd.com

  15. 15

    Meta's 3rd-gen custom AI chips for GenAI inference

    Mar 2026

  16. 165L

    We've built InferX, a specialized runtime environment that fundamentally changes how LLMs are served. The core problem we solve is the latency bottleneck in AI inference, especially with large models. Current systems waste resources or suffer from painfully slow cold starts. InferX's AI-native architecture, with its "snapshot" technology, enables: * *Sub-2s cold starts:* Spin up models instantly. * *High density:* Serve more LLMs on the same GPUs. * *Optimal efficiency:* Maximize GPU utilization. This isn't just another API; it's a new execution layer designed from the ground up for the…

    2025 · github.com

  17. 17

    Run AI jobs from your IDE with a one-click workflow

    Mar 2026

  18. 18RA

    We built RapidFire AI, an open-source Python tool to speed up LLM fine-tuning and post-training with a powerful level of control not found in most tools: Stop, resume, clone-modify and warm-start configs on the fly—so you can branch experiments while they’re running instead of starting from scratch or running one after another. - Works within your OSS stack: PyTorch, HuggingFace TRL/PEFT), MLflow. - Hyperparallel search: launch as many configs as you want together, even on a single GPU - Dynamic real-time control: stop laggards, resume them later to revisit, branch promising configs in…

    Sep 2025 · github.com

  19. 19

    Turn idle GPUs into cash. Get affordable AI for everyone.

    Nov 2025

  20. 20BO

    Read the full blogpost at https://rach.codes/blog/Introducing-Bhumi (click on reader to see the technical breakdown!) AI inference should be fast, but in practice it’s painfully slow. Inference bottlenecks slow down LLM-powered chatbots and AI workflows everywhere. I built Bhumi to fix that. Bhumi is a Python library designed for developers, yet its performance-critical core is implemented in Rust (via PyO3) for near-native speed. This hybrid approach delivers up to 2.5x faster response times across providers like OpenAI, Anthropic, and Gemini—without changing the…

    2025 · bhumi.trilok.ai

  21. 21CR

    hi everyone. how does moving llm call prompts and output structure definitions away from code into configuration land sound? would you use something like this if it was stable and well documented enough? please don't hold back the criticism. i appreciate all feedback (constructive & otherwise).

    2024 · github.com

  22. 22CT

    I had been looking to try <500M parameter language models but you wouldn't find an API to try them anywhere, so I built this cloudflare hosted static website that hosts weights and built an inference runtime for these models that uses WebGPU and runs inference from your browser. These are only so useful in a multi-turn conversation but it's still interesting to see what you can pack in a <250mb model. I tried using ONNX versions earlier, but there were too many quirks of using them with language models and the TPS wasn't too impressive. Inspired by svenflow&#x2F;webgpu-gemma, I put my codex…

    May 2026 · chonklm.com

  23. 23LH

    I work on inference scheduling — KV cache-aware routing, load balancing across GPU workers, that kind of thing. I wanted something like k9s but for my inference stack. Nothing existed, so I built it. llmtop is a real-time terminal dashboard for LLM inference workers. It scrapes the Prometheus &#x2F;metrics endpoints that vLLM, SGLang, and LMCache already expose and shows everything in one view: KV cache usage, queue depth, TTFT&#x2F;ITL latencies (P50&#x2F;P99 from histogram buckets), token throughput, prefix cache hit rates. Color-coded — red means go fix it. ``` brew install…

    Mar 2026 · github.com

  24. 24

    Unified Inference Stack with multi cloud GPU orchestration

    Dec 2025

Ranked by how close each launch is in meaning, then by votes. Refine with a description →