nowfound

Alternatives

Products that do what Inferbench, collect/share datapoints on GPU's inference performance does

Built a community-driven database for inference hardware. All results are submitted by users and validated by volunteers.

  1. 1

    AI models that run on an inference cloud optimized for speed

    May 2026 · generalcompute.com

  2. 2GA

    2021 · inferrd.com

  3. 3
    PHBench400

    Predict the next Series A from a ProductHunt launch

    May 2026 · phbench.com

  4. 4IV

    The video demo runs a 7b Model on a normal gaming GPU. I think it already works quite well (accounting for the limited hardware power). :)

    2024 · github.com

  5. 5
    ZeroGPU309

    The compute efficient layer for AI inference

    Jun 2026 · zerogpu.ai

  6. 6MD

    2012 · sourceforge.net

  7. 7AB

    I created a web page to compare different analytical databases (both self-managed and services, open-source and proprietary) on a realistic dataset. It contains 20+ databases, each with installation and data loading scripts. And they can be compared to each other on a set of 43 queries, by data load time or by storage size. There are switches to select different types of databases for comparison - for example, only MySQL compatible or PostgreSQL compatible. If you play with the switches, many interesting details will be uncovered. Full description:…

    2022 · benchmark.clickhouse.com

  8. 8
    Banana235

    Serverless GPUs for Machine Learning inference

    2022

  9. 9PI

    Deploying vision models is time consuming and tedious. Setting up dependencies. Fixing conflicts. Configuring TRT acceleration. Flashing (and re-flashing) NVIDIA Jetsons. A streamlined, developer-friendly solution for inference is needed. We, the Roboflow team, have been hard at work open sourcing Inference, an open source vision deployment solution. Our solution is designed with developers in mind, offering a HTTP-based interface. Run models on your hardware without having to write architecture-specific inference code. Here's a demo showing how to go from a model to GPU inference on a video…

    2023 · github.com

  10. 10

    Accelerating open machine learning research with Cloud TPUs

    2017

  11. 11

    Calculate the GPU memory you need for LLM inference

    2025

  12. 12

    Fast multimodal-native inference at scale

    Dec 2025

  13. 13WH
  14. 14TA
  15. 15TR
  16. 16DG
  17. 17
    GPU.LAND126

    Affordable cloud GPUs for deep learning

    2021

  18. 18TC

    Hello HN! I’m Jonathan from TensorDock. After 7 months in beta, we’re finally launching Core Cloud, our platform to deploy GPU virtual machines in as little as 45 seconds! https://www.tensordock.com/product-core Why? Training machine learning workloads at large clouds can be extremely expensive. This left us wondering, “how did cloud ever become more expensive than on-prem?” I’ve seen too many ML startups buy their own hardware. Cheaper dedicated servers with NVIDIA GPUs are not too hard to find, but they lack the functionality and scalability of the big clouds. We thought to…

    2022 · tensordock.com

  19. 19

    Benchmarks local LLM engines on your hardware

    15d ago · github.com

  20. 20CF
  21. 21VI

    Most inference UIs that I've come across pretty much just give us a chat-like interface to toy around with models in a single visual conversation thread. Given the fact that we are limited to seeing only one output at a time, it's kind of hard to compare outputs from different models, adjustments made to the prompting, and sampler settings. But even when keeping the generation parameters the same (e.g., to test for reliability in the output) and just going for multiple passes, there is no easy way to have a side-by-side comparison to keep track of the outputs from the multiple "rounds". I…

    2024 · github.com

  22. 22

    Easy to use and fairly priced GPUs for Machine Learning

    2019

  23. 23AG

    This is a vector index I built that supports insertion and k-nearest neighbors (k-NN) querying, optimized for GPUs. It operates entirely in CUDA and can process queries on half a billion vectors in under 200 milliseconds. The codebase is structured as a standalone library with an HTTP API for remote access. It’s intended for high-performance search tasks—think similarity search, AI model retrieval, or reinforcement learning replay buffers. The codebase is located at https://github.com/rodlaf/BinaryGPUIndex.

    2025 · rlafuente.com

  24. 24

    AI-Native Data Infrastructure for Spatial and Physical AI

    Apr 2026 · zibra.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →