nowfound

Alternatives

Products that do what AILatency does

Independent AI API performance benchmarks

  1. 1UD

    Hey HN! I’m the founder of Unify, and we’ve just released our Model Hub, which provides a collection of LLM endpoints with live runtime benchmarks all plotted across time: https://unify.ai/hub A key finding is that static tabular runtime benchmarks for LLMs simply do not work. It’s necessary to take a time-series perspective, and plot the variations through time. We currently have 21 models provided by: Anyscale, Perplexity AI, Replicate, Together AI, OctoAI, Mistral AI and OpenAI, with more on the roadmap. We test across different regions (Asia, US, Europe), with varied…

    2024

  2. 2

    Trace, evaluate, and improve AI agents in production

    Aug 2026 · telerik.com

  3. 3
    Web Bench138

    A 10x better benchmark for AI browser agents

    2025

  4. 4AT

    I recently built a small open-source tool to benchmark different LLM API endpoints — including OpenAI, Claude, and self-hosted models (like llama.cpp). It runs a configurable number of test requests and reports two key metrics: • First-token latency (ms): How long it takes for the first token to appear • Output speed (tokens/sec): Overall output fluency Demo: https://llmapitest.com/ Code: https://github.com/qjr87/llm-api-test The goal is to provide a simple, visual, and reproducible way to evaluate performance across different LLM providers, including…

    2025 · llmapitest.com

  5. 5
    LLM Stats308

    Compare API models by benchmarks, cost & capabilities

    Oct 2025

  6. 6

    Measure the full AI SDLC. From token to production.

    Apr 2026 · waydev.co

  7. 7

    Is your data ready for AI, find out now

    2022

  8. 8

    An open benchmark for AI agents that test APIs

    May 2026 · resources.kusho.ai

  9. 9

    Open-source monitoring for machine learning models

    2021

  10. 10
    cto bench125

    The ground truth code agent benchmark

    Dec 2025

  11. 11BR

    I built BenchFlow, an open-source framework that lets you integrate and evaluate AI tasks using Docker-based benchmarks. You can try it out right now by cloning the repo and running a benchmark in minutes. As an AI researcher, I was frustrated with how much time my team spent setting up benchmark environments rather than actually improving our models. We'd spend weeks configuring environments, only to find inconsistencies when comparing results with other teams. BenchFlow started as an internal tool to standardize our evaluation process, and we decided to open-source it after seeing how much…

    2025 · github.com

  12. 12

    Build the semantic layer that makes AI analytics trustworthy

    Mar 2026 · metabase.com

  13. 13LL

    Hey Folks! I've been building an open source benchmark for measuring local LLM performance on your own hardware. The benchmarking tool is a CLI written on top of Llamafile to allow for portability across different hardware setups and operating systems. The website is a database of results from the benchmark, allowing you to explore the performance of different models and hardware configurations. Please give it a try! Any feedback and contribution is much appreciated. I'd love for this to serve as a helpful resource for the local AI community. For more check out: - Website:…

    2025 · localscore.ai

  14. 14

    AI platform for accurate, standard API & endpoint versioning

    2024

  15. 15

    Run agent benchmarks in minutes, not hours

    Mar 2026 · benchspan.com

  16. 16BY

    we had hundreds of discussions with engineering leaders over the past few months, and everyone's trying to understand where they are in the AI journey. we collected all this data into a benchmark and built a free grader to let you know where you stand. you answer on a 1–5 scale (e.g., autonomy runs from "suggestions only" to "agents own multi-hour workflows across code, infra, and external systems") - takes about 5 minutes. https://agent-benchmarks.com/software-factory/ waiting for your results!

    Jul 2026 · agent-benchmarks.com

  17. 17

    A real-time dashboard of email provider performance

    Mar 2026

  18. 18

    Live AI benchmarks, drift alerts, and smart model routing

    7d ago · aistupidlevel.info

  19. 19

    Easy goverment data access for citizens, optimized for AI

    Apr 2026 · katzilla.dev

  20. 20

    Trace AI requests, workflows, and costs in one timeline

    May 2026

  21. 21

    Data APIs for AI agents & developers — free to start

    Feb 2026

  22. 22MD

    We’re excited to share ML-Dev-Bench, a new open-source benchmark that tests AI agents on real-world ML development tasks. Unlike typical coding challenges or Kaggle-style competitions, our benchmark simulates end-to-end ML workflows including: - Dataset handling and preprocessing - Debugging model and code failures - Implementing new model architectures - Fine-tuning and improving existing models With 30 diverse tasks, ML-Dev-Bench evaluates agents across critical stages of ML development. To complement this, we built Calipers, a framework that provides systematic performance evaluation and…

    2025 · github.com

  23. 23

    Source-backed research protocol for AI agents

    Jul 2026 · rrrrrredy.github.io

  24. 24

    Benchmark AI models for YOUR use case

    Jan 2026

Ranked by how close each launch is in meaning, then by votes. Refine with a description →