nowfound

Alternatives

Products that do what OpenBench does

Benchmark local LLMs without living in the terminal.

  1. 1

    Test-driven development for LLMs

    2023

  2. 2
    LM Studio209

    Discover, download, and run local LLMs (incl. DeepSeek R1)

    2025

  3. 3

    Trace LLM requests + costs with OpenTelemetry monitoring

    Oct 2025

  4. 4

    Open-source benchmarks for cloud browser infrastructure

    Apr 2026

  5. 5

    Live SaaS metric benchmarks from over 600 companies

    2016

  6. 6
    Openlit152

    One click observability & evals for LLMs & GPUs

    2024

  7. 7

    Benchmarks local LLM engines on your hardware

    14d ago · github.com

  8. 8
    ChattyUI149

    Run open-source LLMs locally in the browser using WebGPU

    2024

  9. 9
    Arkor142

    Fine-tune and Deploy Open-weight Models in TypeScript

    Jul 2026

  10. 10IB

    Built a simple web app that tells you which open-source LLMs will work on your hardware. It auto-detects your specs, shows compatible models from Hugging Face, gives realistic performance estimates (tokens/sec), and recommends quantization settings. You can also manually input specs to see "what if I upgraded my RAM?" Made this after wasting time downloading giant models only to find they crawled on my hardware. Hope it saves you some frustration!

    2025 · caniusellm.com

  11. 11TR

    Hi HN, Today I'm showcasing Trunchbull, a benchmarking platform designed for authoring benchmarks and running them against different models. We have direct support for benchmarks that use the harbor authoring system, custom tool authoring via the vercel ai sdk and configuration limits. We've also already imported terminalbench 2.0, as a sort of proof of concept that our harbor task orchestrator works, although you currently need a paid account as we are provisioning sandbox environments. I've made several popular benchmarks publicly available for testing. You dont need an account or your…

    24d ago · trunchbull.dev

  12. 12

    LLM Provider arbitrage to get the best performance for the $

    2025

  13. 13LT

    Measures the ability of various LLMs to navigate a fictional codebase via iterative directory tree expansion and observation. Each model's baseline ability is compared against combinations of various prompt engineering mods to quantify exactly how much they help or hinder the LLM. Interesting findings here: https://github.com/aiwebb/treenav-bench#interesting-findings

    2024 · github.com

  14. 14LB
  15. 15

    Reproducible benchmarks for evaluating AI models

    12d ago · github.com

  16. 16SA

    Writeup: https://www.deeptempo.ai/blogs/the-36-percent-false-positive...

    Jul 2026 · github.com

  17. 17BI
  18. 18IB

    I was overspending on GPT-4o. It was really hard to compare different models I could switch to, so I built this LLM comparison tool. It shows leaderboards, pricing, and performance data across 100+ LLMs (including all major providers and open-source models). Key features: - Live pricing comparisons - Benchmark Scores (MMLU, HumanEval, GPQA, etc.) - Context length vs cost analysis - Speed/throughput tests across providers - Quality vs price visualizations - Open source (all data verifiable) Try it out: https://llmstats.com I'd like to know your opinion :) Tech stack: Next.js,…

    2025 · llm-stats.com

  19. 19GB

    Hey HN, We’re excited to share PySpur, an open-source tool that provides a graph-based interface for building, debugging, and evaluating LLM workflows. Why we built this: Before this, we built several LLM-powered applications that collectively served thousands of users. The biggest challenge we faced was ensuring reliability: making sure the workflows were robust enough to handle edge cases and deliver consistent results. In practice, achieving this reliability meant repeatedly: 1. Breaking down complex goals into simpler steps: Composing prompts, tool calls, parsing steps, and branching…

    2024 · github.com

  20. 20OH

    I'm Fenil, co-founder/CEO of OpenFunnel (YC F24), building this with my co-founder/CTO Aditya. We're launching OpenBenchmarks (https://openbenchmarks.com), open-source, reproducible benchmarks for SaaS APIs, starting with the category we know best: GTM APIs. ## Why we built this More and more B2B software evaluation will/already runs through reasoning models inside agentic workflows rather than through people. And buyers increasingly pick vendors that are API-first and ship MCPs, so they can wire them into internal workflows. Strong reasoning models are skeptical of…

    Jul 2026 · openbenchmarks.com

  21. 21CB

    I built a small benchmark to test CLI coding agents on blind bug detection. A challenger agent injects bugs and writes ground truth (`bugs.json`). A different reviewer agent audits the repo without seeing ground truth, and an LLM matcher scores bug-to-finding assignments. Current run: 50 repos, 150 challenges, 450 reviews, 2,603 injected bugs. Weighted detection: Claude 58.05%, Codex 37.84%, Gemini 27.81%. LLM-judge benchmarks are easy to get wrong, so I’d really appreciate critical feedback on benchmark fairness, scoring/matching methodology, and obvious failure modes I’m missing. Full…

    Feb 2026 · github.com

  22. 22FR

    2019 · app.f0cal.com

  23. 23

    Competitive software benchmarks grounded in public evidence

    11d ago · signalbench.win

  24. 24OB

    Today, we're launching the Open Benchmarks Grants: a $3M commitment to fund open-source and academic teams building benchmarks for AI agents. In partnership with HuggingFace, PrimeIntellect, FactoryHQ, Together, Harbor, and PyTorch, the grants provide funding, data development support, and research collaboration. Our ability to measure AI has been outpaced by our ability to develop it, and we believe this evaluation gap is one of the most important problems in AI. Open benchmarks are one of the most important levers for advancing AI safely and responsibly—but the academic and open-source…

    Feb 2026 · benchmarks.snorkel.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →