nowfound

AI · March 21, 2025

BR

BenchFlow – run AI benchmarks as an API

I built BenchFlow, an open-source framework that lets you integrate and evaluate AI tasks using Docker-based benchmarks. You can try it out right now by cloning the repo and running a benchmark in minutes. As an AI researcher, I was frustrated with how much time my team spent setting up benchmark environments rather than actually improving our models. We'd spend weeks configuring environments, only to find inconsistencies when comparing results with other teams. BenchFlow started as an internal tool to standardize our evaluation process, and we decided to open-source it after seeing how much…

In plain words

BenchFlow is an open-source framework that lets AI researchers integrate and evaluate AI tasks using Docker-based benchmarks accessible as an API. It standardizes benchmark environments across different machines and teams, reducing setup time and inconsistencies in results. Users can run benchmarks in minutes by cloning the repository. The tool provides a unified interface for any AI task and eliminates dependency conflicts through containerization, addressing the common problem of researchers spending weeks on environment configuration instead of model improvement.

written from the facts on this page · September 2026

From the sources

In the maker’s words, at launch

I built BenchFlow, an open-source framework that lets you integrate and evaluate AI tasks using Docker-based benchmarks. You can try it out right now by cloning the repo and running a benchmark in minutes. As an AI researcher, I was frustrated with how much time my team spent setting up benchmark environments rather than actually improving our models. We'd spend weeks configuring environments, only to find inconsistencies when comparing results with other teams. BenchFlow started as an internal tool to standardize our evaluation process, and we decided to open-source it after seeing how much time it saved us. Unlike other benchmarking tools that focus on specific domains, BenchFlow provides a unified interface for any AI task. The Docker-based approach ensures consistent environments across different machines and teams. You don't need to worry about dependency conflicts or environment setup - just implement a simple interface and you're ready to go. How to try it out? check our link but here's a preview of that 1. pip install benchflow 2. load a benchmark and define how to call your agents/models 3. run it and get the result Available benchmarks you can try today: - MMLU-PRO: Test your model's knowledge across 57 subjects - Bird: Evaluate business intelligence reasoning capabilities - WebArena: See how your agent performs on web-based tasks - MedQA-CS: Test medical question answering abilities The framework handles all the containerization, task distribution, and result collection, so you can focus on improving your models rather than managing infrastructure. I'd love to hear your feedback and see how you use it. What benchmarks would you like to see added next? Please give us a star if you can, thanks! GitHub: https://github.com/benchflow-ai/benchflow Website: https://benchflow.ai/ Benchmark Hub: https://benchflow.ai/benchmarks Inspo: https://github.com/ServiceNow/BrowserGym

More ai this month

the category →
  • I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.

    AI · 17d ago · simedw.com

  • Astute585

    Automate your B2B brand going viral, with new media creators

    AI · 18d ago · company-app.joinastute.com

  • Grok Bot547

    AI teammates that you can give real work to

    AI · 25d ago · x.ai

  • Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    AI · 26d ago · cactuscompute.com

  • Make your software self-driving

    AI · 30d ago · coldtea.ai

  • Soloop472

    Approval-first Agent OS for solo founders

    AI · 30d ago · soloop.io

Launched alongside, March 2025

the whole month →
  • Mimic Human Research & Save Findings in AI Knowledge Base

    AI · 2025 · sider.ai

  • The first AI dev team

    AI · 2025 · atoms.dev

  • Aha1,151

    The world's first AI influencer marketing team

    AI · 2025 · ahacreator.com

  • Fluently976

    Start speaking English as well as your native language

    AI · 2025 · getfluently.app

  • Conversational AI surveys, interviews, user tests, polls

    AI · 2025 · theysaid.io

  • Record your screen, share instantly, look like a PRO

    Growth · 2025 · supercut.ai