nowfound

AI · August 12, 2026

TR

Trunchbull, run real models against any benchmark in your browser

Hi HN, Today I'm showcasing Trunchbull, a benchmarking platform designed for authoring benchmarks and running them against different models. We have direct support for benchmarks that use the harbor authoring system, custom tool authoring via the vercel ai sdk and configuration limits. We've also already imported terminalbench 2.0, as a sort of proof of concept that our harbor task orchestrator works, although you currently need a paid account as we are provisioning sandbox environments. I've made several popular benchmarks publicly available for testing. You dont need an account or your…

What it does

Provider-neutral AI model benchmarking. Detect regressions before they reach production.

Deploy tools from GitHub, run them against managed AI models, and inspect every decision under controlled budgets. Give us a GitHub URL. Trunchbull validates the contract, builds the repository, and deploys it into an isolated Cloudflare Worker—ready for any benchmark configuration. Run terminal and code-execution evaluations inside controlled, image-backed environments. Trunchbull exposes files, commands, processes, and interactive terminals while keeping each workload inside an isolated runtime. Define the maximum steps, token usage, and dollar spend before a run begins. Limits are enforced across model responses and tool calls—not checked after the damage is done. A score tells you what…from trunchbull.dev

In the maker’s words, at launch

Hi HN, Today I'm showcasing Trunchbull, a benchmarking platform designed for authoring benchmarks and running them against different models. We have direct support for benchmarks that use the harbor authoring system, custom tool authoring via the vercel ai sdk and configuration limits. We've also already imported terminalbench 2.0, as a sort of proof of concept that our harbor task orchestrator works, although you currently need a paid account as we are provisioning sandbox environments. I've made several popular benchmarks publicly available for testing. You dont need an account or your credit card information to access these: [Public benchmark demos](https://trunchbull.dev/sandboxes) [GSM8K math reasoning](https://trunchbull.dev/try/gsm8k) [SkateBench](https://trunchbull.dev/try/skatebench) [ARC-Challenge](https://trunchbull.dev/try/arc-challenge) [TruthfulQA](https://trunchbull.dev/try/truthfulqa-mc1) [Medical AI Failure Atlas](https://trunchbull.dev/try/medical-ai-failure-atlas) These demos let you pick from a preselected list of models, and will systematically test the selected models against their case scenarios. Any and all feedback is welcome, but i'm particularly interested in: - knowing what kind of benchmark evidence youd like to expect - any improvements on our benchmark run page, anything that can provide clarity or better understanding of the benchmark u just ran. - what you'd prefer to see on the overview page. - improvements on our documentation - better configuration and spend limits.

Does the same job

all alternatives →
  • oqoqo27d ago · oqoqo.ai · ▲340

    Build evals and custom benchmarks for real-world tasks

  • APIEval-20May 2026 · ▲121

    An open benchmark for AI agents that test APIs

  • BenchspanMar 2026 · ▲85

    Run agent benchmarks in minutes, not hours

  • TW
    Terminal-Wrench, a dataset of 331 realistic hackable environmentsApr 2026 · github.com · ▲6

    I want to share a new dataset of 331 reward-hackable environments. These are real environments used in Terminal Bench and adjacent benchmarks. I first got interested in this because, as a reviewer of Terminal Bench, I noticed a lot of our tasks were hackable. I also noticed that many contributors to the benchmark do so because it provides credibility when selling environments to labs. Hence, TBench tasks are, in my opinion, held to a higher quality standard than those being used today for RL. No one is spending hours manually reviewing the $1B in tasks being purchased by major labs. As far…

  • OH
    OpenBenchmarks – Helping agents discover and pick the right SaaS APIsJul 2026 · openbenchmarks.com · ▲6

    I'm Fenil, co-founder/CEO of OpenFunnel (YC F24), building this with my co-founder/CTO Aditya. We're launching OpenBenchmarks (https://openbenchmarks.com), open-source, reproducible benchmarks for SaaS APIs, starting with the category we know best: GTM APIs. ## Why we built this More and more B2B software evaluation will/already runs through reasoning models inside agentic workflows rather than through people. And buyers increasingly pick vendors that are API-first and ship MCPs, so they can wire them into internal workflows. Strong reasoning models are skeptical of…

  • AR
    A registry of agent benchmarks (including many OSS agent trajectories)2024 · explorer.invariantlabs.ai · ▲6

    If you're interested in exploring what LLM-based agent systems these days actually do to solve certain benchmarks such as SWEBench or WebArena, we created a small leaderboard with our team, that allows to view a lot of public and OSS agent results including all the runtime traces (the step-by-step reasoning behind the scenes). Looking at traces is actually quite interesting, as they reveal a lot about the inner working and shortcomings of current agent system, e.g. see https://explorer.invariantlabs.ai/u/invariant/webarena--SteP... for an example trace.

More ai this month

the category →
  • I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.

    AI · 17d ago · simedw.com

  • Astute585

    Automate your B2B brand going viral, with new media creators

    AI · 18d ago · company-app.joinastute.com

  • Grok Bot547

    AI teammates that you can give real work to

    AI · 25d ago · x.ai

  • Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    AI · 27d ago · cactuscompute.com

  • Turn website visitors into qualified pipeline

    AI · 19d ago · clarasdr.ai

  • Kane CLI446

    Natural language browser & mobile app tests from terminal

    AI · 24d ago · testmuai.com

Launched alongside, August 2026

the whole month →
  • TL

    Life & fun · 10d ago · louisabraham.github.io

  • Hey Noah641

    A proactive AI executive assistant for founders

    AI · Aug 2026 · heynoah.io

  • Let agents source clips from terabytes of your local video

    Work · 18d ago · clipto.com

  • SA

    Hello HN! I found that picking out plausible but diverse skin tones for my digital art and game development projects was kind of difficult, and I got curious about if there was a way to define a color space that made it easy. I've built a color picker and procedural generation algorithm based on the space as well as a bunch of other fun js features and demos throughout the page that use the equations. If you find it interesting, I have lots of explanations of how I built it and what properties the space has. The methodology might be a bit shaky, but hopefully the result is as helpful for…

    Life & fun · Aug 2026 · toneyalexander.github.io

  • AdAnt AI608

    Claude for viral, high-converting social ads

    AI · Aug 2026 · adant.ai

  • I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.

    AI · 17d ago · simedw.com