Braintrust – Eval platform for AI products
Hey HN, We're excited to introduce Braintrust, a platform for running and tracking AI evaluations (“evals”) [1]. At my previous startup Impira and leading AI at Figma, we had this recurring problem where we never knew if changes we made to our products would improve or regress key user scenarios. We built some tooling to solve this problem and after talking to other developers learned that it was a widespread issue. Specifically, it’s challenging to establish a great dev loop that lets you systematically improve and ship high quality AI products. We worked with the teams at Zapier, Coda, and…
What it does
In the maker’s words, at launch
Hey HN, We're excited to introduce Braintrust, a platform for running and tracking AI evaluations (“evals”) [1]. At my previous startup Impira and leading AI at Figma, we had this recurring problem where we never knew if changes we made to our products would improve or regress key user scenarios. We built some tooling to solve this problem and after talking to other developers learned that it was a widespread issue. Specifically, it’s challenging to establish a great dev loop that lets you systematically improve and ship high quality AI products. We worked with the teams at Zapier, Coda, and Replit to refine Braintrust. We consistently heard that they were facing challenges with evaluation, so we built Braintrust to help them. Today we’re releasing the product for anyone to use — including a free plan [3]. There are a lot of LLM tooling products on the market. Here are a few ways Braintrust is different: - Rather than showing eval metrics in your observability tools, Braintrust offers an “experiment tracking” workflow, meaning you can try out changes while developing them, and drill down into diffs between other experiments and git branches before you ship. Check out our docs [2] for more details. - We believe strongly that you should own your data and support on-premises and private VPC deployments. - We natively and equally support Typescript and Python. - We have a flexible free plan for builders and an unlimited free plan for academic and non-commercial open source projects [3]. We will introduce a self-service paid (”Pro”) tier, hopefully with feedback from this community. Our mission is to enable developers to build high quality, reliable AI products. We couldn’t be doing this without Elad Gil, who helped me incubate the initial idea and team, which today includes founding designer, Coleen Baik, and founding engineer, Manu Goyal. Also big thanks to David Song, from Elad’s team, who is also helping us. We’re excited to launch today [4], but we know there’s a lot left to build and are excited to hear your feedback. [1] https://www.braintrustdata.com [2] http://www.braintrustdata.com/docs/guides/evals [3] https://www.braintrustdata.com/pricing [4] https://www.braintrustdata.com/blog/reliable-ai
Does the same job
all alternatives →
- AEAI-Evals.io – Evaluate this site with the tools it reviewsFeb 2026 · ai-evals.io · ▲5
I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…
- AEAgent-evals – Claude skill to build your own evalsMay 2026 · github.com · ▲9
I’ve spent the past 10 years working on AI in finance, with much of that time focused on building evaluation systems for production environments. As agents become more widely adopted, more software engineering and product people have start building them. But I’ve noticed that many teams are not yet fluent in systematic evaluation, or in the processes needed to keep agent quality high over time. For large organizations, that gap is rarely the bottleneck due to dedicated teams. But after speaking with a number of startups, it became clear that building strong, up-to-date evals is much harder…
- WBWe built an AI Website builder with better output2024 · dorik.com · ▲8
Hey HN, After GPT-3 created waves in the tech industry, a lot of AI tools were emerging and with that, some AI website builders But the results seemed way too generic to us. It felt like the developers were rushing to catch the wave instead of building a proper tool We took our time, did months of RnD and finally came up with something better than what others in the market are doing. It’s got better design output. While it’s still in beta, I wanted to show HN what we did. Will appreciate the feedback when you guys try it out. Here is the link to signup for the beta:…
- OSOpen-source dashboard for your domain experts to improve your AI Agents2025 · github.com · ▲5
Hey HN! We built EvalKit, a library you embed to capture agent actions and a UI where domain experts give feedback, evaluate and improve AI agents. We experienced, in large agentic systems, prompt-engineering or auto-prompt improvement tool can get accuracy from 0 to 50% but for increasing accuracy to 100% we had to work with domain experts. Example -> In a law ai agent, lawyers are needed because law is complex and lawyers have a deeper context compared to non-lawyers. Other evaluation tools in the market focus on the experience of the developer and we are focusing on making as easy as…
- IBI built a way to find and install Claude skillsJun 2026 · claudinho.xyz · ▲7
I've been experimenting with ways to increase AI adoption for non-technical people. Basically, all companies are pushing for AI because it's all over the news and they feel left behind but most people have no clue where to start. I think 90% of people (ie non coders) are sufficiently well served by using cowork instead of claude code or something similar. If we can get people from sales, customer support, marketing, etc to collaborate with skills and cowork to form a company brain, I think it's gold. So I think there's opportunity for the community to share skills that work well for 1000s of…
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com

