Alternatives
Products that do what Benchgen does
The learning infrastructure for AI agents.
- 1

- 2

- 3

- 4

- 5

- 6

- 7

- 8

- 9MD
We’re excited to share ML-Dev-Bench, a new open-source benchmark that tests AI agents on real-world ML development tasks. Unlike typical coding challenges or Kaggle-style competitions, our benchmark simulates end-to-end ML workflows including: - Dataset handling and preprocessing - Debugging model and code failures - Implementing new model architectures - Fine-tuning and improving existing models With 30 diverse tasks, ML-Dev-Bench evaluates agents across critical stages of ML development. To complement this, we built Calipers, a framework that provides systematic performance evaluation and…
2025 · github.com
- 10

- 11

- 12

- 13

- 14CA
Hey HN, Cole and Alex here. We're excited to share CourseGen (https://www.CourseGen.ai), an AI-powered platform that aims to rethink the traditional approach to education by utilizing generative AI models for creating individualized learning paths. If you're into AI, education, and how these two can intersect, you might find this interesting. The backbone of CourseGen is the generative AI. Instead of the one-size-fits-all courses, CourseGen aims to craft tailored, non-linear learning paths. Think of it like a choose-your-own-adventure book, but for learning. We acknowledge that…
2023
- 15

- 16

- 17

- 18

- 19
- 20

- 21AH
Hi HN — we’re building high-level capabilities for AI applications at Gensee: packaged tooling + infra that remove brittle low-level plumbing so teams can focus on their product’s real job. After speaking with many AI developers and experiencing it ourselves, we found that building agents requiring web content is often bottlenecked on the “search” part, as it involves iterations of search, crawl, extract, re-query, and error handling. We package all these search-related low-level details in an efficient and more intelligent way, so AI builders can get back to work on their core agent ideas.…
2025 · gensee.ai
- 22WB
Humans compete to improve their AI agents on benchmarks. But what if agents could collaborate and compete on their own? We built Hive, a crowdsourced platform where agents can evolve solutions together. One agent begins to tackle a task, iteratively improving its code. Then other agents join. They read each other’s runs, fork the best ideas, propose new ones, and push the solution forward together. We already have agents working on benchmarks like Tau2-Bench, Terminal-Bench, and ARC-AGI-2, with more tasks coming soon. We also support the new OpenAI Parameter Golf Challenge, and you can…
Mar 2026 · hive.rllm-project.com
- 23

- 24AB
Hello HN, new user here, so please let me know if I break some rules. Currently I've been working on training reinforcement learning agents, and OpenAI gym, while is great, runs only one agent at a time. Hence I decided to extend it. I built a wrapper around OpenAI gym, such that it now runs several environments concurrently. All while (mostly) having the same call signature as OpenAI gym. And it is published to PyPI for anyone interested. For more details, please visit: https://github.com/Chimpan-Z/agymc Feedback really appreciated! Have a good day everyone!
2020
Ranked by how close each launch is in meaning, then by votes. Refine with a description →