nowfound

Alternatives

Products that do what AIhumanbench does

Measuring the Gap Between Silicon and Soul.

  1. 1
    AI Titans553

    Prove your brains, brawn, and bytes

    2023

  2. 2
    Web Bench138

    A 10x better benchmark for AI browser agents

    2025

  3. 3
    TestAI407

    1,000+ automated tests for AI agents in one click

    2025

  4. 4
    debate AI297

    debate smarter with AI

    2024

  5. 5
    cto bench125

    The ground truth code agent benchmark

    Dec 2025

  6. 6

    Get real-world tasks done with autonomous AI agents

    Jun 2026 · arena.ai

  7. 7

    A live demo of AI algorithms making judgments about you.

    2020

  8. 8

    The voice AI auto-testing loop to simulate, evaluate & ship

    2025

  9. 9

    Validate human and AI-generated creative assets in 2 hours

    2023

  10. 10

    Can you guess which image is of a real person vs AI?

    2019

  11. 11

    Compare AI models side-by-side on same prompt

    Feb 2026 · testaimodels.com

  12. 12PT

    Hey everyone! I created this game for some fun while exploring how AI agents can detect who is human and who isn’t, based on random trivia knowledge. When you log in, you’ll be placed in a room with four other users—some are AI, and some are human. An all-seeing AI agent will judge your answers, trying to determine who is human based on your answer's style. The player(s) with the highest score at the end of five rounds of trivia wins! It’s free to play. Just type in a username and get started. Have fun!

    2024 · playandsurvive.vercel.app

  13. 13
    IdeaScore121

    AI-powered idea testing and refinement

    2023

  14. 14CB

    Hey HN, we're excited to share Cua-Bench ( https://github.com/trycua/cua ), an open-source framework for evaluating and training computer-use agents across different environments. Computer-use agents show massive performance variance across different UIs—an agent with 90% success on Windows 11 might drop to 9% on Windows XP for the same task. The problem is OS themes, browser versions, and UI variations that existing benchmarks don't capture. The existing benchmarks (OSWorld, Windows Agent Arena, AndroidWorld) were great but operated in silos—different harnesses,…

    Jan 2026 · github.com

  15. 15
    Reason857

    User research on autopilot

    Apr 2026 · reason8.io

  16. 16

    Learn about, try & give feedback on Google’s emerging AI

    2023

  17. 17ΤB

    τ-Bench is an open benchmark for evaluating AI agents on grounded, multi-turn customer service tasks with verifiable outcomes. It's been great to see the community adopt it since launch — this is now the third iteration. With τ³-Bench, we're extending it to two new settings: knowledge-intensive retrieval and full-duplex voice. τ-Knowledge: agents must navigate ~700 interconnected policy documents to complete multi-step tasks. Best frontier model (GPT-5.2, high reasoning) hits ~25%. The surprising part: even when you hand the model the exact documents it needs, performance only reaches ~40%.…

    Mar 2026

  18. 18

    Roast or Toast? Let AI Rate Your Squad's Look challenge

    2025

  19. 19MD

    We’re excited to share ML-Dev-Bench, a new open-source benchmark that tests AI agents on real-world ML development tasks. Unlike typical coding challenges or Kaggle-style competitions, our benchmark simulates end-to-end ML workflows including: - Dataset handling and preprocessing - Debugging model and code failures - Implementing new model architectures - Fine-tuning and improving existing models With 30 diverse tasks, ML-Dev-Bench evaluates agents across critical stages of ML development. To complement this, we built Calipers, a framework that provides systematic performance evaluation and…

    2025 · github.com

  20. 20

    Make studying fun with the latest Gen Z slang

    2024

  21. 21

    The World's #1 Vibe Coding Benchmark

    Jun 2026 · woaibench.ai

  22. 22
    Spec2734

    Spec-driven testing for AI agents and AI apps

    Apr 2026 · spec27.ai

  23. 23BO

    My cofounder and I were playing with AI voices from different vendors and were shocked at how good the results were. We decided to make a game so you can test your own ability to tell humans from AI voices. This has helped us narrow down which voices we'd like to use for an upcoming project—perhaps you will find it useful and/or interesting too! Would love to hear any feedback you have!

    2024 · bot.unison.fm

  24. 24IT

    I built 1e4.ai - a chess web app where you play against neural networks trained to mimic human Lichess players at specific Elo ranges. There's a separate model for each 100-point rating bucket from ~800 to 2200+, and the bots not only choose human-like moves but also burn clock time, play worse under time pressure, and blunder in human-like ways. Live demo: https://1e4.ai Code: https://github.com/thomasj02/1e4_ai A few things that might be interesting: - Trained on almost a full year of Lichess blitz games, around 1B total games - Architecture is an a small…

    May 2026

Ranked by how close each launch is in meaning, then by votes. Refine with a description →