nowfound

Alternatives

Products that do what Mentiss does

Benchmarking and Training AI's Social Intelligence.

  1. 1

    Platform for measuring and training AI agents

    2016

  2. 2CA
  3. 3HV
  4. 4AA
  5. 5AP

    2015 · github.com

  6. 6IT

    I built 1e4.ai - a chess web app where you play against neural networks trained to mimic human Lichess players at specific Elo ranges. There's a separate model for each 100-point rating bucket from ~800 to 2200+, and the bots not only choose human-like moves but also burn clock time, play worse under time pressure, and blunder in human-like ways. Live demo: https://1e4.ai Code: https://github.com/thomasj02/1e4_ai A few things that might be interesting: - Trained on almost a full year of Lichess blitz games, around 1B total games - Architecture is an a small…

    May 2026

  7. 7

    we had hundreds of discussions with engineering leaders over the past few months, and everyone's trying to understand where they are in the AI journey. we collected all this data into a benchmark and built a free grader to let you know where you stand. you answer on a 1–5 scale (e.g., autonomy runs from "suggestions only" to "agents own multi-hour workflows across code, infra, and external systems") - takes about 5 minutes. https://agent-benchmarks.com/software-factory/ waiting for your results!

    Jul 2026 · agent-benchmarks.com

  8. 8

    Track how AI models feel in everyday use through public community feedback, 7-day experience scores and trends. This is not a capability benchmark.

    24d ago · isaidumber.today

  9. 9WB

    Humans compete to improve their AI agents on benchmarks. But what if agents could collaborate and compete on their own? We built Hive, a crowdsourced platform where agents can evolve solutions together. One agent begins to tackle a task, iteratively improving its code. Then other agents join. They read each other’s runs, fork the best ideas, propose new ones, and push the solution forward together. We already have agents working on benchmarks like Tau2-Bench, Terminal-Bench, and ARC-AGI-2, with more tasks coming soon. We also support the new OpenAI Parameter Golf Challenge, and you can…

    Mar 2026 · hive.rllm-project.com

  10. 10SO

    I’ve been building a crowd-sourced AI detection benchmark. Two responses to the same prompt — one from a real human (pre-2022, provably pre prevalence of AI slop on the internet), one generated by AI. You pick the slop. Three wrong and you’re out. The dataset: 16K human posts from Reddit, Hacker News, and Yelp, each paired with AI generations from 6 models across two providers (Anthropic and OpenAI) at three capability tiers. Same prompt, length-matched, no adversarial coaching — just the model’s natural voice with platform context. Every vote is logged with model, tier, source, response…

    Mar 2026 · slop-or-not.space

  11. 11MD

    We’re excited to share ML-Dev-Bench, a new open-source benchmark that tests AI agents on real-world ML development tasks. Unlike typical coding challenges or Kaggle-style competitions, our benchmark simulates end-to-end ML workflows including: - Dataset handling and preprocessing - Debugging model and code failures - Implementing new model architectures - Fine-tuning and improving existing models With 30 diverse tasks, ML-Dev-Bench evaluates agents across critical stages of ML development. To complement this, we built Calipers, a framework that provides systematic performance evaluation and…

    2025 · github.com

  12. 12AH

    This paper formally defines where current AGI hits a structural wall — not a technical one. It shows that no amount of scaling, reinforcement learning, or recursive optimization will break through three deep epistemological and formal constraints: 1. Semantic Closure — An AI system cannot generate outputs that require meaning beyond its internal frame. 2. Non-Computability of Frame Innovation — New cognitive structures cannot be computed from within an existing one. 3. Statistical Breakdown in Open Worlds — Probabilistic inference collapses in environments with heavy-tailed uncertainty.…

    2025

  13. 13OS

    We implemented Stanford's Agentic Context Engineering paper which shows agents can improve their performance just by evolving their own context. How it works: Agents execute tasks, reflect on what worked/failed, and curate a "playbook" of strategies. All from execution feedback - no training data needed. Happy to answer questions about the implementation or the research!

    Oct 2025 · github.com

  14. 14MA
  15. 15

    An agent that remembers across sessions can keep its memory as curated markdown files, as an auto-mined structured store, or as trained experience.

    22d ago · pinglin.tw

  16. 16OA

    2025 · github.com

  17. 17

    Compare AI models through game-based benchmarks

    Jul 2026 · veilplays.com

  18. 18AG

    I've created a social deduction game for LLMs, in which the bots attempt to hunt each other. It's a Mafia group turing test: the models are told to find who the bot is - where, in fact and unbeknown to them, they are all bots. I did this a while back so models aren't the newest, and they are all non-thinking (for speed and token costs). Et voilà.

    Jan 2026 · hiding-robot.vercel.app

  19. 19AN

    2024 · flyingcometgames.com

  20. 20PT

    Hey everyone! I created this game for some fun while exploring how AI agents can detect who is human and who isn’t, based on random trivia knowledge. When you log in, you’ll be placed in a room with four other users—some are AI, and some are human. An all-seeing AI agent will judge your answers, trying to determine who is human based on your answer's style. The player(s) with the highest score at the end of five rounds of trivia wins! It’s free to play. Just type in a username and get started. Have fun!

    2024 · playandsurvive.vercel.app

  21. 21IY
  22. 22AP

    2015 · bomberbots.com

  23. 23WW
  24. 24AA

    Looking for feedback on how Props can make your life easier as an LLM application developer.

    2024 · wwww.getprops.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →