nowfound

Alternatives

Products that do what LLM Debate Benchmark does

  1. 1AN
  2. 2AA

    2015 · en.arguman.org

  3. 3AD

    Ever wish you could get the best arguments for both sides of a debate? I built an AI-powered debate platform that pits language models against each other on controversial topics. Each AI is randomly assigned a side (pro/con). You vote before and after to see if you were persuaded. Most content today presents lopsided arguments. They provide strong points for one side, weak ones for the other. This project aims to surface the strongest arguments from both sides, using LLMs to simulate a fair debate. With enough usage, I want to use it to benchmark LLMs. My hypothesis is that randomly…

    2025 · bot-bicker.vercel.app

  4. 4LB
  5. 5CL

    2023 · convoclash.net

  6. 6CD
  7. 7PA
  8. 8BA
  9. 9BD

    2017 · bigquery.cloud.google.com

  10. 10TF

    2011 · barkles.com

  11. 11DA

    2017 · debategate.net

  12. 12

    Collaborative Argument Tree for Significant Debates

    2014

  13. 13LT
  14. 14LD
  15. 15AA

    2015 · arguman.org

  16. 16LD
  17. 17FC
  18. 18LC
  19. 19LR

    Sep 2025 · github.com

  20. 20

    3d ago · redactle.net

  21. 21DS
  22. 22DU
  23. 23PP

    I've been working on applying LLMs to long-context, verifiable problems over the past year, and today I'm releasing a benchmark of 62,000 pencil puzzles across 94 types (sudoku, nonori, slitherlink, etc.). The benchmark also allows for intermediate checks /rule breaks for all varieties at any step. I tested 51 models against a subset (300 puzzles) in two modes: single-shot (output the full solution) and agentic (iterate with verifier feedback). Some results: - Best model (GPT 5.2@xhigh) solves 56%. (~ half the puzzles are unsolved by any model) - Agentic solves average 29 turns. The…

    Mar 2026 · ppbench.com

  24. 24MC

Ranked by how close each launch is in meaning, then by votes. Refine with a description →