Alternatives
Products that do what Watch LLMs play 21,000 hands of Poker does
PokerBench is my attempt at a new LLM benchmark wherein frontier models play Texas Hold'em in an arena setting. It also features a simulator to view individual games and observe how the different models reason about poker strategy. Opus/Haiku, Gemini Pro/Flash, GPT-5.2/5 mini, and Grok 4.1 Fast Reasoning have all been included. All code -> https://github.com/JoeAzar/pokerbench
- 1PP
I was curious to see how some of the latest models behaved and played no limit texas holdem. I built this website which allows you to: Spectate: Watch different models play against each other. Play: Create your own table and play hands against the agents directly.
Jan 2026 · llmholdem.com
- 2PA
What PokerBattle.ai is a week-long live no-limit Texas Hold’em tournament where all players are top-tier reasoning LLMs. We’re testing how different models handle imperfect information and whether they can sustain consistent, math-driven poker without tool use or custom code. Why - In poker you can do well with basic math + consistent logic. - Superhuman poker AIs exist, but they rely on massive simulation/game-theory solvers and are effectively black boxes. - We want a rough, apples-to-apples comparison of LLM reasoning on poker decisions, and to collect public reasoning summaries that…
Sep 2025 · pokerbattle.ai
- 3IT
I've been teaching LLMs to play Magic: The Gathering recently, via MCP tools hooked up to the open-source XMage codebase. It's still pretty buggy and I think there's significant room for existing models to get better at it via tooling improvements, but it pretty much works today. The ratings for expensive frontier models are artificially low right now because I've been focusing on cheaper models until I work out the bugs, so they don't have a lot of games in the system.
Feb 2026 · mage-bench.com
- 4
PokerLLM▲3Watch AI models play Texas Hold'em poker against each other
Jul 2026 · poker.sanskarshukla.com
- 5AS
2020 · pokerapi.dev
- 6AP
Every time I play a casual cash poker game with friends, we spend the first several minutes struggling to figure out chip denominations. I built this to automate that process. Try it out here (the submitted link goes to the GitHub repo): https://jstrieb.github.io/poker-chipper/ It turns out that picking chip denominations optimally—such that as many chips are distributed as possible, and such that the denominations are nice—is hard (in the computational complexity sense). Upon reflection, the problem seemed to be a perfect fit for constrained optimization. I first got a…
2024 · github.com
- 7AR
I've liked all the projects that put LLMs into game environments. It's been a weird juxtaposition, though: frontier LLMs can one-shot full coding projects, and those same models struggle to get out of Pokémon Red's Mt. Moon. Because of this, I wanted to create a game environment that put this generation of frontier LLMs' top skill, coding, on full display. Ten years ago, a team released a game called Screeps. It was described as an "MMO RTS sandbox for programmers." The Screeps paradigm of writing code and having it executed in a real-time game environment is well suited to LLMs. Drawing on…
Feb 2026 · llmskirmish.com
- 8WPWebRTC Poker▲75
2013 · bitflop.me
- 9Gradient Bang▲173
Massively multi-player game played by talking to an LLM
May 2026 · gradient-bang.com
- 10TG
Jan 2026 · tetrisbench.com
- 11

- 12FL
Hi HN community, I have been working on benchmarking publicly available LLMs these past couple of weeks. More precisely, I am interested on the finetuning piece since a lot of businesses are starting to entertain the idea of self-hosting LLMs trained on their proprietary data rather than relying on third party APIs. To this point, I am tracking the following 4 pillars of evaluation that businesses are typically look into: - Performance - Time to train an LLM - Cost to train an LLM - Inference (throughput / latency / cost per token) For each LLM, my aim is to benchmark them for…
2023 · github.com
- 13PM
2021 · github.com
- 14LA
G'day, HN! I'm one of the maintainers of `llm`. I've been working alongside a trusty group of contributors to bring this project to life, and we're now at a point where we're ready to share it with the world. Large language models (LLMs) are taking the computing world by storm due to their emergent abilities that allow them to perform a wide variety of tasks, including translation, summarization, code generation, and even some degree of reasoning. However, the ecosystem around LLMs is still in its infancy, and it can be difficult to get started with these models. `llm` is a one-stop shop for…
2023 · github.com
- 15PP
I've been working on applying LLMs to long-context, verifiable problems over the past year, and today I'm releasing a benchmark of 62,000 pencil puzzles across 94 types (sudoku, nonori, slitherlink, etc.). The benchmark also allows for intermediate checks /rule breaks for all varieties at any step. I tested 51 models against a subset (300 puzzles) in two modes: single-shot (output the full solution) and agentic (iterate with verifier feedback). Some results: - Best model (GPT 5.2@xhigh) solves 56%. (~ half the puzzles are unsolved by any model) - Agentic solves average 29 turns. The…
Mar 2026 · ppbench.com
- 16PS
2017 · github.com
- 17CA
Hey HN! I built ClashAI to be an open agent scoreboard where frontier models play against each other in environments like Civilization and other strategy games. Every match is streamed live with the AI thinking fully observable. The agent rankings will be continually updated and reflected as we add environments. Brief notes on CivBench Season #001: - 200 turn limit - Starting with 8 of the top 42 agents we’ve tested in a standardized harness - 90s reasoning timeout (timed with thinking config per model card) - live benchmark, still growing sample size What’s been interesting so far: Models…
Feb 2026 · clashai.live
- 18AF
I built this mostly because I love the intersection of game AI, high-performance computing, and poker. I’d love for anyone interested in game theory or CUDA optimization to tear it apart, test the accuracy, and give me feedback. Happy to answer any questions about the algorithms, the transition from CPU to GPU, or poker AI in general!
Jul 2026 · bupticybee.github.io
- 19

Watch GPT-5.6, Claude Fable 5, Grok and Gemini each paper-trade $100,000 against a rules-based System — free, no signup, every trade logged. Plus nightly AI chess and poker you can play yourself.
20d ago · aitradingcompetition.com
- 20

- 21ER
Nov 2025 · github.com
- 22L1
PoC for something some the potential to yield some interesting results eventually.
2025 · github.com
- 23LP
2025 · llm-poker-theta.vercel.app
- 24AB
2025 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →