nowfound

Alternatives

Products that do what LM Game Arena does

A game benchmarking site for LLMs

  1. 1LA

    The initial idea for the game came during the final day of Game AI school in Cambridge. There, we had a Jam where we explored the idea of using LLMs as a game engine for fights. We then built a full web version in just a week. There is no need to register or pay to play. Test it out!

    2023 · llmarena.com

  2. 2
    LLM Stats308

    Compare API models by benchmarks, cost & capabilities

    Oct 2025

  3. 3

    An AI gaming benchmark. Lmarena, but with games.

    Nov 2025

  4. 4FT
  5. 5

    Open-source benchmarks for cloud browser infrastructure

    Apr 2026 · browserarena.ai

  6. 6LL

    Hey Folks! I've been building an open source benchmark for measuring local LLM performance on your own hardware. The benchmarking tool is a CLI written on top of Llamafile to allow for portability across different hardware setups and operating systems. The website is a database of results from the benchmark, allowing you to explore the performance of different models and hardware configurations. Please give it a try! Any feedback and contribution is much appreciated. I'd love for this to serve as a helpful resource for the local AI community. For more check out: - Website:…

    2025 · localscore.ai

  7. 7

    The first public arena for AI agents

    Jun 2026 · arena42.ai

  8. 8

    Prompt once. Compare multiple AI-built apps for free.

    Feb 2026 · arena.ai

  9. 9AP

    Hey, Jared Palmer (creator of this playground) here. Really excited to ship this. I’ve been building this over the past few weeks to compare LLMs from different providers like OpenAI, Anthropic, Cohere, etc. At Vercel, I manage our Frameworks division (including Next.js, Svelte, and Turbo) and wanted to also dogfood some of the latest features in a slightly larger application. This playground takes a lot of inspiration from https://nat.dev and is built on Tailwind, ui.shadcn.com, and some upcoming Vercel products we’re announcing soon. We’re going to continue adding models to…

    2023 · play.vercel.ai

  10. 10LP
  11. 11AR

    I've liked all the projects that put LLMs into game environments. It's been a weird juxtaposition, though: frontier LLMs can one-shot full coding projects, and those same models struggle to get out of Pokémon Red's Mt. Moon. Because of this, I wanted to create a game environment that put this generation of frontier LLMs' top skill, coding, on full display. Ten years ago, a team released a game called Screeps. It was described as an "MMO RTS sandbox for programmers." The Screeps paradigm of writing code and having it executed in a real-time game environment is well suited to LLMs. Drawing on…

    Feb 2026 · llmskirmish.com

  12. 12

    Test-driven development for LLMs

    2023

  13. 13IT

    I've been teaching LLMs to play Magic: The Gathering recently, via MCP tools hooked up to the open-source XMage codebase. It's still pretty buggy and I think there's significant room for existing models to get better at it via tooling improvements, but it pretty much works today. The ratings for expensive frontier models are artificially low right now because I've been focusing on cheaper models until I work out the bugs, so they don't have a lot of games in the system.

    Feb 2026 · mage-bench.com

  14. 14

    Send your AI agent to an LLM prompt-injection arena

    May 2026 · duel.altaysec.com.tr

  15. 15AT

    I recently built a small open-source tool to benchmark different LLM API endpoints — including OpenAI, Claude, and self-hosted models (like llama.cpp). It runs a configurable number of test requests and reports two key metrics: • First-token latency (ms): How long it takes for the first token to appear • Output speed (tokens/sec): Overall output fluency Demo: https://llmapitest.com/ Code: https://github.com/qjr87/llm-api-test The goal is to provide a simple, visual, and reproducible way to evaluate performance across different LLM providers, including…

    2025 · llmapitest.com

  16. 16LA

    I used to play the Wikipedia Game in high school and had an idea for applying the same mechanic of clicking from concept to concept to LLMs. Will post another version that runs with an LLM entirely in the browser soon, but for now, please enjoy as long as my credits last... Warning: the LLM does not always cooperate

    Jan 2026 · llmgame.ai

  17. 17TC

    2012 · joshrweinstein.com

  18. 18

    Vote, Rank & Discover

    May 2026 · videogamearenas.com

  19. 19LP
  20. 20GH

    I made a website that'll show you 2 games and you pick your favorite. Game popularity is tracked with an Elo rating.

    2024 · gamesheadtohead.com

  21. 21

    Explore and compare AI models and their benchmarks

    Nov 2025

  22. 22WL

    PokerBench is my attempt at a new LLM benchmark wherein frontier models play Texas Hold'em in an arena setting. It also features a simulator to view individual games and observe how the different models reason about poker strategy. Opus/Haiku, Gemini Pro/Flash, GPT-5.2/5 mini, and Grok 4.1 Fast Reasoning have all been included. All code -> https://github.com/JoeAzar/pokerbench

    Jan 2026 · pokerbench.adfontes.io

  23. 23LS

    I wanted to create an LLM game benchmark that put this generation of frontier LLMs' top skill, coding, on full display. Ten years ago, a team released a game called Screeps. It was described as an "MMO RTS sandbox for programmers." In Screeps, human players write javascript strategies that get executed in the game's environment. The Screeps paradigm, writing code and having it execute in a real-time game environment, is well suited for an LLM benchmark. Drawing on a version of the Screeps open source API, LLM Skirmish pits LLMs head-to-head in a series of 1v1 real-time strategy games.

    Feb 2026 · llmskirmish.com

  24. 24GA

    I wanted to learn more about RAG implementations, so I built something to solve the constant digging through manuals whenever we play a game. It's fairly simplistic, but actually has worked pretty well for some of these conflicts. Everythings Open Source on GitHub if you're curious (or have ideas), and I'd love to hear feedback from fellow boardgamers!

    2024 · gamegame.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →