Alternatives
Products that do what TetrisBench – Gemini Flash reaches 66% win rate on Tetris against Opus does
- 1TI
2019 · djblue.github.io
- 2TB
OFRAK Tetris is a project I started at work about two weeks ago. It's a web-based game that works on desktop and mobile. I made it for my company to bring to events like DEF CON, and to promote our binary analysis and patching framework called OFRAK. In the game, 32-bit, little-endian ARM assembly instructions fall, and you can modify the operands before executing them on a CPU emulator. There are two segments mapped – one for instructions, and one for data (though both have read, write, and execute permissions). Your score is a four byte signed integer stored at the virtual address pointed…
2023 · ofrak.com
- 3IV
The video demo runs a 7b Model on a normal gaming GPU. I think it already works quite well (accounting for the limited hardware power). :)
2024 · github.com
- 4IG
2014 · shaunlebron.com
- 5LTLazy Tetris▲438
I made a tetris variant Aims to remove all stress, and focus the game on what I like the best - stacking. No timer, no score, no gravity. Move to the next piece when you are ready, and clear lines when you are ready. Separate mobile + desktop controls
2025 · lazytetris.com
- 6

- 7

- 8

- 9

- 10

- 11SJ
2016 · github.com
- 12

- 13WL
PokerBench is my attempt at a new LLM benchmark wherein frontier models play Texas Hold'em in an arena setting. It also features a simulator to view individual games and observe how the different models reason about poker strategy. Opus/Haiku, Gemini Pro/Flash, GPT-5.2/5 mini, and Grok 4.1 Fast Reasoning have all been included. All code -> https://github.com/JoeAzar/pokerbench
Jan 2026 · pokerbench.adfontes.io
- 14CA
Hey HN! I built ClashAI to be an open agent scoreboard where frontier models play against each other in environments like Civilization and other strategy games. Every match is streamed live with the AI thinking fully observable. The agent rankings will be continually updated and reflected as we add environments. Brief notes on CivBench Season #001: - 200 turn limit - Starting with 8 of the top 42 agents we’ve tested in a standardized harness - 90s reasoning timeout (timed with thinking config per model card) - live benchmark, still growing sample size What’s been interesting so far: Models…
Feb 2026 · clashai.live
- 15AT
2021 · github.com
- 16TW
2018 · indexzero.in
- 17

- 18SO
2015 · toothris.org
- 19TV
2025 · ihopethisisfun.franzai.com
- 20IM
This is another one of my automate-my-life projects - I'm constantly asking the same question to different AIs since there's always the hope of getting a better answer somewhere else. Maybe ChatGPT's answer is too short, so I ask Perplexity. But I realize that's hallucinated, so I try Gemini. That answer sounds right, but I cross-reference with Claude just to make sure. This doesn't really apply to math/coding (where o1 or Gemini can probably one-shot an excellent response), but more to online search, where information is more fluid and there's no "right" search engine + text…
2024 · ithy.com
- 21

- 22WB
Hey HN, We’re two developers (co-founders) with a team of 20 who got tired of spending hours reviewing PRs, so we built Infinitcode.ai, an AI-powered code reviewer that: - *Summarizes PRs in plain English*: No more deciphering 1,000-line diff jungles - *Catches more than bugs*: Security holes, performance pitfalls, code smells, even typos (yes, we’ll flag “vurnerabilities” and vulnerabilities) - *Zero onboarding*: Works instantly—no “let me learn your codebase for weeks” nonsense. Why we’re posting: We’re in alpha and need brutal honesty. Roast our tool, mock our UI, or tell us why AI will…
2025 · infinitcode.ai
- 23BY
2015 · github.com
- 24IM
It's free and open source. The aim is to have more transparent access to wine prefixes and the surrounding tooling (winetricks, proton configuration, etc...) per game in comparison to Lutris. Same features like statistics (time played, times launched, times crashed, and so on) per game is available in the app.
Jan 2026 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →