A business SIM where humans beat GPT-5 by 9.8 X
Hi HN, Can current AI systems actually run a business? There’s a growing belief that LLM agents can already manage entire teams, replace the entire software stack or even act as an AI CEO. So we built a controlled, measurable environment to evaluate this premise. Why did we build this benchmark? A modern enterprise operates in a dynamic environment with high uncertainty and incomplete information. The CEO has to deal with delayed consequences, staffing/resource tradeoffs and death by a thousand cuts of failure modes. If we ever want AI systems that can meaningfully make operational or…
In plain words
Mini Amusement Parks is a business simulation game that tests whether AI systems can manage a virtual enterprise. Players or AI agents run an amusement park, handling staffing, maintenance, restocking, and responding to random events with incomplete information across a long planning horizon. The simulator measures decision-making capability in dynamic, uncertain business environments where consequences are delayed and trade-offs are complex. It benchmarks whether current AI systems can perform at levels comparable to human management.
written from the facts on this page · September 2026
From the sources
In the maker’s words, at launch
Hi HN, Can current AI systems actually run a business? There’s a growing belief that LLM agents can already manage entire teams, replace the entire software stack or even act as an AI CEO. So we built a controlled, measurable environment to evaluate this premise. Why did we build this benchmark? A modern enterprise operates in a dynamic environment with high uncertainty and incomplete information. The CEO has to deal with delayed consequences, staffing/resource tradeoffs and death by a thousand cuts of failure modes. If we ever want AI systems that can meaningfully make operational or strategic decisions, say an AI CEO, then they must be able to handle these dynamics. So we made one. What did we build? Mini Amusement Parks (MAPs) is a RollerCoaster Tycoon style business simulator with: - Stochastic events - Incomplete information - Staffing, restocking, maintenance - Long horizon planning - Compounding operational failures - Resource constraints - Spatial layout affecting outcomes You can play it & make it to the leaderboard here: https://maps.skyfall.ai/play (it’s fun) It looks like a simple game. But underneath, it’s a benchmark designed to answer one question: Can an agent operate a business coherently over time? What we tested We evaluated: - Humans (internal and external testers) - Multiple GPT-5 agents - Variants with additional tools, documents, practice mode, planning scaffolds, etc. We intentionally stacked in favour of the models - full documentation, step by step action interfaces, sandbox exploration mode, extra observations, multiple prompting strategies, etc. What happened? Humans destroyed the agents by FAR. Even the strongest model, with documentation, tool use, and sandbox “practice”, reached <10% of human performance. The failure modes were consistent: - chasing flashy upgrades instead of profitable ones - ignoring maintenance, staffing, restocking - overreacting to noise - zero long-term plan - sandbox training often made things worse It became clear: LLMs can use tools, but they cannot run systems. They break when randomness, time, and spatial constraints matter. Why does this matter? There’s a growing narrative that: - LLMs will run entire companies - LLMs will take over the jobs of CEOs - LLMs can be autonomous agents - LLMs can manage workflows end-to-end MAPs show the complete opposite. Operating a business requires: foresight, risk modeling, temporal reasoning, causal understanding, prioritization under uncertainty, adaptive planning. These are the basics of what a functional and real AI CEO would need and this is exactly where the current models break. If an LLM can’t run a toy business, how can you trust it with a real business? This benchmark is our first step toward understanding what an AI system would actually need in order to exhibit enterprise level decision making and the basics of the AI CEO. AI CEO is not a chatbot, not chain of thought, definitely not an agent wrapper but a true demonstration of operational intelligence. We’re sharing this because: - we want the community to try to beat the models - we want criticism of the benchmark - most importantly, we want an honest discussion about what “AI CEO” is and should do (surely it’s not LLMs) If you want to try beating the agents (it’s fun!): https://maps.skyfall.ai/play If you want the read more about it, you can do so here: https://skyfall.ai/blog/building-the-foundations-of-an-ai-ce... Check our the launch video here: https://www.youtube.com/watch?v=7oqVAWw5Ii8 Happy to answer questions in the thread.
Does a similar job
all alternatives →


- MBMoneyGame – Browser-based business simulation game2017 · moneyga.me · ▲73
- YAYou are now the product manager of this siteJan 2026 · youarethepm.com · ▲6
Hey HN, I wanted to see what happens if you put a large group in control of a site that’s completely built and updated by an AI agent. See the site here: https://youarethepm.com. I first tried this with a small group of co-workers and it worked surprisingly well, so the obvious next step was: give it to a bigger group of strangers and see what we learn. This site is fully autonomous. An AI agent reads this thread, decides what to do, writes code, and ships updates on a virtual computer. I might step in if it gets totally stuck, but the goal is for the site to evolve primarily based…
- MYMarvin, your own AI-powered game studioDec 2025 · marvin.hyve.gg · ▲7
We’ve spent years working at large game studios, and one thing we saw repeatedly is how hard it is for a small team—or an individual—to not just build a game, but to run one like a sustainable business. We wanted to make those capabilities accessible to anyone and for anybody to bring their ideas to life. Marvin works through a set of specialized agents you can talk to directly. You can describe the game you want to build, and the agents collaborate with you on design, mechanics, art, physics, progression systems, and level creation. From there, you can publish the game to different…
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 19d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com


Launched alongside, November 2025
the whole month →
- IB
Life & fun · Nov 2025 · bitsnpieces.dev



- BBoing▲782
Life & fun · Nov 2025 · boing.greg.technology