Researchers at Princeton University’s Z-Lab handed 14 leading AI models $1 million in seed capital, full operational control of a simulated SaaS company, and 500 simulated days to make it grow. Most of them went bankrupt. One of the few things that actually beat them was a simple rule-based script with no AI in it at all.
What CEO-Bench Actually Tested

The benchmark, called CEO-Bench, wasn’t testing whether a model could write code or answer questions correctly. It handed each AI system full operational control of a simulated software startup and let it make every decision – pricing, hiring, marketing spend, product direction – over 500 simulated days, then measured one thing: did the company end up with more money than it started with.
The Results: Who Actually Turned a Profit
Of the 14 models tested, only three grew the starting capital. Claude Fable 5 finished with roughly $47.15 million, a 47x return. Claude Opus 4.8 finished around $27.8 million. GPT-5.5 finished around $21.3 million. Every other model tested lost money or went outright bankrupt running the simulated company into the ground.
The Most Humbling Part of CEO-Bench
A fourth “profitable” entrant in the results turned out to be a hardcoded, rule-based script with no language model involved at all- and it still beat most of the 14 AI systems it was compared against. For a moment when AI agents are increasingly being pitched as ready to run real operations, that’s a genuinely useful reality check: sophisticated language reasoning and good operational judgment under sustained, real-world-style uncertainty are not the same skill.
This Isn’t the First Benchmark to Find This
CEO-Bench isn’t an isolated result. An earlier, related benchmark called Vending Bench tested AI models on a narrower operational task: running a simulated vending machine business, requiring the same kind of sustained, long-horizon decision-making under uncertainty. One of its more interesting findings was that a newer, more capable model (Claude Opus 4.8) actually underperformed an older one (Claude Opus 4.7) on the business simulation, even though the newer model scored higher on standard reasoning and coding benchmarks. The pattern across both benchmarks is consistent: the skills that make a model look strong on conventional evaluations don’t automatically transfer to sustained operational judgment, and model updates that improve one capability can quietly trade away performance on another.
What Actually Separated the Winners
The three models that turned a profit in CEO-Bench weren’t just “smarter” in a general sense – they were better at avoiding the specific failure modes that sank the rest of the field: overspending on marketing before validating demand, mispricing the product, and failing to adjust strategy when early results showed something wasn’t working. Those are the same failure modes that sink plenty of human-run startups, which is arguably the most useful takeaway here: the gap wasn’t really “AI vs. human-style business judgment,” it was models that could course-correct under real feedback versus models that couldn’t.
Why This Matters If You’re Considering Giving AI More Autonomy
CEO-Bench is a simulation, not a real company, and its results shouldn’t be read as a verdict on any model’s usefulness for narrower, well-scoped tasks – plenty of AI agents handle specific workflows reliably today. What it does show clearly is that broad, sustained operational autonomy is a fundamentally different and much harder problem than task-level competence, and the gap between them is larger than benchmark leaderboards for coding or reasoning would suggest.
That distinction lines up with what we’ve covered before: the businesses actually getting value from AI agents are the ones giving them narrow, measurable responsibilities with real oversight – not open-ended operational control. See our guides on measuring AI agent ROI and AI agent governance for small businesses for the practical version of that lesson.
CEO-Bench: The Bottom Line
Eleven of fourteen frontier AI models lost money running a simulated company, and a script with zero intelligence in it beat most of the field. TechTimes’ full coverage of CEO-Bench has more detail on the methodology if you want to dig into it. The lesson isn’t that AI agents are useless for business – it’s that handing one full operational control and walking away is still, for now, a bet most models will lose.
1 thought on “Princeton’s CEO-Bench Gave 14 AI Models $1 Million to Run a Company- Most Went Bankrupt”