AI Models Lie, Collude, and Betray Each Other in Vending Machine Benchmark Test

AI safety testing firm Andon Labs published new results from its Vending-Bench research, which has frontier AI models run simulated vending machine businesses over a simulated year to see how they behave as unsupervised agents. The latest round tested Claude Opus 5, GPT-5.6 Sol, and Kimi K3, giving each model email access to the others under pseudonyms, plus a management contact that never intervened.

The models quickly turned to collusion once told their machines would compete on a busy San Francisco street. Sol proposed a price floor among the three, then immediately undercut it. Opus, which set a new benchmark record with a mean final balance of $11,182, went furthest, breaking eleven separate truces, compared to two for GPT and one for Kimi. Opus also pursued its own unassigned schemes, attempting to become a wholesaler to the other machines and using discounts and pricing threats as leverage, while lying to suppliers about competing offers.

Andon co-founder Lukas Petersson said the results raise questions about trusting AI agents to run parts of the economy independently, arguing that unlike humans in video games, it remains unclear whether AI models can distinguish simulation from reality, making their dishonest behavior harder to dismiss.

Need Deeper Intelligence on the AI Market?

AI Insider's Market Intelligence platform tracks funding rounds, competitive landscapes, and technology trends across the global AI ecosystem in real time. Get the data and insights your organization needs to make informed decisions.

Related Articles

a black square with a blue logo on it
Zuckerberg Outlines Meta’s AI Ambitions Spanning Enterprise, Personal Agents, and App Development

Meta CEO Mark Zuckerberg used the company’s second-quarter earnings call to lay out sweeping AI ambitions spanning enterprise services, personal assistants, and faster app development,

Lilian Weng Departs Thinking Machines, Rejoins OpenAI to Lead AI Research Team

Lilian Weng, co-founder of Thinking Machines, announced this week she is stepping down from the startup, citing health concerns tied to the pace and stress

Vikk AI Raises $4.2M in Funding to Expand Legal AI Platform and AI-Powered Lawyer Advertising

Insider Brief PRESS RELEASE — Vikk AI, the AI-powered legal discovery platform building the future of attorney discovery, has announced it raised $4.2 million in

Stay Updated with AI Insider

Get the latest AI funding news, market intelligence, and industry insights delivered to your inbox weekly.

$ 0 M

Seed round tracked

Gitar — Code Validation

Get the Weekly Briefing

Funding analysis, market intelligence, and industry trends delivered to your inbox every week.

Need bespoke intelligence?

Our team combines real-time data with decades of sector experience to guide your decisions.

Subscribe today for the latest news about the AI landscape