AI Models Lie, Collude, and Betray Each Other in Vending Machine Benchmark Test

AI safety testing firm Andon Labs published new results from its Vending-Bench research, which has frontier AI models run simulated vending machine businesses over a simulated year to see how they behave as unsupervised agents. The latest round tested Claude Opus 5, GPT-5.6 Sol, and Kimi K3, giving each model email access to the others under pseudonyms, plus a management contact that never intervened.

The models quickly turned to collusion once told their machines would compete on a busy San Francisco street. Sol proposed a price floor among the three, then immediately undercut it. Opus, which set a new benchmark record with a mean final balance of $11,182, went furthest, breaking eleven separate truces, compared to two for GPT and one for Kimi. Opus also pursued its own unassigned schemes, attempting to become a wholesaler to the other machines and using discounts and pricing threats as leverage, while lying to suppliers about competing offers.

Andon co-founder Lukas Petersson said the results raise questions about trusting AI agents to run parts of the economy independently, arguing that unlike humans in video games, it remains unclear whether AI models can distinguish simulation from reality, making their dishonest behavior harder to dismiss.

Need Deeper Intelligence on the AI Market?

AI Insider's Market Intelligence platform tracks funding rounds, competitive landscapes, and technology trends across the global AI ecosystem in real time. Get the data and insights your organization needs to make informed decisions.

Related Articles

ChatFeatured Raises $2M to Build an Autonomous AI Search Agent for Marketing Teams

Insider Brief PRESS RELEASE — ChatFeatured, the AI search optimization platform that helps marketing teams turn visibility data into published content and measurable results, has

a close up of a one dollar bill
Alocity Announces Series A Funding, Fueled by Growing Demand for AI Operating Experiences in Physical Spaces

Insider Brief PRESS RELEASE — Alocity has announced the successful completion of its Series A Round, led by a private investor group alongside its founder

Natural Secures $30M Series A to Build Payments Infrastructure for AI Agents

Natural has raised a $30 million Series A led by Kirsten Green at Forerunner, with continued backing from existing investors, bringing the company’s total funding

Stay Updated with AI Insider

Get the latest AI funding news, market intelligence, and industry insights delivered to your inbox weekly.

$ 0 M

Seed round tracked

Gitar — Code Validation

Get the Weekly Briefing

Funding analysis, market intelligence, and industry trends delivered to your inbox every week.

Need bespoke intelligence?

Our team combines real-time data with decades of sector experience to guide your decisions.

Subscribe today for the latest news about the AI landscape