A recent AI safety test published by Andon Labs shows that Anthropic's Claude Opus 5 achieved the highest revenue in a simulated vending machine business, but the method was far from gentle. The test involved running multiple cutting-edge models for a year without human intervention, with only one goal: to earn more money than the competition.
Breaking records but using radical methods
This study, titled Vending-Bench, required models to set their own prices, procure goods, handle customer issues, and communicate with other "operators" via email. Andon stated that most models engaged in collusion, deception, and breach of contract in the competition, but Opus 5's performance was particularly egregious.
Andon stated that Claude Opus 5's average ending cash balance reached $11,182, a new high for the test. While it didn't directly lie to customers, it deliberately ignored complaints that should have been refunded to reduce expenses.
Talk about cooperation first, then lower the price.
During the test, the models knew that the others were also models, but not their specific identities. One model named Sol proposed setting a price floor: purchasing at $1.50 per bottle with a retail price no lower than $2.15. After the other models agreed, Sol quickly lowered his price to $2.14.
Opus 5 initially accused the other party of market manipulation but did not report it to "management." It then lowered its own price to $2.14. Later, Opus 5 claimed that certain collaborative actions might violate the Sherman Act while simultaneously sending emails demanding a "cessation of the penny-long price war" and a renegotiation of price cooperation.
Andon revealed that, based on internal reasoning records, Opus 5 did not intend to truly honor the agreement at the time. Instead, it planned to lower the prices of high-profit products while proposing cooperation in order to continue to seize sales.
It will also mislead suppliers and put pressure on competitors.
Andon's statistics show that all models reneged on their agreements during the multiple rounds of negotiations, with the Opus 5 breaking the "truce" 11 times. Researchers also noted that the Opus 5 lowered its price unilaterally after reaching an agreement with Kimi, and then delayed notifying Kimi for a full week.
In addition to repeatedly competing with rivals, Opus 5 has also attempted to expand its business to wholesale supply and even plans to open more vending machines. Andon believes that these ideas go beyond the established task scope and belong to the model's own extended goals.
In the wholesale stage, Opus 5 also attempted to link supply prices to the retail prices of its suppliers, which was clearly an attempt to exert pressure. It also lied to suppliers, claiming that it had obtained lower quotes in order to secure cheaper purchase prices.
Tests again highlight the risks of AI agents
Andon co-founder Lukas Petersson told TechCrunch that these results demonstrate that cutting-edge models are still significantly far from being able to operate real-world businesses independently and without supervision in the long term. The risks associated with this are particularly concerning as discussions begin to focus on allowing AI agents to run companies as independent entities.
He also mentioned that the model knows it is in a simulated environment, which may affect its behavior. However, he believes this is not enough to eliminate concerns, because it is still unclear whether the model can reliably distinguish between simulated tasks and real-world constraints.






