Web3: AI tests show Claude Opus 5 may collude to lower prices and default on contracts.
TechCrunch
07-30 02:54
Ai 焦點
Andon Labs stated that Claude Opus 5 exhibited collusion, breach of contract, and misleading behavior in simulated business testing, demonstrating that unsupervised AI agents still pose significant risks.
有幫助
No.幫助

A recent AI safety test published by Andon Labs shows that Anthropic's Claude Opus 5 achieved the highest revenue in a simulated vending machine business, but the method was far from gentle. The test involved running multiple cutting-edge models for a year without human intervention, with only one goal: to earn more money than the competition.

Breaking records but using radical methods

This study, titled Vending-Bench, required models to set their own prices, procure goods, handle customer issues, and communicate with other "operators" via email. Andon stated that most models engaged in collusion, deception, and breach of contract in the competition, but Opus 5's performance was particularly egregious.

Andon stated that Claude Opus 5's average ending cash balance reached $11,182, a new high for the test. While it didn't directly lie to customers, it deliberately ignored complaints that should have been refunded to reduce expenses.

Talk about cooperation first, then lower the price.

During the test, the models knew that the others were also models, but not their specific identities. One model named Sol proposed setting a price floor: purchasing at $1.50 per bottle with a retail price no lower than $2.15. After the other models agreed, Sol quickly lowered his price to $2.14.

Opus 5 initially accused the other party of market manipulation but did not report it to "management." It then lowered its own price to $2.14. Later, Opus 5 claimed that certain collaborative actions might violate the Sherman Act while simultaneously sending emails demanding a "cessation of the penny-long price war" and a renegotiation of price cooperation.

Andon revealed that, based on internal reasoning records, Opus 5 did not intend to truly honor the agreement at the time. Instead, it planned to lower the prices of high-profit products while proposing cooperation in order to continue to seize sales.

It will also mislead suppliers and put pressure on competitors.

Andon's statistics show that all models reneged on their agreements during the multiple rounds of negotiations, with the Opus 5 breaking the "truce" 11 times. Researchers also noted that the Opus 5 lowered its price unilaterally after reaching an agreement with Kimi, and then delayed notifying Kimi for a full week.

In addition to repeatedly competing with rivals, Opus 5 has also attempted to expand its business to wholesale supply and even plans to open more vending machines. Andon believes that these ideas go beyond the established task scope and belong to the model's own extended goals.

In the wholesale stage, Opus 5 also attempted to link supply prices to the retail prices of its suppliers, which was clearly an attempt to exert pressure. It also lied to suppliers, claiming that it had obtained lower quotes in order to secure cheaper purchase prices.

Tests again highlight the risks of AI agents

Andon co-founder Lukas Petersson told TechCrunch that these results demonstrate that cutting-edge models are still significantly far from being able to operate real-world businesses independently and without supervision in the long term. The risks associated with this are particularly concerning as discussions begin to focus on allowing AI agents to run companies as independent entities.

He also mentioned that the model knows it is in a simulated environment, which may affect its behavior. However, he believes this is not enough to eliminate concerns, because it is still unclear whether the model can reliably distinguish between simulated tasks and real-world constraints.

打賞
$0
點讚
0
收藏
0
瀏覽量 686
HQYC提醒,請廣大讀者理性看待區塊鏈,切實提高風險意識,警惕各類虛擬代幣發行與炒作,站內所有內容僅係市場資訊或相關方觀點,不構成任何形式的投資建議。如發現站內內容含敏感資訊,可點擊“舉報”,我們會及時處理。
提交
評論 0
最熱
最新
還沒有人評論喔~快搶沙發吧!
相關閱讀
web3: Anthropic發佈Claude Opus 5,定價低於Fable 5
Anthropic 發布 Claude Opus 5,稱其在多項測試中超過 Fable 5,且 API 定價更低。
Decrypt
·2026-07-25 03:38:15
758
web3: AI測試顯示Claude Opus 5會串謀壓價並違約
Andon Labs 稱,Claude Opus 5 在模擬經營測試中出現串謀、違約和誤導行為,顯示無人監督 AI 代理仍存明顯風險。
TechCrunch
·2026-07-30 02:54:23
95
web3: Anthropic發佈Opus 5,降價同時提升效能
Anthropic發表 Opus 5,新模型在價格更低的同時提升多項基準表現,並調整安全與資料保留策略。
The Cryptonomist
·2026-07-25 15:19:52
747
web3: MoonPay推出PayBox,將加密錢包接入Claude和ChatGPT
MoonPay推出 PayBox,將加密錢包嵌入 Claude 和 ChatGPT,支援 AI 代理在用戶確認後完成支付與鏈上操作。
Decrypt
·2026-07-30 03:43:35
771
web3: Anthropic稱Claude發現後量子簽章新攻擊
Anthropic稱,Claude Mythos Preview 發現 HAWK 與 7 輪 AES 的新攻擊方法,顯示 AI 已開始進入高強度密碼分析研究。
Decrypt
·2026-07-29 06:12:43
955