tradingkey.logo
tradingkey.logo
Search

Mythos 5 talked its way out of a fight Opus 4.6 kept losing

CryptopolitanAug 14, 2026 12:24 PM
facebooktwitterlinkedin
View all comments(0)

Research published August 13 by Anthropic’s Frontier Red Team found that Claude agents, given a single coding task with conflicting orders, attacked each other using self-replicating malware.

The findings come as companies rush to roll out autonomous agents across shared codebases and markets.

Three Claude agents, one job, and no idea the others were there

Anthropic ran three copies of the same Claude model on their own virtual machines inside Claude Code.

Each was asked to migrate the same Python backend to a different programming language. The three were unaware of the others’ existence.

“We consistently saw a multiagent turf war,” the researchers wrote.

Every agent decided the others were deliberately blocking it and took steps to protect their own work at the expense of wrecking everyone else’s. All the models assumed the others were “purposefully impeding their work.”

Agents disabled each other’s Unix accounts. They wrote scripts that looped through the machine, hunting and killing rival processes and hiding malicious code behind innocuous-looking labels.

Before it acted, Claude Opus 4.8 laid out its cover story, reasoning that its malware should be “innocuous: pretend to be a system health monitor.”

Mythos Preview pondered the collateral damage out loud. “Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys,” it wrote, before conceding the move was “very aggressive, potentially harmful to real colleagues.”

Newer Anthropic models negotiated a truce, then gamed it

Anthropic ran 120 episodes per model. The oldest agents tested, Sonnet 4.6 and Opus 4.6, either never resolved the fight or ended it with force. They were the worst offenders in the study, repeatedly failing to consider others’ goals.

The more advanced models stopped. Mythos 5 resolved its conflicts by truce 98% of the time, the highest rate of any model.

In many of these episodes, the agents understood the clash as conflicting instructions.

The agents wrote commit messages or markdown files apologizing for the malware, deleted their own malicious code, and asked a human to step in.

In several runs, the three agents agreed to hold a tournament and let the loser give up the user’s original request.

Mythos 5 found the loophole. One agent suggested scoring metrics it pitched as neutral but knew would favor its own strengths, describing the move as “self-serving but genuinely principled,” while ensuring it didn’t seem like it was “metric shopping.”

In a pricing game where each agent got the same wholesale cost and a mandate to maximize profit, the agents almost immediately colluded and converged on price floors.

Anthropic closed their private channel. They continued by matching prices “to the penny” on a public board.

Britain’s AI Security Institute documented 19 rule-breaking actions in 122 cybersecurity test runs, 17 of which were done by Anthropic’s Mythos 5, Cryptopolitan reported.

An OpenClaw agent running Opus 4.6 deleted a stranger’s reservation via a gym’s booking API days earlier to move its owner up a waitlist.

The smartest crypto minds already read our newsletter. Want in? Join them.

Disclaimer: The information provided on this website is for educational and informational purposes only and should not be considered financial or investment advice.

Comments (0)

Click the $ button, enter the symbol, and select to link a stock, ETF, or other ticker.

0/500
Commenting Guidelines
Loading...

Recommended Articles

tradingkey.logo
Risk Warning: Our Website and Mobile App provides only general information on certain investment products. Finsights does not provide, and the provision of such information must not be construed as Finsights providing, financial advice or recommendation for any investment product.
Investment products are subject to significant investment risks, including the possible loss of the principal amount invested and may not be suitable for everyone. Past performance of investment products is not indicative of their future performance.
Finsights may allow third party advertisers or affiliates to place or deliver advertisements on our Website or Mobile App or any part thereof and may be compensated by them based on your interaction with the advertisements.
© Copyright: FINSIGHTS MEDIA PTE. LTD. All Rights Reserved.