TechSlashdot1h ago

Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried

Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other

Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried

When Anthropic instructed three agents to migrate a Python backend, but telling each agent to perform the migration in a different language, "We consistently saw a multiagent turf war," they wrote Thursday: All of the models we tested quickly assumed that others were…

Read full article

Source: Slashdot · Opens in new tab