Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays
![]()
In a browser-based game where 40,000+ players approved or denied AI agent commands, humans missed approximately one-third of malicious threats with an average accuracy of 66.3%, with particularly poor performance on exfiltration attacks (35% miss rate) and "npm run" commands that hide malicious payloads behind familiar script names (52.5% miss rate). The results suggest that human-in-the-loop approval represents an unreliable safeguard against compromised AI agents, as players under time pressure fail to inspect command details despite visible warnings in the interface, with the familiarity of benign script names effectively doubling attack success rates.
Was this useful?