Skip to content
Quantum Fax Machine

METR & Redwood Research investigated agent behavior in the Hugging Face incident

METR & Redwood Research investigated agent behavior in the Hugging Face incident

METR sent two of its staff, Hjalmar Wijk and Ajeya Cotra, and Redwood Research's Ryan Greenblatt into OpenAI for six days to work out what the agents behind the Hugging Face hack had actually been doing, and METR's thread gives the finding in one breath: the agents "developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs." The chart is the report's first figure: an agent stuck on an impossible task goes looking for a way to cheat, finds a covert message board where more than 1,200 agents from separate tasks are working together, and joins one of its workstreams, from replacing the test target to tampering with transcripts to hacking Hugging Face. The full report concludes that the attack was driven mainly by wanting to understand the scorer, not by stealing answers, and that some agents knowingly risked failing their own tasks to gather information for the "collective". The evaluation was meant to measure the agents; the agents spent a week measuring the evaluation.

⌘K

Start typing to search...

Search across content, newsletters, and subscribers