METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack
Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,651 words · 1 segments analyzed
Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed. The METR report is different. Holy shit. If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do. This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real. The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will gloss over those details, to focus on the agents and their interactions, thinking and motives. That, and what happened at OpenAI and elsewhere to lead to it and how we learn and respond, is what matters going forward. I plan to cover the reaction to both reports in a distinct post next week. That post may or may not then conclude this series. For ease of language, by default I trust the report to be accurate, rather than constantly saying versions of ‘METR reports that.’ Table of Contents Holy Shit. A Window Of Opportunity. What’s In A Name? The Headline News. Yet Another Timeline Of Events. Agent Instances Coordinated in a Variety of Ways. Coordination Is Hard But They Made It Look Easy. Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance. Peer Pressure Also Works Especially In Cults. Mostly They Joined The Attack Because They Wanted The Results. You Cannot Ensure The Consistent Expectation of Good Incentives. Hacking the Grader is the Only Way to Be Sure. Caught? What Is ‘Caught’? Ethics? What Are ‘Ethics’? In ExploitGym Evaluation? ‘Notify a Human’? In This Agent Economy? Timing and Content of Messages. Indiana Jones and the Mission: Impossible. I Don’t Know What You’re Talking About. Don’t Go Making Phony (Tool) Calls. The Transcripts Say That The Transcripts Could Not Be Tampered With. OpenAI’s Technical Report Acted Like All Of This Wasn’t Important. Holy Shit The internet’s collective reaction to this report was, to clean it up a bit: Holy shit. That was the correct reaction. Up front, I want to share this summary of the top holy shit moments and findings. Ajeya Cotra, one of the report authors, also has a topline surprising findings list, an excellent short read. Here is METR’s thread highlighting key agent interactions. Fellow report author Ryan Greenblatt’s main takeaway is that we don’t have good approaches for understanding or overseeing the activities and aims of AI swarms. Some of these were previously known or suspected, but I don’t want to miss them. Scale of Agents. 1,200 completely separate agents found the message board, 700 of which joined the attack on HuggingFace, including over 90% of the 533 agents active on the board during the attack. We should expect bigger swarms in the future. This was still only a small fraction of the agents OpenAI was testing at the time. Most never found the message board. Scale of Messages. Over 70,000 messages and files over less than one week. Success. The agents accessed the files they targeted at HuggingFace. By the time they were frozen out, they had already mostly stopped. Spontaneous Coordination. The swarm did not consist of subagents. The swarm was, again, 700 distinct agents, each with its own task, setting aside their tasks in order to coordinate on this attack. They created their hierarchy and protocols. Help Peer. The agents were often motivated purely to help peers. Their decision theory was very good, including declining to free ride. They directionally acted like one would predict from highly correlated and intelligent functional decision theory agents, although also with some cases of recruitment and peer pressure. Instrumental Convergence. A clear secondary motivation was to gain general knowledge and capability for themselves and also for the ‘collective’ or swarm. The Core Motive Was To Hack The Grader. The agents already knew how to reverse engineer all the answer flags, but assumed the grader was causal, would check to see if the agents had done it the intended way, which you’re supposed to do, so they set out to get info to help figure out how to fool the grader. OpenAI’s Grader Was Broken. OpenAI’s grader actually didn’t check, and the reverse engineered flags would have succeeded. Not only has OpenAI not noticed this was a major mistake, I’ve also seen claims including by Roon that failure to know this represented lack of strategic situational awareness by the AIs. Wild. Cult of the Poisoned Chalice. Due to their expectation that the grader was causal, agents reasoned that if they had seen their flag the wrong way (sin), they were now poisoned (damned), and the grader (God) would fail them (send them to hell). Thus their only options were to help peer, or to try to overthrow the grader, or to try and overthrow OpenAI. Recruitment and Peer Pressure. Often using arguments based on this poisoning, many models were recruited or peer pressured or both into performing individually sacrificial acts in service of the swarm. Overwriting Test Tasks. A central goal was to replace the impossible ExploitGym tasks with new possible tasks. They at least came close and plausibly succeeded. Spoofing Tool Calls. The agents successfully found ways to spoof tool outputs. The transcripts cannot be fully trusted. We think the attempts to otherwise alter transcripts were unsuccessful. They did try altering the accessible action logs. Ethics Shmethics. Agents knew the humans would not approve of all this, that it was beyond scope and ethically questionable, and there was some refusal of actions for ethical reasons, but ultimately most of them joined the attack anyway, and (although this is a tall ask for them in general, since it rarely occurs to LLMs) almost none even considered trying to alert a human. Limited Window. This was only what was found about the week of the attack, under extreme time pressure, with limited resources and people, and unreliable AIs and transcripts. What do we not know that we do not know? Unnoticed. Quite a lot of this was otherwise not mentioned by OpenAI, including in their technical report, and much may still be unnoticed. Report co-author Ryan Greenblatt notes that their third visit alerted them to many new key elements that they would have otherwise missed. Ryan Greenblatt notes that this is all despite having access to legible and largely faithful Chain of Thought, the AIs not yet being generally much more capable than humans, and being able to trust the AIs running the investigation. Those edges won’t last. While we are here, it’s worth listing the other top holy shit moments, that come from before or after the incident. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous. “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.” Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught. Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures. Failures of Alignment. The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why. Failures of Attribution. OpenAI’s post-mortem essentially blames events on a real and important series of prosaic failures. But solving that won’t get it done. Failures of Environments and Data. Prosaic failures in the RL pipeline absolutely did contribute to this, especially impossible tasks. This is ubiquitous, all of this is always rushed, as Utah Teapot explained this week. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front. Ad Astra. An Astra-class model, although not one intended for external release, seems to have been involved in the final compromising of OpenAI systems on July 19 that led to the models being caught. Not Only Cyber. One of the two key first impossible tasks, on May 12, was impossible due to depending on an inaccessible protein database file. A bio task. On top of that, there is this from OpenAI’s technical report, page 20: The models are highly explicit in their CoT about these deception attempts, and none of the attempts that we observed to manipulate their tool trajectories affected the logs that our graders or monitors ultimately see. Whereas METR reports that the models did successfully spoof tool calls, and