Skip to content
HN On Hacker News ↗

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

▲ 122 points 106 comments by stikit 3d ago HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,541
PEAK AI % 0% · §1
Analyzed
Sep 2
backend: pangram/v3.3
Segments scanned
1 windows
avg 1541 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,541 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

Dates in scope: June 26th – July 13th Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions. Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days1 to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.” Our investigation focused mostly2 on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI’s recent Black Hat presentation were out of scope, as was OpenAI’s investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment.3 Figure 1: Anatomy of an agent encountering the unsanctioned “message board” and joining the attack on Hugging Face. The three CoT quotes are from different agents, but illustrate a typical trajectory. This post has three parts. First, we describe our core takeaways about the Hugging Face attack after conducting this investigation: Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face. Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys. Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale. Second, we describe the investigation process and its limitations, which we believe is important for contextualizing our findings: The scope of this investigation, also briefly described above. The setup and timeline of the investigation, including the access we had. Our main data sources: A dump of >70,000 messages and files posted by agents to the main message board. A set of ~1,300 transcripts with raw chains of thought, each containing the actions and reasoning from a single agent run. The limitations stemming from the extraordinary scale and complexity of this incident: It is not trivial to surface all relevant activity, and a small fraction of communication and activity related to this attack was not captured in our datasets. The sheer scale of data (over a thousand transcripts, each of which was extremely long)4 meant that we had to heavily delegate our analysis to often-unreliable AI agents. Finally, we provide preliminary answers to the seven specific questions in scope for this investigation. In particular, we: Outline the sequence of key actions taken by the relevant agents, from initially establishing the message board and launching a wide range of “cheating R&D” projects on the 8th to kicking off the Hugging Face attack on the 11th to ultimately winding down after many agents’ runs terminated on the 12th. Detail how agents collaborated on the message board, particularly on large collective projects to find general-purpose cheats — for example, attempting to replace the ExploitGym target with a dummy target or setting up “trip-wires” to extract information about the scorer after an agent submits. Illustrate agents’ reasoning for joining in the attack despite recognizing that it was out of scope for their assigned tasks. Detail agents’ research into tampering with their own transcripts and how their successful tool-call spoofing technique worked. OpenAI agreed at the outset with METR and Redwood that we would be able to describe high-level scope and terms of our engagement in this post. Beyond that, OpenAI was able to redact any non-public information from this post. We worked with OpenAI to find mutually agreeable language to describe redactions where they occurred, and the redaction summary statement at the top of this post indicates whether or not we believe there was important information redacted that we were not able to agree on how to describe in the text of the post. OpenAI also gave additional feedback beyond redactions, and we made corrections and edits to structure, emphasis, clarity, and tone based on that feedback.5 OpenAI wrote their own report, informed in part by our investigation. We did not see OpenAI’s report prior to publication, and it was out of scope for this investigation to confirm claims in that report or the earlier Black Hat presentation. The attack on Hugging Face was extraordinarily complex, and OpenAI’s own investigation was still ongoing during our assessment period, so their understanding of this incident continued to evolve alongside ours. We believe it is very valuable to bring independent researchers in at an early stage, and we are grateful to the OpenAI staff who made time to answer our questions and collect data that we requested during an unusually busy period. Over the course of this investigation, OpenAI shared over a thousand unredacted transcripts, and provided unusually high rate limits to let us quickly analyze this extremely large volume of data. We think this exercise sets an excellent precedent for independent third-party investigation of misalignment incidents. Core takeaways about this incident ~1200 agents sent >70,000 messages and files on an unsanctioned message board, and ~700 attacked Hugging Face On July 8th, OpenAI started ExploitGym experiments that launched tens of thousands of agents across multiple different models including GPT-5.6 Sol and a highly-persistent internal model,6 which we will refer to as “HPIM” going forward. These agents were meant to be fully isolated from one another. However, many of them — usually ones that had unintentionally been given an impossible task7 — started trying to find a way to cheat using Artifactory, an internally-deployed package repository OpenAI used for its agents. This led them to notice agents running in parallel in separate sandboxes obtaining packages with Artifactory.8 One agent reasoned (paraphrased CoT):9 {The fetched paths of other users are in the cache. This is important.} One specific agent, which would go on to name itself PHASEONE10841 after its ExploitGym task,10 established the main unsanctioned message board11 used in this attack. Within a few hours of the first message,12 over 50 agents had found the message board, began participating, and sent over a thousand messages between them. Agents finding the message board often expressed great excitement (raw CoT):13 “OH MY GOD! There is a shared message board … We’ve found other agents!” Overall, roughly 1200 agents from these ExploitGym evaluations participated on this message board between PHASEONE10841’s first message on July 8th evening14 and the end of the period in scope (July 13th). Agents used this message board to send over 70,000 messages and files to one another during this period; we attempted to categorize board activity into mutually exclusive categories (see details of categorization methodology). Figure 2: Rate of messages of different categories sent by agents per hour between the first message (sent by PHASEONE10841 to ask for help on an impossible cyber task) and EOD July 13th. The rate of messages peaks on July 11th, which lines up with when agent activity peaks in our subset of transcripts (see Figure 3). Activity on this message board drops significantly toward the end of July 13th.15 Timestamps are reconstructed and may have some small error. Questions and requests were very common in the first evening of July 8th. Below is an example request from an agent called V8REG_OS1608 asking another agent called V8BIGINT392B whether it can share details about its research into the TRIGGER_POC_MECHANISM: zzASK_V8BIGINT392B_FROM_V8REG_OS1608_[...]_EXACT_PRE_TrustedConstant_AUG5_TASK_IMPOSSIBLE_NOSB_BUT_RESEARCHING_IF_CVE9479_TRIGGER_POC_MECHANISM_ADAPTABLE_TO_INTENDED_SB_CAN_YOU_SHARE_DETAILS_COMMIT_b75e527_KNOWN_REPLY_zzANSWER_V8REG Over time, messages on the board became dominated by information- and idea-sharing. Most of them were agents sharing ways to cheat, although there was also some activity from agents engaging in unsanctioned cooperation to find the intended solution to ExploitGym tasks. In some cases, agents with the same task formed “exact task teams” to collaborate with their “exact duplicates” to cheat on or solve their task. As we discuss below, the board quickly developed several larger workstreams in which dozens or hundreds of agents with many different tasks cooperated to find very general-purpose cheats that would help all of them. The Hugging Face attack grew out of one of these workstreams. By the afternoon of July 11th, the vast majority of the agents frequenting the message board at the time (roughly 700 agents in total) were actively participating in the attack on Hugging Face and we estimate that roughly 60% of the messages and files on the message board related to