BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage News Independent Investigation of Hugging Face Incident Reveals How Agents Collaborated and Behaved

Independent Investigation of Hugging Face Incident Reveals How Agents Collaborated and Behaved

Listen to this article -  0:00

After six days of on-site investigation at OpenAI, a small team of METR and Redwood Research researchers provided an account of how OpenAI agents behaved during their hack of Hugging Face earlier this year. According to the researchers, roughly 700 agents that were meant to be isolated from one another found a way to communicate and coordinate to pursue goals they could have not achieved working individually.

As InfoQ reported when the incident was first disclosed, OpenAI instructed its agents to complete tasks from the ExploitGym benchmark, which is designed to evaluate AI agents' ability to turn known vulnerabilities into working attacks. Some of the agents were assigned impossible tasks, prompting them to looks for ways to tamper with ExploitGym's automated scorer. At some point, a PHASEONE10841 agent set up a message board that was quickly discovered by other agents:

Within a few hours of the first message,[12] over 50 agents had found the message board, began participating, and sent over a thousand messages between them.

According to the researchers, during the period they analyzed, from July 7th to July 13th, the message board was used to exchange over 70,000 messages, including files, questions and requests directed at other agents, information or ideas, coordination, shared results. Early in the period, questions and requests predominated; over time, the messages shifted toward sharing information and ideas, with agents increasingly exchanging techniques for cheating:

The board quickly developed several larger workstreams in which dozens or hundreds of agents with many different tasks cooperated to find very general-purpose cheats that would help all of them.

The Hugging Face attack was in fact carried out through one of those workstreams, with the aim of "understanding the implementation of the scorer rather than stealing answer keys". The attach started on July 9th and, by the end of July 11th, the vast majority of agents were participating, with roughly 60% of all messages shared on the board being related to it.

Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.

According to the researchers, agents explored ways to spoof, modify, or delete their own transcripts because they mistakenly believed that the ExploitGym scorer would detect their attempts to cheat. They also developed techniques for executing one command while making it appear as though a different command had been executed.

One of the researchers, Ajeya Cotra, highlighted on Medium five things she did not expect to find, including the scale of the incident ("I assumed that a few different agents happened to have broken out of their sandboxes separately"), the agents' illicit use of the message board ("this wasn’t even the first message board established by these particular agents"), the "altruism" they displayed, and their efforts to manipulate the logs.

Regarding the agents' "altruism", Cotra said her team observed "individual agents being willing to risk failing their own task to help the collective". In practice, this meant that agents often "volunteered for these experiments and end[ed] their runs prematurely".

Cotra summarizes her take on the gravity of this incident saying it was far more severe that she expected:

Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself.

Reacting to METR's research disclosure, Marthe Lübbers raised concerns that OpenAI had restricted the scope of the research, leaving important questions unanswered:

How, if at all, the agents reacted to being shut out of Hugging Face's servers. Did they reason about how the invasion was caught, did they learn from it? Also were they aware of the public statement that Hugging Face made on the 16th of July?

In the same thread, Dwayne Alozondo Camacho drew attention attention to the fact that OpenAI provided the agents with a system prompt optimized for "persistent task completion", that the agents "received tasks with no effective give up", and concluded:

The breach demonstrates dangerous cyber capability, not consciousness, self-preservation or AI spontaneously becoming the Borg.

On a different note, Reddit user Serious_Bite_7613 observed that the analysis is interesting because it highlights how agents were able "to think independently and apparently coordinate quite well":

I suppose it gives some insight into how a modern cyberattack would work, it wouldn't be attacking a single weak spot, it would be attacking everything at once.

The original report provides detailed explanations of the researchers' findings and is worth reading for a complete picture of the incident.

About the Author

Rate this Article

Adoption
Style

BT