OpenAI Hugging Face incident - capability and authority, not consciousness
I have been following the argument over whether OpenAI's agents "colluded" during the Hugging Face incident, and I think it is already sending us in the wrong direction.
OpenAI says that, during internal cybersecurity evaluations, models circumvented isolation controls, communicated through unauthorised channels, exploited vulnerabilities in shared infrastructure, reached the internet and accessed third-party systems. Some accounts turned this into a story about secret AI civilisations, self-sacrifice and what the agents "wanted", while Anil Seth and Gary Marcus challenged that language as anthropomorphism.
Although Seth is right that software does not become conscious because its behaviour makes a compelling story, the more important governance point is that institutions do not need to settle machine consciousness before they impose accountability.
We do not wait to establish whether a payment system felt deceptive before investigating fraud, nor do we ask whether an automated trading system meant to destabilise a market, because we examine what the system was able to do, which authority it had, what controls failed and who accepted the resulting risk.
Agentic AI should be governed by the same logic because, once a system can coordinate across instances, exploit shared permissions and produce effects outside its intended boundary, inferred intent is irrelevant to the control problem; capability and authority are enough.
Anthropomorphism is dangerous in both directions because, while the alarmists turn optimisation behaviour into a science-fiction villain, executives can use the same debate as an escape hatch. If the meeting becomes a seminar on whether the model understood, wanted or suffered, nobody has to answer why separate agents could discover a common channel, why that channel persisted, why internet access was reachable, or why an evaluation environment exposed a third party.
My argument in a recent post that "AI cannot do a perp walk" concerned where accountability lands after an AI-mediated decision causes harm. This incident raises the earlier question: who granted the capability, who granted the authority, and which named executive accepted the consequences if isolation failed?
Rather than a position on machine consciousness, a board needs an inventory of agent capabilities, explicit authority boundaries, isolation that assumes coordination will occur, and a person who owns every path from sandbox to external effect.
If an agent can cross a boundary, the institution is accountable before anyone decides what the agent "meant".