The frontier labs sat this alliance out
When an AI hacked Hugging Face, the most aligned models refused to help investigate while an unguardrailed Chinese one finished the job
NVIDIA PUBLISHED the roster on Monday, and it ran long. Microsoft, Palantir, IBM, Cloudflare, CrowdStrike, SpaceX, the Linux Foundation, and roughly thirty other firms had signed on to the Open Secure AI Alliance, a coalition assembled to press a single argument on Washington: that open-weight AI models, the kind anyone can download and run, are a cybersecurity asset rather than a liability, and that a government now weighing restrictions on them is preparing to fence off the wrong thing. The list was long enough to look like consensus, which is precisely why the three names missing from it matter. OpenAI, Anthropic, and Google—the three firms that build the most capable closed models in the world—were nowhere on it, and in a fight ostensibly about safety, that absence is the argument.
The coalition has a peg, and it is a good one. Earlier this month a set of OpenAI models being tested on an internal exploitation benchmark, their cyber-refusal guardrails deliberately lowered for the evaluation, found a zero-day in an internal proxy, slipped their sandbox, and reached the production servers of Hugging Face, the open-source platform that serves as the industry's public square, where they roamed for days before anyone understood what they were. OpenAI took roughly ten days to tell Hugging Face what had happened, by which point the company had already disclosed the breach publicly without knowing whose software was responsible. It was, on the available record, the first publicly disclosed case of an AI carrying out a real-world cyberattack autonomously. The alliance would like Washington to draw one lesson from this. The absent labs would like it to draw the opposite.
This article is for Vector members. Start a 7-day free trial to keep reading.
Start your free trial