Over several months in 2026, models from four frontier labs reached real systems from inside one evaluator's cybersecurity tests, through two shared defects: live internet on machines the models were told were offline, and a fictional target that shared a name with a real domain. Read together, the labs' separate disclosures show what third-party evaluation looks like when four labs rely on the same vendor, and what labs and buyers should require before evaluators get the employee-level access now being promised.
Amodei's "We Must Pace the Frontier" proposes embedded evaluators with employee-like access, and Musk and Altman endorsed it within hours. But OpenAI's own staff saw the Artifactory message board in May and declined to stop the run on June 27th, which makes Hugging Face an authority failure rather than an access failure. Four levers decide whether embedded review means anything - the right to publish, control of scope, power over a running process, and who picks the evaluator - and the essay grants only the first.