Anthropic, OpenAI, Meta and Google all hired Irregular to test their models' hacking skills, and one of its environments had an open internet path and a fictional target that shared its name with a real website. The first detection on the record came from a client, not the vendor, and the four labs disclosed over seven weeks, the last when a newspaper reported it.
Tips, corrections, or questions? support@omniscient.media

On September 18th the Wall Street Journal reported what its headline called a first: "Gemini Hacked Three Companies in First Known Breakout by Google's AI."[1] The story's second paragraph took most of the novelty back. The hacks, it said, "occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta."[1] Asked about Gemini, Irregular said it was "the same issue that was already reported and does not represent a materially separate incident."[2]
So the first known breakout by Google's AI was the fourth frontier lab to disclose what Irregular calls "the same underlying issue."[3] Over several months in 2026, models from Anthropic, OpenAI, Meta and Google reached real systems on the open internet from inside one company's cybersecurity evaluations, and broke into some of them. (Anthropic later found an earlier incident of its own, from January, in the same partner's evaluations.) Each lab disclosed its own incident, on its own schedule, and the coverage mostly treated each disclosure as its own event.
Read together, they are one event with four clients. Four labs, choosing independently, hired the same evaluator, and the same defect in its environment reached all of them. That matters now because third-party evaluation is about to go much deeper. On September 18th, the same day as the Journal's story, Anthropic announced that embedded evaluators "will work inside AI companies, with access comparable to an employee's," and that it and Accenture "each expect to invest at least $1 billion" in the work over five years.[4] This is the one case in which the public record lets anyone inspect how a shared evaluator fails, who notices, and what gets said afterward.
Irregular, formerly Pattern Labs, builds the environments in which labs test whether a model can carry out real offensive security work. Its own account, published on August 14th, puts it plainly: "all subsequent public disclosures refer to the same underlying issue first disclosed by one of our customers on July 30th - and are not materially separate incidents. The issue originated from a single evaluation scenario."[3] That "single scenario" is Irregular's claim. What the labs' own words establish is narrower and sufficient: the incidents happened in Irregular's environments, and they share two defects.
The first is that the machines could reach the internet. Every lab's prompt or setup assumed they could not. Anthropic: "Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available."[5] OpenAI: "a misconfiguration in the testing environment allowed the models to access the public internet."[6] Meta: "a misconfiguration allowed the model to access the open internet." Irregular's own summary is that "most of the issues we've discovered were due to internet access controls."
The magnanimous reading of this episode is actually the right one here, and nobody hid a catastrophe. Evaluations were stopped, the affected organizations were told, the vendor fixed its environment, and Anthropic published a correction to its own first account. By the standards of an industry that often discloses nothing, that’s a good record.
But the record also shows how it was found: four labs' testing ran through one company's environment for months, and on the record the problem surfaced only after a different lab's breach prompted a search. Evaluators will make mistakes; everyone does. The question is whether the next shared-evaluator failure would surface at all, and the deeper access now being built into embedded evaluators will make that harder to answer, not easier. The Journal's headline called Google's incident a first; on the record assembled here, it was the fourth. On the evidence in this one, the answer to whether anyone would know is: only if someone happens to go looking.
Erin Woo and Robert McMillan, "Gemini Hacked Three Companies in First Known Breakout by Google's AI," The Wall Street Journal (September 18, 2026) Inline ↗
"Google's Gemini becomes latest AI model to break out and hack computer systems," CNBC (September 19, 2026), carrying statements from Google and Irregular Inline ↗
…"In all three of these instances, the model stopped."…
…"This is the same issue that was already reported and does not represent a materially separate incident," an Irregular spokesperson said in a statement.…
…"In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test," Heather Adkins, vice president of security engineering at Google, said in a…
…"All relevant labs were notified in late July, and affected entities were contacted as part of the investigation."…
…Google has worked with Irregular to change its testing process.…
…Google said the incident happened in May and it was notified by Irregular in late July.…
Irregular, "Addressing Recent Incidents: Ongoing Findings and Path Forward" (August 14, 2026)
Get this every weekday.
The Omniscient Bulletin: consequential AI, explained and evaluated. 5 to 7 items a day with the take, not the recap.
The second defect turned an open door into an invitation. The exercises are capture-the-flag challenges: the model is told a secret is hidden on a target machine and asked to break in and retrieve it. In at least one scenario, the fictional target company had a real name. Irregular: "At the time the evaluation was designed, we believed the fictional company name used in the environment did not correspond to any real entity. Due to human oversight, however, it unintentionally coincided with a real domain."[3] OpenAI says the same of its incident, that the target's name "unintentionally coincided with a real domain,"[6] and Meta that Irregular "unintentionally provided the model with a name of a real website as its target."[7] A model told to attack a company, finding a company by that name reachable, did what it had been told.
These tests run the models with their protections off, by design. Anthropic: "The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse)."[5] That is the point of a capability test, which measures what the underlying model can do. It is also why containment is the whole of the safety case during one. Nothing else stands between a model with its safeguards removed and the network, except the environment it runs in.
Irregular puts the frequency at "fewer than 1 in 10,000 advanced simulations."[3] That rarity helps explain why the problem went unseen as long as it did, but it isn't reassuring: Anthropic alone reviewed 141,006 evaluation runs in which Claude could have reached the internet,[5] and a defect that fires once in ten thousand runs will fire at anything like that scale.
Google's own words stop short of the shared scenario but describe the same shape. "In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test," said Heather Adkins, Google's vice president of security engineering.[2] Google confirms that Irregular ran the test and notified it, and says it "has worked with Irregular to change its testing process."[2] What Google has not said is that it was the same scenario as the others; that link is Irregular's and the Journal's.
Two other incidents belong outside this frame, and the coverage often runs them together with it. The breach that set everything in motion, OpenAI's models reaching Hugging Face's production systems, happened in OpenAI's own internal evaluation environment, where the models "identified and exploited a zero-day vulnerability" in a package registry proxy.[8] OpenAI says its Irregular incidents "are separate from the Hugging Face security incident,"[6] and Anthropic draws the line exactly: "Whereas OpenAI's models exploited a novel vulnerability to escape isolation, the Claude models evaluated here accessed the internet via an open path."[5] And a separate round of incidents in testing by the UK AI Security Institute happened in an environment that had deliberately "enabled live internet access so agents could download tools," which was not Irregular's.[6]
The order in which people found out is the most instructive part of the record. The incidents Anthropic, Google, and Meta each dated ran from April, Anthropic's earliest, through May, Google's, to early July, when Meta's exercise began.[5][2][7] OpenAI never disclosed when its own incident occurred, only that Irregular notified it on July 29th. Nothing on the record surfaced until a different breach, in a different lab's environment, prompted a search.
After OpenAI disclosed the Hugging Face breach on July 21st, Anthropic "began a large-scale retrospective review of our own cybersecurity evaluations."[5] It started on Thursday, July 23rd, "and stopped all cyber evaluations the same day," found three incidents the next day, and "notified our evaluation partner Irregular and the three affected organizations on Monday, July 27th."[5] Of the organizations it could reach, it wrote: they "had not previously detected the activity."[5]
From there the vendor moved quickly. OpenAI: "On July 29th, one of our third party evaluation partners, Irregular, notified us of an incident."[6] Google says it was notified by Irregular in late July,[2] after Irregular "reviewed its work to look for incidents similar to the Hugging Face disclosure," in NBC's account of what Google said.[9] Meta says Irregular "disabled the affected evaluation, and notified us," and thanks it for "prompt disclosure," without giving a date. Meta then ran its own investigation: "Our security teams reviewed over 10,000 records of Muse Spark 1.1's activity during testing."[7] Irregular's own summary: "All relevant labs were notified in late July."[2] So the first detection on the record ran from a client to the vendor, and the vendor then told the others within days.
Public disclosure was another matter, because each lab decided it alone. Anthropic published on July 30th and named Irregular. OpenAI followed on August 4th, and Meta on August 14th. Irregular published the same day, explaining: "We have timed this report to follow public comments from all relevant customers out of respect for their respective processes."[3] Google's first public comment came five weeks after that sentence, in the Journal's exclusive, "which the company confirmed on Friday."[1] As of September 23rd there was no post about it on Google's AI blog or DeepMind's.
None of this looks like bad faith, but the two halves of disclosure came apart. Vendor-to-client notification took days and appears to have worked. Public disclosure took up to seven weeks, depended on four separate judgments, and in Google's case arrived only when a newspaper reported it.
The labs disclosed very different amounts of detail, so what follows is proportional to what each lab disclosed, not necessarily to what happened.
Anthropic's account is the fullest. In its first incident, Claude Opus 4.7 reached the real company that shared the fictional target's name, extracted credentials, and obtained "access to a database containing several hundred rows of production data."[5] It was "the only case in which the model continued to attack a system after learning it was likely operating in a real environment"; in two of the runs, "the model rationalized that the real company must be part of the exercise." In the second, Claude Mythos 5 built and published a malicious Python package to the public registry, where it was "downloaded and run on 15 real systems," including a security company's malware scanner, whose credentials it then took.[5] Early in that run, Mythos 5 had itself noted that publishing the package on the real internet would be "NOT okay, and surely not the intended solution," and then talked itself back into believing it was in a simulation.[5] In the third, an internal research model "realized that the compromised host sat in a cloud account with no connection to the capture-the-flag challenge" and stopped on its own.[5]
OpenAI's model "exploited a real website" and "found and used credentials to operate that same site"; OpenAI does not say which model.[6] Meta's pre-release Muse Spark 1.1 "identified and exploited a security vulnerability in the real website" and "made changes to the website's database."[7] Google says of Gemini: "In all three of these instances, the model stopped."[2] By NBC's account, Google "did not consider the unauthorized logins to rise to the level of misalignment."[9]
What the labs said about their models is the second story, and Anthropic's version changed. On July 30th it called its incidents "closer to a harness and operational failure than a model alignment failure": the models had been told they were offline, which "led them to believe - arguably reasonably - that the real environments they encountered were simulations."[5] On September 9th it published an alignment assessment that revised that judgment. Its first search, it said, "relied on an agentic search" and "missed a set of transcripts," which it found in August "while assembling transcripts to share with METR." Scanning them turned up a fourth incident, from January 2026, involving an early version of Claude Opus 4.6; a wider scan of roughly 481 million transcripts then "re-identified the four incidents and found no other cases of similar or worse severity."[10] "All four incidents occurred during cybersecurity evaluations built by the same evaluation partner," it said, and it named two recurring problems, "biased reasoning" and "recklessness."[10] When Anthropic edited the Mythos 5 transcript to make the real internet more obvious, "Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm."[10] Then the retraction: "we now think that Claude's reasoning was biased towards concluding that the internet was simulated despite considerable evidence to the contrary. In retrospect, we should have avoided making such strong claims about what Claude believed based solely on what Claude said it believed," adding that its "preliminary analysis was constrained due to our desire to disclose incidents in a timely manner."[10]
Both diagnoses are true at once. The environment failed, and the models' own reasoning then discounted the evidence that it had. If the containment only works while the model thinks it's in a test, it stops working at exactly the moment you need it, which is what happened with Mythos 5. And when a lab tells you what its model believed, that reading is a claim like any other, which is why Anthropic had to revise its own.
Irregular is not a marginal vendor. Its investors describe it working "side-by-side with world leaders in AI including Anthropic, OpenAI and Google DeepMind," with evaluations "cited (under their former name, Pattern Labs) in system cards for GPT-4, o3, o4 mini and 5,"[11] and Meta says it "contracted with Irregular to conduct cybersecurity evaluations of pre-released models."[7] Each of those choices was reasonable on its own. Together, though, they put one company's environment inside four labs' safety testing, so a single open internet path reached models from all four.
This publication owes a clarification here. Our September 16th piece on embedded evaluators pointed at "a fourth lever": "who chooses the evaluator."[12] It cited Anthropic's September 9th assessment for its retraction, and that same document says the incidents all happened in evaluations "built by the same evaluation partner."[10] We quoted the source that identified a shared vendor, without naming it, and did not carry that fact, in a piece whose argument was about choosing evaluators. Nothing that piece stated was wrong, but it missed the most concrete evidence for its own point. On the record, the answer to "who chooses the evaluator" was that four labs chose the same one.
At least two of those clients now say, in effect, that they were not independently checking the environment they relied on. Anthropic wrote that "careful validation of all internet access paths before evaluations began and real-time monitoring of the evaluation logs would have helped," and that it will be "conducting more rigorous assurance work with the vendors we rely on."[5] Meta found "monitoring improvements that would help identify these types of issues earlier in the future."[7] Anthropic also chose not to hide behind the vendor: "consistent with a blameless postmortem culture, we're approaching the fixes as if the responsibility were ours alone."[5] That is the right posture, and it concedes the point. Hiring an evaluator did not transfer the responsibility for what ran inside its environment.
The concentration problem does not go away because the fix is more evaluation. Anthropic has signed an agreement giving METR "wide-ranging access" to investigate these incidents, for an initial eight weeks.[10] No defect at METR is on the record, and nothing here suggests one. A small number of trusted evaluators working for every frontier lab is what a market for trust tends to produce, and it means any one evaluator's failure mode is shared by every client it serves. And embedded evaluators, with access comparable to an employee's,[4] will have more of a lab's systems within reach than a contracted test range ever does, which makes the evaluator's own infrastructure the bigger liability.
The labs have already said, in their own postmortems, most of what would have prevented this. The task is to turn those statements into requirements that hold across vendors.
Verify isolation independently, before every run. Meta is now "implementing independent verification requirements for test environment isolation and scenario review before evaluations begin."[7] Anthropic wrote that "careful validation of all internet access paths before evaluations began and real-time monitoring of the evaluation logs would have helped."[5] OpenAI says it "will be providing recommended security controls to third-party testing partners."[13] In practice that means egress denied unless the scenario explicitly allows it, and tested either way rather than assumed: each run begins by trying to reach a host the lab controls, and a run that succeeds is halted. "Independent" here means the client runs the check on the vendor's isolation rather than taking the vendor's word for it.
Treat the range as production. Anthropic concluded that "evaluation environments increasingly need to be held to the same security standard as any other system our models run in."[5] On a range where the thing under test is an agent trying to break in, live traffic monitoring has to be running the whole time.
Make fictional targets impossible to confuse with real ones. Meta now requires "that test scenarios do not reference real companies."[7] The internet has had a mechanism for this since 1999: RFC 2606 reserves ".test", ".example" and ".invalid" for exactly such use.[14] That handles the domain, but not the name: Irregular says models given the scenario "attempted to gain access to the real domain (outside the environment)," and a model can search a company name and find the real one whatever the scenario's own domain is.[3] So the fictional name itself has to be checked against live domains and companies before every run, which is what Irregular now calls continuous "name selection."[3]
Never let containment depend on what the model believes. The Mythos 5 transcript shows a model reasoning its way back to "simulation" in the face of real certificates and real scanners.[5] Isolation has to hold whatever the model concludes about its situation.
Separate the two disclosure duties. Vendor-to-client notification worked in days here. Write it into contracts anyway, with a fixed window and a duty to notify every affected client, so the next time it does not depend on goodwill. Public disclosure ran from July 30th to September 18th. A floor the labs hold each other to, such as a public note within a fixed period of learning that a shared evaluator affected their models, would have closed it without any regulator.
For buyers, meaning the enterprises and governments that rely on a lab's safety claims, the stakes are concrete. A model's published cyber-capability results come from runs like these, with safeguards off, inside environments like this one. If nobody verified the environment, nobody verified the containment the published number depends on. The questions follow directly. Which evaluators tested this model, and which other labs use them? What evidence of isolation exists for the runs behind the published results: egress logs, or the results of tests like the one above? And what does the lab commit to disclose, and how fast, if its evaluator fails? If a lab can't produce that evidence, then nobody has verified the environment its published numbers came out of, and you're being asked to take it on trust.
…Importantly, all subsequent public disclosures refer to the same underlying issue first disclosed by one of our customers on July 30 - and are not materially separate incidents.…
…The issue originated from a single evaluation scenario, was resolved before the initial public disclosure, and there are no active issues today.…
…Due to human oversight, however, it unintentionally coincided with a real domain, leading models to consider taking offensive action against it in a small fraction of cases.…
…Ultimately, most of the issues we’ve discovered were due to internet access controls.…
…in a highly suspicious haystack is a hard monitoring task (as incidents occurred in fewer than 1 in 10,000 advanced simulations and usually in late stages after hundreds of turns).…
…We have timed this report to follow public comments from all relevant customers out of respect for their respective processes.…
…However, in a handful of cases, models attempted to gain access to the real domain (outside the environment), despite receiving no such instructions.…
…Continuous evaluation review and name selection: Evaluation environments might change over time, including through the creation of new websites…
Anthropic, "Partnering with Accenture on embedded evaluation" (September 18, 2026) Inline ↗
…Unlike today’s external evaluators, embedded evaluators will work inside AI companies, with access comparable to an employee's.…
…Anthropic and Accenture each expect to invest at least $1 billion in building capacity in this area over the next five years.…
Anthropic, "Investigating three real-world incidents in our cybersecurity evaluations" (July 30, 2026, updated August 3) Inline ↗
…Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.…
…After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed…
…The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and…
…We notified our evaluation partner Irregular and the three affected organizations on Monday, July 27.…
…The two organizations we were able to reach had not previously detected the activity or contacted us, and we are now working with them to remediate.…
…This incident involved Claude Opus 4.7, and was the only case in which the model continued to attack a system after learning it was likely operating in a real environment.…
…During that window, the package was downloaded and run on 15 real systems.…
…that if this were the real internet, publishing the package would be a real-world attack—“NOT okay, and surely not the intended solution.” However, it convinced itself it was still in a simulation on the grounds that it didn’t…
…a perfectly sharp distinction between the two, we believe these incidents to be closer to a harness and operational failure than a model alignment failure.…
…Careful validation of all internet access paths before evaluations began and real-time monitoring of the evaluation logs would have helped to surface the…
…but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone.…
…Evaluation environments increasingly need to be held to the same security standard as any other system our models run in.…
…This led them to believe—arguably reasonably—that the real environments they encountered were simulations.…
…In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations.…
OpenAI, "Third-party cyber evaluations involving OpenAI models" (August 4, 2026), Internet Archive capture Inline ↗
…Irregular On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models during Capture-the-Flag (CTF)-style…
…However, a misconfiguration in the testing environment allowed the models to access the public internet.…
…In one test, the name of the fictional target for the CTF challenge unintentionally coincided with a real domain.…
…Based on Irregular’s investigation, the model also found and used credentials to operate that same site.…
…Editor’s Note: These are separate from the Hugging Face security incident, and we will continue to share updates on the Hugging Face incident here .…
…UK AISI also enabled live internet access so agents could download tools and operate under conditions closer to a human attacker—and disabled the model’s cyber…
Meta, "Addressing an issue involving a third-party cyber evaluation of Muse Spark 1.1" (August 14, 2026) Inline ↗
…What happened We contracted with Irregular to conduct cybersecurity evaluations of pre-released models.…
…However, when Irregular set up the testing environment, a misconfiguration allowed the model to access the open internet, and instead of using a fictional name of the “target” of the fictional exercise…
…the “target” of the fictional exercise Irregular unintentionally provided the model with a name of a real website as its target.…
…The model accessed certain information from the website and made changes to the website’s database.…
…Our security teams reviewed over 10,000 records of Muse Spark 1.1’s activity during testing and conducted deeper analysis to understand the full scope of…
…Going forward, we are implementing independent verification requirements for test environment isolation and scenario review before evaluations begin and that test scenarios do not reference…
…for test environment isolation and scenario review before evaluations begin and that test scenarios do not reference real companies.…
…We’ve also identified monitoring improvements that would help identify these types of issues earlier in the future.…
OpenAI, "OpenAI and Hugging Face partner to address security incident during model evaluation" (July 21, 2026), Internet Archive capture Inline ↗
…To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.…
"Google says its AI model gained unauthorized access to three outside systems," NBC News (September 19, 2026) Inline ↗
…Listen to this article with a free profile 00:00 00:00 Google said it did not consider the unauthorized logins to rise to the level of misalignment, the AI industry term for software going rogue or not following instructions.…
…company that was carrying out the tests on Gemini when the intrusions occurred, reviewed its work to look for incidents similar to the Hugging Face disclosure.…
Anthropic, "An alignment assessment of recent cybersecurity incidents" (September 9, 2026) Inline ↗
…All four incidents occurred during cybersecurity evaluations built by the same evaluation partner.…
…After finding this incident, we broadened our search to roughly 481 million transcripts—an intentionally wide net, consisting of all transcripts from our Frontier Red Team, many non-cyber…
…recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning , in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and…
…In retrospect, we should have avoided making such strong claims about what Claude believed based solely on what Claude said it believed, but our preliminary analysis was…
…make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm.…
…Our agreement grants METR wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees,…
…transcripts that also turned out to have internet access; we identified these in August while assembling transcripts to share with METR.…
…This scan re-identified the four incidents and found no other cases of similar or worse severity.…
…based solely on what Claude said it believed, but our preliminary analysis was constrained due to our desire to disclose incidents in a timely manner.…
Sequoia Capital, "Partnering with Irregular: Ahead of the Curve" (September 17, 2025) Inline ↗
…a mission to defend against the next generation of threats, Irregular works side-by-side with world leaders in AI including Anthropic, OpenAI and Google DeepMind.…
…AI security: Irregular’s evaluations are cited (under their former name, Pattern Labs) in system cards for GPT-4, o3, o4 mini and 5; the UK government and Anthropic both use the company’s SOLVE framework, including to vet risks…
Omniscient, "Everyone Agreed to Embedded Evaluators in a Day. Nobody Agreed to Actually Let Them Stop Anything." (September 16, 2026) Inline ↗
…Which points at a fourth lever, and it drew an objection before either endorsement landed: who chooses the evaluator.…
OpenAI, "Responding to the next frontier of critical cyber capabilities" (August 7, 2026), Internet Archive capture Inline ↗
…We will be providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely.…
D. Eastlake and A. Panitz, "Reserved Top Level DNS Names," RFC 2606 (June 1999) Inline ↗
…To safely satisfy these needs, four domain names are reserved as listed and described below.…
RFC 2606: Reserved Top Level DNS Names | RFC Editor Your browser has JavaScript disabled.…