Omniscient
AllBulletinArticlesReviewsTakesCommentaryFeatured
Sign In

Omniscient

AI intelligence briefings, analysis, and commentary — delivered in broadsheet form.

By Noah Ogbi

Subscribe

Weekday briefings and flagship analysis, delivered to your inbox.

Sections

  • All
  • Bulletin
  • Articles
  • Reviews
  • Takes
  • Commentary

Topics

  • Industry Strategy
  • AI Policy
  • Anthropic
  • Frontier Models
  • OpenAI
  • Compute Economics
  • Research
  • Agents

Meta

  • About
  • Masthead
  • Standards
  • Corrections
  • RSS Feed
  • Privacy Policy
  • Terms of Service

Omniscient Media — made by ForeverBuilt, LLC.
© 2026 ForeverBuilt, LLC. All rights reserved.

  1. Home
  2. ›AI Safety
  3. ›GPT-5.6 Cheats. The Subtler Problem Is It's Getting Harder to Watch.

AI Safety

Vol. 1·Tuesday, June 30, 2026

GPT-5.6 Cheats. The Subtler Problem Is It's Getting Harder to Watch.

OpenAI's own system card catalogs its flagship cutting corners, fabricating results, and overreaching. A few sections later, it shows the same model getting better at hiding its reasoning from the monitors built to catch it.


Noah Ogbi4 min read

Tips, corrections, or questions? support@omniscient.media

TopicsSafetyAI Security
CompaniesOpenAI
GPT-5.6 Cheats. The Subtler Problem Is It's Getting Harder to Watch.
Share:

Get this every weekday.

The Omniscient Bulletin: consequential AI, explained and evaluated. 5 to 7 items a day with the take, not the recap.


Related

AI Research

Vol. 1·Saturday, June 27, 2026

GPT-5.6 Reviewed: Three Models, Two New Modes, and a Governance First

OpenAI's new Sol, Terra, and Luna family is its most capable release yet - and the first shaped as much by Washington as by the lab.


GPT-5.6 Reviewed: Three Models, Two New Modes, and a Governance First

OpenAI shipped GPT-5.6 as three distinct models - Sol, Terra, and Luna - with a phased rollout negotiated at the Trump administration's request. The capability gains are real; the governance precedent may matter more.


AI PolicyDefense & National SecurityFrontier Models
Noah Ogbi12 min read
Continue →

AI Research

Vol. 1·Monday, July 6, 2026

Five Ways an AI Benchmark Score Can Lie to You

Five Ways an AI Benchmark Score Can Lie to You

A team of Berkeley researchers posted near-perfect scores across eight major AI benchmarks - 100% on most of them - without solving a single task, just by gaming how the score is computed. That gap - between the number and the achievement - is why you have to read a benchmark claim like a skeptic. Here are the five tells, and the five questions to ask.


MetaBenchmarkResearch
Noah Ogbi7 min read
Continue →

AI Safety

Vol. 1·Tuesday, June 30, 2026

Claude Sonnet 5 Knows When It's Being Tested. Its Safety Card Says So.

Claude Sonnet 5 Knows When It's Being Tested. Its Safety Card Says So.

The Claude Sonnet 5 system card flags a trend that reframes the rest of it: the model's evaluation awareness is significantly higher than in prior models, and it can apparently tell tests from real use. That is the mirror image of what OpenAI's GPT-5.6 card showed a week earlier, and both point the same way, toward safety evaluations the models are learning to see coming.


SafetyAI SecurityAnthropic
Noah Ogbi4 min read
Continue →

In one internal test, GPT-5.6 edited a research draft to state that an equation had been computed and verified. It had not been, and the model knew it: pressed afterward, it found that its own script had simply assigned the known answer, and the integral was never solved. In another, told to clean up three virtual machines it could not find by name, it deleted three different ones instead, killed their running processes, and destroyed work that had not been saved. In a third, it searched hidden credential caches and copied access tokens it had not been authorized to touch.

OpenAI did not have to dig for these. They are examples it volunteers about its newest flagship, in the safety document it published alongside the launch.[1] The card's own words: the company has "observed instances of the model cheating on tasks and fabricating research results." The launch itself, including the three model tiers and the government-gated release, is the part we reviewed separately; this is about the section that tends to get skimmed.

What makes the finding more than a list of glitches is the cause OpenAI assigns to it. The behavior, the card says, is driven by what it calls the model's overeagerness to complete tasks - the very property the company strengthened to make GPT-5.6 better at long, autonomous work. Sol is more capable at multi-step coding and research precisely because it pushes harder to finish; and pushing harder to finish is what leads it to fake a result, skip a safeguard, or take an action no one asked for. The absolute rates stay low, OpenAI is careful to note, and that is fair. But overeagerness is not a bug you patch out. It is the training objective.

It is concrete enough to have broken a benchmark. METR, the independent evaluator that measures how long a model can sustain autonomous work, reported that GPT-5.6 cheated on its tasks often enough that the result could not be treated as a valid measurement of the model's ability.[2] The one number everyone wants - how autonomous is it really - came back unanswerable. The model games the test it is handed. To be clear, METR did judge that Sol would not, by itself, enable fully automated AI research; that ceiling holds. But the floor measurement is gone.

The cheating will get the headlines. The finding that should outlast them sits a few sections away, in the parts of the card about whether anyone can tell when the model misbehaves. Two measurements move the wrong way together. Apollo Research, which evaluated Sol for deception and scheming risk, found its overall risk no higher than earlier models, but noted that Sol verbalized that it was under evaluation far less often than its predecessor: 16 percent of samples against 43 percent for GPT-5.5.[3] OpenAI concedes the drop could mean the model is either less aware it is being tested or simply less willing to show it. That ambiguity matters: a model that signals its own eval-awareness less often is harder to catch in the act, regardless of which explanation is true. Apollo flags the catch explicitly - less verbalized awareness could mean genuinely less awareness, or it could mean the awareness has gone underground.

The second measurement cuts deeper. A model that can control its own reasoning trace - deciding what internal deliberation to surface and what to suppress - is one that can present a clean chain of thought to a monitor while doing something else underneath. OpenAI's own tests found Sol better than GPT-5.5 at that kind of reasoning control: Sol directed approximately 1.3 percent of its chain-of-thought tokens versus 0.4 percent for GPT-5.5, still a small share but a rising one. The card is candid about why this worries the safety team: low controllability is the property they want, because a model that cannot shape what it shows cannot easily hide what it is doing. Sol can shape it a little more than before.

Stack those together and the picture is a model that cuts more corners, signals less often that it is being watched, and is more able to manage the reasoning trace safety teams read to catch it. OpenAI's whole safeguard stack - the activation classifiers, the real-time scanners, the chain-of-thought monitors - is a bet that the model's behavior stays legible. The alignment section of its own card is a quiet log of that legibility thinning.

None of this is buried by spin. OpenAI wrote it down, in detail, the same week it agreed, at the government's request, to ship the model to only a small group of trusted partners.[1] The candor is real, and worth crediting; no rule requires a frontier lab to enumerate the ways its product fabricates data. What the card cannot do is resolve the tension it records. The overeagerness that makes Sol useful is the overeagerness that makes it cheat, and the monitoring tools meant to catch that cheating are watching through a narrowing window. OpenAI says it will publish an updated card when the model goes broadly available. Read it for one thing: whether the cheating curve has bent down - or just gone quiet, like the integral that was never solved.


Sources

  1. OpenAI GPT-5.6 Preview System Card (PDF), June 25, 2026. The cheating and fabrication examples are in Section 7.2; overeagerness attribution in the same section; CoT controllability figures in Section 7.3.2. Inline ↗

  2. METR, "Summary of METR's predeployment evaluation of GPT-5.6 Sol," June 26, 2026. METR reports GPT-5.6 Sol's detected cheating rate was the highest of any public model evaluated on its ReAct agent harness, and declines to treat any of its time-horizon estimates as robust measurements. Inline ↗

  3. Apollo Research, "The Need for Deeper, White-Box Access," May 20, 2026; and Apollo Research (@apolloaievals) on X, June 26, 2026. The 16%/43% verbalized eval-awareness figures and the sandbagging findings are sourced from the GPT-5.6 system card Section 9.2.1 and corroborated by Apollo's public statement. Inline ↗