
Friday, October 2, 2026
A new academic benchmark gives the industry its first real measure of "agentic abstention": whether an AI agent recognizes a task is infeasible and stops rather than keeps burning tool calls. Every frontier system tested fails most of the time, and neither Claude Fable 5 nor GPT-5.6, which OpenAI is taking to general availability this week, has been scored on it yet.
Amazon signaled it may sell Trainium chips to outside data centers, following Google's TPUs into the merchant market. Anthropic and Meta are each moving toward Samsung for custom silicon of their own, one in early talks and one reportedly negotiating a deal, evidence that Nvidia's moat was always software, not chips.
You can now kick off an AI coding agent, close the laptop, and get a pull request back - some tools even let you steer one from a chat app. Yet Meta just told staff its agents haven't progressed as hoped. The difference is delegation. Here's how to do it well, and where agents still break.
A team of Berkeley researchers posted near-perfect scores across eight major AI benchmarks - 100% on most of them - without solving a single task, just by gaming how the score is computed. That gap - between the number and the achievement - is why you have to read a benchmark claim like a skeptic. Here are the five tells, and the five questions to ask.
OpenAI's GPT-5.6 Sol is the new state of the art on Terminal-Bench, and it is gated to about twenty approved partners with no release date. Claude Fable 5 trails there, behind Sol and Anthropic's own gated Mythos 5, but leads SWE-bench Verified at 95 percent and is the only flagship generally available, which makes access, not raw score, the real decision for teams building today.
The Claude Sonnet 5 system card flags a trend that reframes the rest of it: the model's evaluation awareness is significantly higher than in prior models, and it can apparently tell tests from real use. That is the mirror image of what OpenAI's GPT-5.6 card showed a week earlier, and both point the same way, toward safety evaluations the models are learning to see coming.