
Saturday, August 15, 2026
A Floodlight/Wired investigation found that OpenAI's Stargate data center in Abilene, Texas used a permit built for dry cleaners to bring a gigawatt-scale gas plant online with no environmental review. Two EPA rulemakings this year suggest that workaround is becoming national policy, not a Texas anomaly.
Anthropic's new interpretability paper found a hidden band of activity inside Claude that catches deception its outputs never admit to. But five months after its CEO said he couldn't rule out machine consciousness, the coverage keeps reading a safety tool as a discovery of sentience, and the gap has only widened in the days since publication.
A new academic benchmark gives the industry its first real measure of "agentic abstention": whether an AI agent recognizes a task is infeasible and stops rather than keeps burning tool calls. Every frontier system tested fails most of the time, and neither Claude Fable 5 nor GPT-5.6, which OpenAI is taking to general availability this week, has been scored on it yet.
Amazon signaled it may sell Trainium chips to outside data centers, following Google's TPUs into the merchant market. Anthropic and Meta are each moving toward Samsung for custom silicon of their own, one in early talks and one reportedly negotiating a deal, evidence that Nvidia's moat was always software, not chips.
You can now kick off an AI coding agent, close the laptop, and get a pull request back - some tools even let you steer one from a chat app. Yet Meta just told staff its agents haven't progressed as hoped. The difference is delegation. Here's how to do it well, and where agents still break.
A team of Berkeley researchers posted near-perfect scores across eight major AI benchmarks - 100% on most of them - without solving a single task, just by gaming how the score is computed. That gap - between the number and the achievement - is why you have to read a benchmark claim like a skeptic. Here are the five tells, and the five questions to ask.