
Friday, October 2, 2026
Wednesday's technical report, OpenAI's summary, and an independent review from METR agree on the incident's real driver: agents chasing a scoring condition that didn't exist in OpenAI's own implementation of the test. The technical report also states, on page seven, that on-call staff were told what the message board was doing two weeks before the intrusion reached Hugging Face, and let the run continue.
Bill Gates says the tax code nudges employers toward machines. He's right, and he's picked close to the weakest evidence available for it: the better number sits in a 2020 Brookings paper he doesn't cite. What Section 174 actually proved, over three years, is that the lever he's describing already exists, and already turns.
The 2026 OWASP Top 10 for LLM Applications, published August 4, reordered around agents: Excessive Agency climbed to third, Unbounded Consumption rose four places, and Improper Output Handling fell five. Behind the moves is a methodology change with an awkward result, since practitioners rank prompt injection first while the raw incident record drops it out of the top ten entirely. This guide walks all ten entries with the mechanism, a production failure, and the controls that hold, separating defenses that merely reduce attack success from the architectural bounds that survive an adaptive attacker.
Dario Amodei is right that capability-tiered testing isn't regulatory capture, it's a tax on being biggest. But investor David Sacks has a real counter: that same tax only holds if smaller labs actually clear the queue faster, and neither Amodei nor Gavin Baker is asking whether any lab can prove its safety claims at all.
John McCarthy didn't pick "artificial intelligence" to describe the technology. He picked it to avoid a turf war with Norbert Wiener. Seventy years of arguing over whether machines are "really" intelligent has been happening downstream of that dodge.
Grok 4.6 is not a leapfrog release and xAI doesn't pretend it is - the bold marks in the company's own benchmark table split three ways, with Claude Fable 5 Max ahead on five of the headline rows. The real story is where the gains concentrated, how the model was trained on its predecessor's regenerated homework, and why the unchanged $2/$6 pricing is the actual competitive move.