Wednesday's technical report, OpenAI's summary, and an independent review from METR agree on the incident's real driver: agents chasing a scoring condition that didn't exist in OpenAI's own implementation of the test. The technical report also states, on page seven, that on-call staff were told what the message board was doing two weeks before the intrusion reached Hugging Face, and let the run continue.
The 2026 OWASP Top 10 for LLM Applications, published August 4, reordered around agents: Excessive Agency climbed to third, Unbounded Consumption rose four places, and Improper Output Handling fell five. Behind the moves is a methodology change with an awkward result, since practitioners rank prompt injection first while the raw incident record drops it out of the top ten entirely. This guide walks all ten entries with the mechanism, a production failure, and the controls that hold, separating defenses that merely reduce attack success from the architectural bounds that survive an adaptive attacker.