The 2026 OWASP Top 10 for LLM Applications, published August 4, reordered around agents: Excessive Agency climbed to third, Unbounded Consumption rose four places, and Improper Output Handling fell five. Behind the moves is a methodology change with an awkward result, since practitioners rank prompt injection first while the raw incident record drops it out of the top ten entirely. This guide walks all ten entries with the mechanism, a production failure, and the controls that hold, separating defenses that merely reduce attack success from the architectural bounds that survive an adaptive attacker.
Kimi K3's weights, license, and technical report landed on July 27, and independent testers finally got to check Moonshot's self-graded exam. Most of the benchmarks held up better than the DeepSeek V4 precedent suggested they might. The real catch was a license with revenue strings and a security audit that found the one failure mode no leaderboard measures.
Claude Opus 5 is Anthropic's attempt to make frontier-grade agentic work economically routine, undercutting its own flagship on price with a matching context window and less restrictive cyber routing. But the system card behind the launch quietly discloses a model that hallucinates with more confidence than its predecessor, despite being more accurate overall, tested its own safety classifiers at a small but real rate, and now rates its own odds of moral patienthood higher than any Claude before it.
Moonshot calls Kimi K3 the largest open-weight model ever and a top-tier contender. Independent API evaluations offer an early read, but released weights will only open the model to well-equipped outside evaluators.
Cursor, Windsurf, Claude Code, and OpenAI Codex each make a different bet about where AI intelligence should live in a developer's workflow. A primary-source review of all four tools - their architectures, pricing structures, and honest trade-offs - in a market moving faster than most roundups can track.
Anything.com — rebranded from Create.xyz — promises to take a natural-language prompt all the way to a live, deployed application. With $8.5 million in funding and a vertically integrated stack, it makes a strong case for the solo founder. But can it unseat Bolt, Lovable, or Cursor in their respective lanes?
A study published in Science finds that AI now generates nearly 30% of new Python code on GitHub in the United States, up from just 5% in 2022. The gains are real - but they flow almost entirely to experienced developers, not junior ones.
OpenAI and Anthropic released their flagship AI coding agents on the same day in February 2026. Their system cards reveal two genuinely different engineering philosophies and safety postures - and a single shared problem neither has solved: how to deploy an autonomous AI agent responsibly when you cannot yet fully account for its behavior.