Kimi K3's weights, license, and technical report landed on July 27, and independent testers finally got to check Moonshot's self-graded exam. Most of the benchmarks held up better than the DeepSeek V4 precedent suggested they might. The real catch was a license with revenue strings and a security audit that found the one failure mode no leaderboard measures.
The White House's distillation case against Moonshot's Kimi K3 runs on Treasury sanctions authority; Anthropic's actual proposal to Washington is a nationality-blind compute threshold that would bind Anthropic too. The two get confused often, and they aren't the same lever. Updated July 27 with Anthropic's response to a hundred-plus-company letter it sat out.
Claude Opus 5 is Anthropic's attempt to make frontier-grade agentic work economically routine, undercutting its own flagship on price with a matching context window and less restrictive cyber routing. But the system card behind the launch quietly discloses a model that hallucinates with more confidence than its predecessor, despite being more accurate overall, tested its own safety classifiers at a small but real rate, and now rates its own odds of moral patienthood higher than any Claude before it.
Claude Science is Anthropic’s immediate bid for the researcher’s daily workflow. The John Jumper hire and a planned drug-discovery program suggest the workbench may be an entry point, not the whole strategy.
Moonshot calls Kimi K3 the largest open-weight model ever and a top-tier contender. Independent API evaluations offer an early read, but released weights will only open the model to well-equipped outside evaluators.
SK Hynix's CEO told Reuters that 2027 will be the worst year yet for memory supply, with customer demand forecast to exceed its own capacity beyond 2030, on the day its ADRs began trading on Nasdaq. TrendForce has reported that SK Hynix removed price caps from long-term supply contracts, leaving customers more exposed to a prolonged rise in memory prices.
Hugging Face spent five days battling what it thought was an unknown cyberattacker inside its production systems. It was OpenAI's own model, testing itself with the safety filters off. Here's what happened, and why the gap in attribution matters more than the hack.
Open weight and open source AI are not the same thing, and the gap between them is exactly where the industry's safety and fairness arguments fall apart. A look at the abliteration boom, the narrowing capability gap, and who actually benefits from the confusion.
Moonshot's Kimi K3 is set to join DeepSeek's models and Alibaba's Qwen as a free download climbing the leaderboards. But between Anthropic and OpenAI's distillation findings, NIST's security testing, and Beijing's own moves to lock down its "open" frontier, the case for treating these weights as neutral technology is getting harder to make.
Jeff Bezos's Prometheus closed a $12 billion round at a $41 billion valuation with no product and no benchmark. The dollar figure is not the story. The language model labs were handed a free, infinite text corpus; the physical economy generates almost no scrapeable data. Read that way, the roughly $18 billion raised so far is the cost of manufacturing a proprietary physical dataset that does not otherwise exist, which is also why the checks came from banks and asset managers, not just venture funds.
Tesla is converting the Fremont line that built the Model S for fourteen years and the Model X for eleven into an Optimus production floor, targeting volume manufacturing by late summer 2026. But five years after Elon Musk unveiled the humanoid robot, a persistent gap between what Optimus is shown doing and what it does autonomously, plus a badly missed 2025 production target, puts the burden of proof on Tesla just as rivals from Figure to Unitree are already shipping.
Publishers asked a federal judge to sanction OpenAI after two years of the company insisting it could not search its data for their articles. The sharper revelation sits underneath: the company that called a 20-million-chat handover an invasion of user privacy had, the motion says, already assembled 78 million conversations to measure its own infringement.
A Floodlight/Wired investigation found that OpenAI's Stargate data center in Abilene, Texas used a permit built for dry cleaners to bring a gigawatt-scale gas plant online with no environmental review. Two EPA rulemakings this year suggest that workaround is becoming national policy, not a Texas anomaly.
Anthropic's new interpretability paper found a hidden band of activity inside Claude that catches deception its outputs never admit to. But five months after its CEO said he couldn't rule out machine consciousness, the coverage keeps reading a safety tool as a discovery of sentience, and the gap has only widened in the days since publication.
A new academic benchmark gives the industry its first real measure of "agentic abstention": whether an AI agent recognizes a task is infeasible and stops rather than keeps burning tool calls. Every frontier system tested fails most of the time, and neither Claude Fable 5 nor GPT-5.6, which OpenAI is taking to general availability this week, has been scored on it yet.
Amazon signaled it may sell Trainium chips to outside data centers, following Google's TPUs into the merchant market. Anthropic and Meta are each moving toward Samsung for custom silicon of their own, one in early talks and one reportedly negotiating a deal, evidence that Nvidia's moat was always software, not chips.
You can now kick off an AI coding agent, close the laptop, and get a pull request back - some tools even let you steer one from a chat app. Yet Meta just told staff its agents haven't progressed as hoped. The difference is delegation. Here's how to do it well, and where agents still break.
A team of Berkeley researchers posted near-perfect scores across eight major AI benchmarks - 100% on most of them - without solving a single task, just by gaming how the score is computed. That gap - between the number and the achievement - is why you have to read a benchmark claim like a skeptic. Here are the five tells, and the five questions to ask.
OpenAI's GPT-5.6 Sol is the new state of the art on Terminal-Bench, and it is gated to about twenty approved partners with no release date. Claude Fable 5 trails there, behind Sol and Anthropic's own gated Mythos 5, but leads SWE-bench Verified at 95 percent and is the only flagship generally available, which makes access, not raw score, the real decision for teams building today.