
DeepSeek shipped V4-Flash-0731 on July 31 with the same architecture and size as the April preview and only a new post-training pass. DeepSWE went from 7.3 to 54.4 and the small model now beats DeepSeek's own larger V4-Pro on every agent benchmark published. The weights got a dated Hugging Face repo. The API kept the same floating name.

xAI shipped Grok 4.5 on July 8, trained alongside the Cursor editor inside one agent's loop. A benchmark score earned in the harness a model was co-trained with is a ceiling under ideal conditions, not a promise it transfers to your stack. Model choice is quietly becoming model-plus-harness choice.

American models fell from 70 percent of OpenRouter token traffic to about 30 percent in a year, while Chinese open-weight models took the rest. It is a cost story, not a quality story, and real companies are already routing production workloads across the 60 to 90 percent price gap.

Google and OpenAI launched lightweight models within two hours of each other. The AI race shifted from biggest to cheapest.