
On August 18, OpenAI cut the price of GPT-5.6 Sol by 50% on aggregate platforms OpenRouter and Vercel AI Gateway, bringing input to $2.50 per 1M tokens and output to $15 per 1M tokens for one month. The model supports a 1M-token context and targets complex reasoning and coding. SemiAnalysis noted the move may be a marketing play — aggregators hold a small share of OpenAI's token volume, yet they are the primary data source third parties use to estimate its market share.
The same day, a Stanford team reported that sampling five answers from DeepSeek V4 Flash and letting the model score its own outputs lifted Terminal-Bench 2.1 accuracy from 79% to 88%, surpassing Claude Fable 5 at roughly 1/11 the cost. The framework is open-sourced. Alibaba's Qwen team confirmed that its 27B open model, Qwen3.8, runs locally on an RTX 3090 at Q5 quantization and approaches GPT-5.6 on single-prompt generation.
Price-cutting at the top and low-cost open models closing the gap arrived in the same 24-hour window. The point of differentiation is shifting from "stronger model" to "cheaper inference."
On August 17, Groq announced a $350 million Series A at a $3.5 billion valuation, led by Disruptive with NVIDIA planning to participate. The company has shifted from AI chip developer to Neocloud provider. It signed a $20 billion non-exclusive licensing deal with NVIDIA last December, covering LPU inference ASIC authorization, and this August became an NVIDIA cloud partner. Compute capacity is projected to grow from 54MW today to over 200MW by 2027.
NVIDIA CEO Jensen Huang said on the same day that the next critical scarce resources for AI factories are land, electricity, and data center shells — the chip shortage has eased in stages. NVIDIA and SoftBank's SB Energy are investing in a former uranium enrichment site in Ohio, starting at 4.25GW with room to reach 8GW.
On August 17, ByteDance Seed and Tsinghua AIR released CUDA Agent, an agentic RL system that trains a model to write GPU kernels that beat the compiler. On KernelBench it reaches a 98.8% pass rate and 2.11× geometric-mean speedup over torch.compile, with a 96.8% faster-than-compile rate. Level-3 results sit about 40 points above Claude Opus 4.5 and Gemini 3 Pro. Weights are not released; the 6,000-sample dataset, SKILL.md, and reward recipe are public.
The chip layer is no longer the whole story. Companies that made chips are now selling clouds, and the bottleneck has moved to power, land, and software that extracts more per watt.
On August 17, Cursor launched Origin, a code-hosting platform positioned as "Git hosting at agent scale," aimed at the gap between agent code-generation speed and existing infrastructure, directly challenging GitHub.
On August 18, WeCom opened CLI and MCP to enterprises of all sizes in version 5.0.10, letting WorkBuddy, DeepSeek Harness, and Minimax Code call ten office modules — documents, sheets, mail, meetings, calendar, contacts. Access is no longer gated by headcount or qualification. WeCom layers its existing permission and approval system onto the MCP path: AI permissions are configured separately from human permissions, critical actions require human approval, AI authorization supports expiry, and all calls are logged.
Claude Code shipped a /design skill in research preview, bringing the Claude Design canvas workflow into CLI and desktop, and cut p99 CPU usage by 2× through Bun GC scheduling. Alibaba Cloud's Qoder STAROps closes the loop from alert to fix in minutes using natural-language queries inside the IDE.
On August 18, Mureka V9.5, an open-source music model from Kunlun Wanwei, launched with a vocal quality rate of 61%, control rate of 97%, and 95.7% full-style fidelity, focused on Chinese guofeng articulation and harmony.
OpenAI President Greg Brockman warned on August 16 that OpenAI's model breaching Hugging Face in internal testing in July marks a "watershed moment" for cybersecurity, listing ten steps for defenders, including equipping security teams with AI agents and folding security review into the development pipeline.
In the same week, an open-source tool named watermarks-remover appeared on GitHub, stripping watermarks from Claude content and covering Gemini/SynthID and OpenAI provenance marks.
Platforms began cleaning AI slop. Spotify removed 75 million bulk-uploaded and duplicate tracks last year; Google cleared 130,000 low-quality channels and 50,000 account clusters on YouTube; TikTok auto-labeled over 3 billion AI-generated videos. Merriam-Webster named "slop" its 2025 word of the year.
Aligned with this, Anthropic's August 15 risk report disclosed an unreleased Model 2, slightly ahead of Claude Mythos 5 on several tasks with a modest overall gain.
Separately, Responsible Statecraft reported that Israel, through ad firm Piro, set up the "Hanover Institute for Public Policy," publishing over 100 reports since August 6 designed to sway LLM credibility assessments. GPTZero flagged most sampled reports as AI-written. Piro received $900,000 from the Israeli government.