
On August 13, OpenAI and Cerebras launched Ultrafast Mode, a service tier powered by Cerebras' Wafer-Scale Engine, running GPT-5.6 Sol at up to 750 output tokens per second. Cerebras says this is 11x faster than Fable 5 and 5x faster than Opus 4.8 on Fast mode, with no quality degradation. On Humanity's Last Exam, a 2,500-question benchmark answerable only by PhD-level specialists, Ultrafast finished in 11 hours 11 minutes, while Claude Fable 5 needed 78 hours 27 minutes. On GDP-Val, a benchmark for economically valuable knowledge work, Ultrafast delivered 5.6x end-to-end speedup. The service is available first in the OpenAI API to a select group of customers.
On August 13, Google released Gemini 3.7 Flash, aimed at coding and agent use cases, just three weeks after 3.6 Flash. Through December 31, the model is priced at half of 3.6 Flash — $0.75 per million input tokens and $3.75 per million output tokens. Google is also folding the model into Gemini Spark for AI Pro and Ultra subscribers across 160-plus regions.
On August 14, DeepSeek announced peak-valley API pricing effective August 17, with off-peak hours at half the peak rate. Some models see price increases of more than 3x. The flagship deepseek-v4-pro sees a notable per-million-token price hike during peak hours. The move reverses DeepSeek's earlier posture as a low-cost provider.
On August 13, Lenovo reported Q1 FY2026 results: AI-related revenue grew 60% year over year to $8.5 billion, 35% of total revenue. Chairman Yang Yuanqing told media that compute allocation is shifting from early-stage "80% training" toward a target of "80% inference, 20% training," with the current split around 50/50.
The same day-group of announcements — a 750-token-per-second frontier model, a half-price Flash model, and a Chinese provider raising prices while others cut — show the competition has moved from "who is strongest" to "who is fastest and cheapest."
On August 14, DeepSeek released the official version of V4-Pro, a 1.6-trillion-parameter model with weights released under MIT license, native support for OpenAI Responses API, and targeted adaptation for Codex. The Harness plugin ecosystem entered public testing.
The same day, GLM-5.3 from Zhipu was reported with a 743B base model scoring 84.5 on CyberGym, above GPT-5.6 Sol, built for cyber defense and coding. Meituan launched LongCat-2.0, a 1.6T-parameter MoE model with a 1M context window, scoring 70.8 on Terminal-Bench 2.1, focused on agentic coding.
On August 14, Grok 4.6 topped the CursorBench real-world coding leaderboard, ahead of Claude Opus 5 and GPT-5.6 Sol. xAI also open-sourced the weights of Phoenix, its X-ranking algorithm, a Grok-based transformer whose synthetic-score weights favor propagation intent and genuine conversation, with likes treated as the cheapest signal.
On August 13, a Meta/Oxford study found multimodal pretraining needs only 5% image-generation data, with the optimal mix at 70% language, 25% image understanding, and 5% image generation — cutting image-generation tokens by 5x.
The open-weight push from DeepSeek, Meituan, and xAI, alongside specialized benchmarks for coding and cyber defense, marks a shift toward capability-specialized models competing on narrow, measurable tasks rather than general capability claims.
On August 13, Fortune reported Anthropic plans an October IPO at a $2 trillion valuation, which would surpass SpaceX's $1.77 trillion IPO on June 12 as the largest ever. Anthropic's annualized revenue passed $47 billion in May, and it filed confidentially with the SEC in June.
On August 14, Nikkei reported Tencent will acquire Meta's stake in AI developer Manus to become its largest shareholder. Manus, launched in March 2025 by Butterfly Effect, moved its headquarters to Singapore and cut Chinese operations after US investment. In April, China's foreign-investment security review banned the Meta-Manus acquisition; on August 11, Manus announced it was detaching from Meta and restoring independent operation.
On August 14, media reported Google DeepMind is restructuring away from frontier-model research toward cost-efficient Flash-class models, with layoffs potentially reaching one-third or more. Demis Hassabis stepped down as CEO on August 5, becoming chairman and Alphabet chief scientist. The restructuring is tied to delays of Google's next flagship Gemini model and internal disputes over development priorities.
On August 14, Reuters reported Apple has trained a China-market-specific LLM with Alibaba, ending its prior reliance on third-party models. Apple Intelligence is expected to launch in China within months, after the July 15 filing of the "Apple Intelligence" on-device generative AI service with regulators.
On August 13, Musk told SpaceX employees that the SpaceXAI (formerly xAI) data center power capacity would grow about 7x to 10GW by end of 2027, from the current 1.4GW across Memphis and Southaven. He estimated 10GW could yield $300 billion to $500 billion in annual revenue.
The same day brought two $2-trillion-scale signals and two China-market moves: Anthropic's IPO valuation, and Apple's China-model strategy — while Tencent absorbed Manus and Google retreated from the frontier.
On August 13, at the USENIX Security Symposium in Baltimore, researchers from the University of Birmingham and Durham University disclosed "Download More RAM," an attack exploiting a missing write-protection flaw in DDR4/DDR5 SPD chips to bypass Windows 11 defenses without physical access. The flaw lets software rewrite memory-configuration data, making Windows misreport RAM at 2x capacity and creating memory aliases to read or modify protected regions. It can disable antivirus, compromise locked enterprise systems, and bypass kernel anti-cheat. Microsoft assigned CVE-2026-23670 and added mitigation in the April 2026 security update; systems with Secure Boot enabled are protected. Corsair, G.Skill, and ADATA each had at least one affected consumer product line.
On August 6, a paper accepted to the ACM AI Leadership Summit found commercial AI detectors cannot reliably distinguish AI editing from full drafts. Light "refine abstract only" edits were flagged at 64–80% (Pangram/GPTZero), while unmodified 2023–2025 originals were flagged at 9–15%. After "Undetectable AI" humanization, fewer than 4% of AI-labeled rewrites remain flagged. Honest AI editing carries a higher sanction risk than humanizer-assisted evasion.
Also this week, a technical analysis argued text AI watermarks will always be trivial to remove, ahead of the EU AI Act Article 50 requirement that outputs be "detectable as artificially generated," enforceable from August 2026. SynthID-style token sampling can be stripped by paraphrasing through any weak LLM, and Unicode homoglyph watermarks can be removed by replacing characters. The analysis notes Claude Code previously used homoglyph steganography to tag certain requests.
On August 13, OpenAI launched Computer History for the ChatGPT Mac app, replacing the Chronicle screenshot-based preview with macOS accessibility event logging (clicks, typing, shortcuts, app switches). Events are stored locally for up to 48 hours, then deleted after server-side memory generation. Privacy-mode activity is excluded. The feature ships to Pro, Business, and Enterprise users, excluding the EEA, Switzerland, and the UK.
On August 14, ChatGPT gained the ability to directly edit Google Docs, Sheets, and Slides without switching tabs, rolling out to Plus, Pro, Business, and Enterprise users.
The hardware-level attack, the failing detectors, and the trivially-removable watermark share a common premise: the infrastructure for verifying AI content and securing systems is lagging the models themselves.