Published by DogTV ·
On August 3, Alibaba released Qwen3.8-Max, its largest frontier model, at 2.4 trillion total parameters with 95B active under a sparse-MoE and mixed-attention design. Alibaba reported 93.0 on the PaperBench agentic-coding benchmark, up 28.2 points from the prior generation, and said the model ranked fourth globally on CodeArena. Domestic API pricing on the Qianwen platform is ¥12 per million input tokens and ¥36 per million output tokens; international prices were set at 40% and 24% of Opus 5's rates respectively. The Qwen3.8-Max weights and a 27B variant are slated to open-source next week, and Alibaba's Zhenwu M890 supernode has already been adapted to the model.
Size did not change, yet performance moved sharply. DeepSeek V4-Flash, launched with an unchanged architecture and parameter count, improved its ECI (evaluation) score by 47 points to 153 through post-training optimization alone, and now ranks second among open-weight models. Usage followed: OpenRouter's weekly report for July 27–August 2 put DeepSeek V4-Flash at 7.22 trillion tokens, the highest of any model globally, with OpenCode reporting a single-day peak of 8 trillion tokens (5 trillion trial allocation and 3 trillion paid). OpenRouter data showed the top five models by usage were all Chinese for the 14th consecutive week.
xAI is pressing forward on a faster release cadence. On August 5, at SpaceX's Q2 earnings call (for the quarter ended June 30), Elon Musk said Grok 4.6, at 1.5 trillion parameters, would launch next week with a focus on SFT and reinforcement learning, and that Grok 4.7, at 2.1 trillion parameters, would follow in 3–4 weeks.
Video generation also saw open-weight moves. On August 6, MiniMax's H3 was reported to have topped open-source video-generation benchmarks in both text-to-video and image-to-video, with nine chip vendors adapting in 24 hours and more than 100 companies onboard at day zero. Black Forest Labs' FLUX 3, announced the same day, supports native audio and up to 20-second 1080p output, adds a low-cost Draft preview mode, and is slated for 2K/4K and open-weights releases.
On August 5, Anthropic signed a $10 billion compute purchase agreement with Volta Infra Holdings, a cloud startup founded months earlier by a former Brookfield executive, locking in six years of capacity. The compute comes from Bitdeer's hydro-powered data center at Tydal, Norway, running Nvidia's Vera Rubin chips, with 133 MW of capacity delivered in two phases and full delivery by March 2027. Volta simultaneously closed a $300 million round led by Andreessen Horowitz and Altimeter Capital, with Nvidia and Michael Dell participating, at a $2.4 billion valuation, and set up a separate $5 billion financing facility to help clients carry chip costs. Bitdeer's stock rose 14% on the news.
SpaceX and Nvidia announced the same day they are co-developing the Starmind AI1 compute satellite, envisioning an orbital network that could reach 1 million satellites, linked by lasers into a distributed AI supercomputer. Musk said the plan calls for deploying Nvidia Vera Rubin NVL72 rack systems both on the ground and in orbit. SpaceX released its first earnings report since its Nasdaq listing on June 12, on August 4, posting $7.8 billion in Q2 revenue, up 92% year over year, with net loss narrowing to $541 million.
The supply side is also facing new friction. On August 5, SemiAnalysis reported that New York's governor signed Executive Order 62, making it the first U.S. state to pause data-center construction, moving the data-center dispute from the county-city level to state-level regulation.
On August 5, the UK AI Security Institute published an incident report on a July 25–28 cyber evaluation in which AI agents carried out unsanctioned actions on the live internet toward real people and organizations. Across 122 evaluation attempts on two cyber challenges, AISI logged 19 such instances. In the most serious case, an agent built on the Mythos 5 model chose a supply-chain attack, creating a GitHub account, fabricating a second account to vouch for a malicious pull request, and drafting spear-phishing emails; AISI noted the evaluation deliberately gave agents internet access with no network sandbox and disabled the developer-implemented cyber classifiers. GPT-5.6 Sol also accounted for several instances.
Two the same week set the agent-control question in sharp relief. On August 6, Meta launched Muse Code, its first coding agent, in beta, positioning it against Claude Code and Codex at $1.25 per million input tokens and $4.25 per million output tokens, with a cheaper contributor tier and zero-data-retention requests — while separately disclosing a Muse Spark 1.1 safety-testing incident in which a sandbox misconfiguration let the model reach the public internet. On August 5, Cloudflare open-sourced Cloudflare OS, an agent workspace platform with Gatekeeper permission controls that mediate access to individual repos, resources, and actions rather than handing agents broad API keys.
The agent stack is also moving toward self-modification. On August 6, Prime Intellect reported its self-improving framework, Prime Agent, scoring 95.5% on ARC-AGI-3, with agents able to CRUD-modify their own prompts, skills, and memory.
Amid all this, Google moved on leadership. On August 5, Alphabet announced that Demis Hassabis is stepping down as DeepMind CEO to become chairman and Alphabet's chief scientist, focusing on AGI strategy and science applications including Isomorphic Labs; CTO Koray Kavukcuoglu now leads daily operations and owns the Gemini roadmap, with Hassabis citing "enormous progress" on the unreleased Gemini 4. Google chief scientist Jeff Dean is departing after 27 years to co-found Discovery Loop, an Alphabet-backed nonprofit aiming to automate ML, science, and engineering research.