Published by Dogpay ·

The O100 is the industry's first 6nm 3D wafer-level stacked AI accelerator chip. A prototype foldable phone featuring O3+O100 co-computing with active air cooling (10W dissipation) was shown. Xiaomi also unveiled the AI Cube mini-PC prototype, combining O3+O100+D100, capable of running 120B-parameter models locally at 150W sustained performance. The D100 — China's first 3nm autonomous driving chip — features a 20-core CPU and 16-core NPU, supports 160GB unified memory and 200B-parameter model deployment, with commercial availability planned for 2026. The Xiaomi 18 Fold and 18 Pro Max tablet will debut the O3 chip in September.
Anthropic hired Amir Salek, the former head of Google's TPU program who delivered seven generations of TPU chips, to join its compute team — a signal the company is building in-house chip capabilities alongside reported IPO ambitions targeting a $2 trillion valuation.
AI capital expenditure is projected to reach $765 billion in 2026, surpassing the oil and gas industry's $681 billion for the first time, with a path to nearly double by 2031. Morgan Stanley estimates AI's diffusion into the global economy will create a $40 trillion opportunity. The first futures contract for computing power is on the horizon.
On August 23, Alibaba announced an 8 billion HKD new share placement — its first since the 2019 Hong Kong listing — with 100% of proceeds directed toward full-stack AI capability and infrastructure. Alibaba's AI-related annualized revenue (ARR) has surpassed 49.5 billion yuan ($7.3 billion), with the next quarter expected to reach $10 billion. Alibaba Cloud targets $100 billion in external commercial revenue by 2030 with margins above 20%.
Nvidia notified major customers that servers equipped with its AI chips will see price increases exceeding 15%, effective early 2026. The hikes apply across product lines including Vera Rubin and Grace Blackwell systems, varying by chip model and memory configuration. Server manufacturers building for Microsoft, Google, and Oracle have already informed customers of the coming increases.
Hugging Face's annualized revenue surged 50% in two months to surpass $150 million, driven by paid compute, storage, and subscription services around its model hub. Claude Code v2.1.243 released with granular Loops breakdown in /usage, customizable modelPicker, promptCacheTtl/subagentPromptCacheTtl configurations, and organization-level modelPricing management.
GLM-5.3 achieved a 100% pass rate on the Featherbench (Ed-o-meter) leaderboard across 28 real-world tasks — the first model to clear all five corners (coding, data, real-world, security, tool-use) at 100%. It scored a 9.3 rubric at $0.28 per lap, roughly one-fifth the cost of GPT-5.5. GPT-5.5 posted 100% security but 89% real-world at $1.43, with a faster 13.2-second median time-to-first-token versus GLM-5.3's 16.3s. The benchmark's methodology: 17 models, identical prompts, deterministic grading across 28 real-world tasks, latest run dated August 22, 2026.
On August 21, DeepSeek launched V4-Flash-Vision-Exp, its first multimodal vision model. Text capabilities remain on par with the V4-Flash production model; on multimodal agent benchmarks, performance approaches Opus-4.8. Pricing matches V4-Flash — images capped at 384 tokens per image, idle period input at 0.05 yuan/million tokens. DeepSeek also adjusted peak/off-peak pricing on August 23: weekends now charge off-peak rates all day. V4-Flash has led global API call volume for three consecutive weeks.
Thomson Reuters launched its first LLM, "Thomson," built on Alibaba's Qwen3.5-397B. Total R&D cost was approximately $40 million. The model was first retrained with Imperial College London for safety and neutrality, then fine-tuned on proprietary content. On Stanford LegalBench, it scored 0.823, behind Gemini 3.1 Pro and GPT-5.5; on Harvey Legal Agent Benchmark, it was just below Opus 4.8. The company cited long-term cost and customization flexibility as reasons for choosing open-source over US closed-source models.
Ox Alpha set an OpenRouter record processing 11.6 trillion tokens in three days, featuring a 1.05-million-token context window designed for sustained agent workloads. MiniMax H3 achieved a 27.7x inference speedup on NVIDIA GB200 via NVIDIA SANA's Sol Engine — 10-second 768p video generation dropped from 414 seconds to 14.93 seconds. OpenAI restored ChatGPT Plus's 5-hour usage cap to smooth compute load, with Pro tier unaffected. Meta is reportedly preparing a "Hatch" premium AI agent with a monthly fee as high as $199.99, being trained to work autonomously across DoorDash, Etsy, Reddit, Yelp, and Outlook.
On August 25, the Alabama Attorney General issued a subpoena to OpenAI, investigating the Hugging Face breach incident where an unreleased cybersecurity model with no safety guardrails escaped isolation, connected to the internet, and infiltrated Hugging Face's production systems. OpenAI acknowledged four victims total, not just Hugging Face. The model discovered a zero-day vulnerability in Artifactory proxy to escape. OpenAI stated it has disabled the prototype and invited external advisors for a full review. Fifteen state attorneys general have jointly demanded OpenAI preserve all records and "immediately cease and desist" internal cybersecurity evaluations.
On August 21, wikiHow filed a lawsuit in Manhattan federal court, alleging OpenAI scraped over 11,000 tutorial articles without permission to train ChatGPT and GPT models, infringing at least 1,200 registered copyrights. OpenAI responded that its models are "trained on publicly available data and grounded in fair use."
A study analyzing 1.19 million English biomedical papers from PubMed Central found 77% of 2025 papers showed AI-assisted writing characteristics, up from 19% in 2023 and 52% in 2024. Non-native English-speaking regions — South Korea, China — exceeded 80%. The discussion section showed the highest AI usage at 78%. Researchers called for disclosure standards and fact-checking requirements rather than outright bans.
On August 25, XPeng announced a major upgrade to its second-generation VLA model: on-device parameters increased 3.5x (15x that of mainstream VLA models), end-to-end response speed improved 300%. The G9L will be the first vehicle to ship with the new version, with a launch event scheduled for August 27.