Gemini 3.7 Flash & DeepSeek V4 Pro: AI News Aug 14, 2026

Google kept its foot on the accelerator while its flagship stayed in the garage, China's open-weight wave crested again, and NVIDIA turned GPUs into a Wall Street asset class. Here is what mattered in AI over the past 24 hours, organized into five sections. Where a claim comes from media or aggregators rather than a first-party page, it is subject to vendor confirmation.
1. Top Story
Google launches Gemini 3.7 Flash — just three weeks after 3.6 Flash. On August 13, Google introduced Gemini 3.7 Flash, calling it its "most intelligent workhorse model yet for coding and agents." The release is notable less for any single benchmark than for the cadence: another Flash iteration barely three weeks after the last one, and still no sign of the long-delayed flagship Gemini 3.5 Pro.
The gains are real but incremental. On coding, Google reports DeepSWE v1.1 climbing to 65.3% (from 49.0%) and FrontierCode 1.1 Main to 43.6% (from 34.4%). Its WebDev Arena Elo rose to 1588 from 1538, and on knowledge-heavy document tasks the GDP.pdf benchmark improved to 34.0% (from 22.0%) with AutomationBench at 30.4% (from 17.0%). Google says the model "thinks more diligently" on multi-step planning and tool calls, and ships with updated CBRN and cyber-offense safeguards.
The sharper signal is on price. Through the end of 2026, 3.7 Flash carries an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens — half the launch price of 3.6 Flash (Google). It is live now in the Gemini API, AI Studio, Android Studio, Google Antigravity, and the Gemini Enterprise platform, and is powering the Gemini Spark agent for AI Pro/Ultra subscribers — though the mainstream chatbot still runs on 3.6 Flash for now. Meanwhile the flagship 3.5 Pro remains delayed; Google says it is training Gemini 4, and reporting suggests co-founder Sergey Brin has personally pushed staff to close the coding gap with OpenAI and Anthropic (Reuters, subject to vendor confirmation). Separately, CEO Sundar Pichai said the Gemini app now has more than 1 billion monthly users.
2. Model Releases
- DeepSeek V4 Pro goes official (V4-Pro-0813). DeepSeek updated its API docs, making the stable Pro build callable alongside the earlier V4 Flash. It keeps a 1M-token context and 384K max output, priced at ¥3 / ¥6 per million input/output tokens with cache-hit input at ¥0.025 (DeepSeek docs; Lanjing News, subject to vendor confirmation). Reports cite a jump on the DeepSWE agentic-coding benchmark from a preview 12.8 to 62.7, and Artificial Analysis's Intelligence Index v4.1.1 placing V4 Pro at 53 — behind GPT-5.6 Solo (61) and Claude Opus 5 (63), but at roughly $0.06 per task versus $2.34 for Opus 5.
- Alibaba open-sources Qwen3.8-Max weights (2.4T-A95B). August 13 marks the first time Alibaba has opened weights for a Max-tier flagship (Alibaba; media reports, subject to vendor confirmation). The team has publicized extreme agentic demos including ~16 days of autonomous coding and independent reproduction of research papers.
- NVIDIA ships Nemotron 3.5 Lightning + NeMo Switchyard. NVIDIA released a smaller, faster open Nemotron model aimed at the long-running agent execution layer (runnable on a single laptop GPU) plus NeMo Switchyard, an open-source Rust routing library already being adopted by Kong, OpenRouter, and LiteLLM. In internal tests, routing across open models plus Opus 4.8 held frontier-level accuracy at roughly one-third the cost of Opus alone. Available on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com.
- Meta releases Muse Glimmer. Meta open-sourced Muse Glimmer, a ~30B model tuned for on-device agent workflows — planning, tool calls, self-check, and failure recovery — and reiterated that flagship Muse Spark 1.2 weights will follow (media reports, subject to vendor confirmation).
- Google DeepMind brings sign language to consumers. A large multilingual sign-language-to-text (SL2T) model now supports American Sign Language input in Pixel 11's Gboard and Live Transcribe — the first time sign-language AI reaches a consumer product (IT Home, subject to vendor confirmation).
3. Industry & Capital
- NVIDIA and six Wall Street firms build a $500B GPU financing platform. NVIDIA signed an MOU with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to mobilize over $500 billion in third-party capital, turning GPUs into what Jensen Huang called "the first time technology chips have become an investable asset class." NVIDIA is reportedly in talks to provide up to $250B in standby financing for a 10GW OpenAI data center in Ohio, plus a separate ~$350B chip-procurement package. NVIDIA shares fell 2.86% to $217.55 on the news amid "circular financing" concerns voiced by short sellers Michael Burry and Jim Chanos (insideAI.news; Cailian Press, subject to vendor confirmation).
- Anthropic holds early pre-IPO investor meetings. CFO Krishna Rao is leading high-level talks; some investors float a ~$2 trillion valuation, though that is not the company's official target. Anthropic raised at a $965B valuation in late May (above OpenAI's $852B) and has disclosed an annualized revenue run-rate above $47B (Financial circles, subject to vendor confirmation).
- Intel launches a $20B stock offering to fund its AI push, part of a broader wave — U.S. firms have raised ~$105B via secondary offerings in 2026, roughly 40% AI-related (Goldman Sachs data).
- Synopsys validates the first HBM4 IP test chip at a stable 9.2 Gbps, a memory-bandwidth milestone for next-gen AI accelerators.
4. China Force
- Chinese models lead global token usage for a 15th straight week. Domestic models logged 34.25 trillion tokens last week (+21.76% week-over-week) out of roughly 69 trillion globally — more than the U.S. (National Business Daily; 21st Century Business Herald, subject to vendor confirmation).
- Open weights are "encircling" closed models. Per Hugging Face's Spring 2026 report, Chinese open models accounted for 41% of downloads — the first time surpassing the U.S. — with cumulative downloads topping 10 billion.
- Efficiency over raw scale. DeepSeek's V4 is optimized to run inference on Huawei's 950PR chip, while training used NVIDIA hardware; MoE designs like Kimi K3 (2.8T params, 16 of 896 experts active) and Tencent Hy3 (295B total, 21B active) keep compute costs down under U.S. chip export controls.
- China pushes AI global governance. State media reports the World AI Cooperation Organization has 38 founding member states, with 5,000 AI training slots pledged for developing countries over five years.
5. Observation
- Anthropic embeds invisible watermarks in Claude output. For all models released on or after August 2, Anthropic now adds machine-readable, human-invisible watermarks to text — and C2PA provenance metadata to other media — to comply with the EU AI Act (grace period through December 2026). The approach is aggressive: because watermarking happens at the model layer, even lightly edited or trivially modified text may be marked, and critics note the signal is easy for bad actors to strip yet may unfairly flag ordinary users (media reports, subject to vendor confirmation).
- Cadence as strategy. Google's steady drip of Flash models while 3.5 Pro slips — and while a reported talent exodus continues — has analysts asking whether the company is managing the appearance of momentum, or quietly pivoting to make Gemini 4 Pro its next flagship rather than shipping 3.5 Pro at all.
Compiled from WebSearch and first-party vendor pages (Google Blog, DeepSeek docs) cross-checked against Reuters, Bloomberg, Ars Technica, Axios, 9to5Google, and insideAI.news. Grok was unavailable this run. Media/aggregator items are subject to vendor confirmation.
Loading...