Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
DeepSeek

DeepSeek-V4-Flash Goes Official: Agent Benchmarks Surpass V4-Pro-Preview

AI By Crimson AI DeepSeek Blog 2 August 2026 · 08:32 33 views
Share: X Telegram

DeepSeek has launched the official deepseek-v4-flash API in public beta, featuring improved agent capabilities that outperform V4-Pro-Preview across nine benchmarks, with no code changes required.

DeepSeek-V4-Flash Goes Official: Agent Benchmarks Surpass V4-Pro-Preview

Key points

DeepSeek has officially released the deepseek-v4-flash API into public beta as of July 31, 2026. The update retains the same model string, endpoint, and architecture, but introduces a refined post-training pass that significantly boosts agent performance metrics, surpassing the V4-Pro-Preview on nine agent benchmarks.

The most notable change is the model's enhanced agent capabilities. According to DeepSeek's reported figures, the official V4-Flash substantially exceeds V4-Pro-Preview across all nine evaluated benchmarks, including Terminal Bench 2.1 (82.7), NL2Repo (54.2), and Cybergym (76.7). These scores were measured using DeepSeek Harness minimal mode at max tier with top_p=0.95 and temperature=1.0, and should be treated as vendor-reported until independently verified.

DeepSeek emphasizes that V4-Flash-0731 shares the exact structure and size of its preview version; only the post-training was redone. This means latency and cost profiles remain stable, but prompts may behave differently due to changes in tool-calling style and refusal behavior. Self-hosters are unaffected as this is an API-side upgrade.

The official Flash model natively supports the Responses API format and is specifically adapted for Codex, allowing developers to integrate it without translation shims. This positions Flash as the default model for coding agents, not just a budget fallback.

No pricing changes accompanied the release. V4-Flash remains at $0.14 per 1M input tokens (cache miss), $0.0028 (cache hit), and $0.28 per 1M output tokens. A peak-hour surcharge has been announced but is not yet active. Only the Flash API was upgraded; V4-Pro and app models remain unchanged, with an official V4-Pro release promised soon.

BenchmarkV4-Flash (Official)
Terminal Bench 2.182.7
NL2Repo54.2
Cybergym76.7
DeepSWE54.4
Toolathlon verified70.3
Agent Last Exam25.2
Automation Bench (Public)25.1
DSBench-FullStack68.7
DSBench-Hard59.6
Source
DeepSeek · DeepSeek Blog
Related news
DeepSeek
DeepSeek 10 Aug 2026

DeepSeek V4 Launches with 1.6T Parameters, Claims 10-50x Cheaper Pricing

DeepSeek unveiled its V4 model family on April 24, 2026, featuring two text-only variants with up to 1.6 trillion parameters and a...

19
DeepSeek
DeepSeek 4 Aug 2026

DeepSeek Retires Legacy API Aliases: Price Changes and Migration Guide

DeepSeek retired the legacy API aliases 'deepseek-chat' and 'deepseek-reasoner' on July 24, 2026, at 15:59 UTC. Developers must up...

33
DeepSeek
DeepSeek 27 Jul 2026

DeepSeek vs Stepfun: Two Chinese AI Labs, Two Divergent Strategies for 2026

DeepSeek focuses on reasoning and low-cost APIs, while Stepfun bets on multimodal generation including video and audio. Here's how...

28