On the deepseek v4.1 flash vs v4 pro benchmarks published so far, the new Flash model wins the only direct head-to-head available: Terminal-Bench 2.1, where V4.1 Flash posts 90.6 against the 87.9 that V4 Pro published in August 2026 — and it does so as the smallest model in DeepSeek's new architecture family, at reduced API prices. That result comes straight from DeepSeek's own API changelog for the 10 September 2026 release, and it explains the boldest line in the launch: DeepSeek says V4.1 Flash has surpassed V4 Pro on performance, cost and speed after internal and external testing. Vendor numbers deserve scrutiny, so this page lays out exactly what has been published for each model, what overlaps, what does not, and what you should run yourself before moving production workloads.
📺 Watch: NEW DeepSeek V4.1-Flash Is a GAME CHANGER!
🔥 Get the Agent OS as a free bonus: AI Profit Boardroom members get the full Agent OS zip, prompt libraries, daily tutorials and weekly live coaching calls. → Get inside · Want AI SEO help 1-on-1? Book a free SEO strategy session →
One note on sourcing before the numbers. Every figure on this page comes from DeepSeek's own published material: the API changelog entry for the V4.1 Flash release of 10 September 2026, and the launch announcement DeepSeek posted to its open platform on 9 September 2026. The V4 Pro figures are the ones DeepSeek published at that model's general-availability release in August 2026. Nothing here is estimated, and where the two models have not been scored on the same test, that gap is flagged rather than papered over. If you want the launch timeline itself — what was announced, when, and how the rollout was staged — the DeepSeek V4.1 Flash release date write-up covers that side of the story; this page is about the scores.
DeepSeek V4.1 Flash vs V4 Pro Benchmarks: The Published Numbers
Per the API changelog, V4.1 Flash ships with five headline benchmark scores. GPQA Diamond, the graduate-level science test, comes in at 90.9. Codeforces competitive programming lands at a 3471 rating. Terminal-Bench 2.1, the agentic terminal-work benchmark, scores 90.6. DeepSWE v1.1, the software-engineering suite, posts 74.2. And Humanity's Last Exam with tools enabled reaches 63.9.
V4 Pro's published card, from its August 2026 general-availability release, is built around different strengths: a mixture-of-experts design with 1.6 trillion total parameters and roughly 49 billion active per token, a 1-million-token context window, up to 384K output tokens in a single generation — and a Terminal-Bench 2.1 score of 87.9, which at the time sat a tenth of a point behind Claude Fable 5's 88.0 on DeepSeek's own published comparison.
Terminal-Bench 2.1 is therefore the one clean overlap between the two published sets, and on it the new small model beats the old flagship by 2.7 points: 90.6 versus 87.9. That is the concrete evidence behind DeepSeek's claim that Flash has surpassed Pro — not marketing hand-waving, but also not a full picture, because the remaining four V4.1 Flash scores have no published V4 Pro equivalent to compare against.
| Benchmark | DeepSeek V4.1 Flash (10 Sept 2026) | DeepSeek V4 Pro (Aug 2026) |
|---|---|---|
| Terminal-Bench 2.1 | 90.6 | 87.9 |
| GPQA Diamond | 90.9 | Not published on this suite |
| Codeforces (rating) | 3471 | Not published on this suite |
| DeepSWE v1.1 | 74.2 | Not published on this suite |
| HLE (with tools) | 63.9 | Not published on this suite |
All scores are DeepSeek's own published results, not independent replications. Third-party numbers usually appear within days of a DeepSeek release, and the sensible move is to treat this table as the company's opening claim rather than settled fact.
If you want to turn model releases like this into actual income — agents, automations and content systems built on whichever model wins the week — the AI Profit Boardroom is where the working playbooks live → Get the full model-to-money system. Prefer 1-on-1 help with your SEO first? Book a free SEO strategy session.
Why a Flash Model Is Beating the Flagship
The changelog is explicit that V4.1 Flash is not a distilled V4: it is "the smallest model in our new architecture family", built with native multimodal visual understanding and designed, in DeepSeek's words, for a higher capability ceiling, faster inference, higher throughput and scaling to larger models. In other words, the V4.1 Flash vs V4 Pro comparison is really a new architecture generation against the old one, with the new generation's entry model already clearing the old flagship on the shared agentic benchmark.
The multimodal piece matters more than it looks. The experimental V4 Flash Vision Exp model was DeepSeek's first public step into visual understanding in August; V4.1 Flash folds that capability in natively rather than as a separate experimental endpoint, and the changelog notes the old vision model's name now routes to V4.1. One model, text and vision, agent-grade terminal scores — that combination is what makes this release more than a price cut.
📺 Watch: How to use DeepSeek V4.1 Flash API for FREE!
Pricing, Model Names and the Billing Switch
On the API, the new model is simply called deepseek-flash — change the model name in your request and you are on V4.1. The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1 as well, so unmodified integrations move over automatically. Per the changelog, API prices have been reduced with the release, the earlier V4 models are retired, and V4 Pro continues to be served with unchanged billing.
The launch announcement added one more wrinkle worth knowing: from launch, requests sent to V4 Pro are answered by V4.1 Flash and billed at V4.1 Flash unit pricing until V4.1 Pro arrives. So if your stack points at the Pro endpoint today, you may already be getting Flash answers at Flash prices — check your logs before you assume your baseline is still the old flagship. DeepSeek's pricing moves have history here: the V4 pricing update and off-peak discounts were already among the most aggressive in the market before this release cut prices again.
Which Model Should You Use Now?
Ranked for practical use today, based on what has been published:
- DeepSeek V4.1 Flash for almost everything. It posts the better Terminal-Bench 2.1 score, it is cheaper per the changelog's reduced pricing, it handles vision natively, and it is the architecture DeepSeek says it will scale the rest of the family from. For agent work, coding runs and high-volume automation, it is the default.
- DeepSeek V4 Pro for long-context holdouts. Pro's published card still lists the 1M-token context window and 384K output generation — specs V4.1 Flash's changelog entry does not restate. If your workloads lean on enormous single-pass context or very long outputs, verify Flash matches those limits before you migrate. Billing for Pro itself is unchanged, so there is no cost penalty for waiting.
- Wait for V4.1 Pro if you need a new flagship. The announcement makes clear a Pro-tier successor is coming; the current release is the family's smallest member. If Flash already beats the old Pro on agentic benchmarks, the successor is the one to watch.
Whichever way you go, the model is only half the decision — the harness you run it in is the other half. The DeepSeek V4 Pro best harness guide walks through the options for the outgoing flagship, and the general DeepSeek harness guide covers the setup that applies to Flash as well.
📺 Watch: DeepSeek V4.1 Flash + Hermes Agent (FREE!)
How to Verify the Benchmarks Yourself
Published scores are a starting point, not a verdict — the gap between a launch chart and your actual workload is where migration mistakes happen. The practical approach: take three or four tasks you actually run — a coding ticket, a research brief, a long document extraction — and run them through deepseek-flash and your current model side by side before switching anything that matters. The Goldie Bench write-up covers how these model brains compare in hands-on tests and gives you a repeatable format for exactly this kind of side-by-side. And if you want the model wired into a full working setup rather than a chat window, the Agent OS is the free system built for dropping whichever model wins your tests into real workflows. A useful first project: the DeepSeek V4 tutorial translates directly to V4.1, since the API surface is unchanged apart from the model name.
FAQs on the V4.1 Flash vs V4 Pro Numbers
Does DeepSeek V4.1 Flash actually beat V4 Pro?
On the one benchmark both models have published — Terminal-Bench 2.1 — yes: 90.6 against 87.9, per DeepSeek's own figures. DeepSeek's broader claim that Flash surpassed Pro on performance, cost and speed comes from its internal and external testing and has not yet been independently replicated across a full shared suite.
What benchmarks did DeepSeek publish for V4.1 Flash?
Five headline scores in the 10 September 2026 changelog entry: GPQA Diamond 90.9, a Codeforces rating of 3471, Terminal-Bench 2.1 at 90.6, DeepSWE v1.1 at 74.2, and Humanity's Last Exam with tools at 63.9.
Is V4 Pro being retired?
Per the changelog, the previous V4 models are retired while V4 Pro continues with unchanged billing — but the launch announcement stated that V4 Pro requests are answered by V4.1 Flash at Flash unit pricing until V4.1 Pro arrives. Check your API logs to confirm which model is actually answering your Pro-endpoint calls.
Is DeepSeek V4.1 Flash multimodal?
Yes — the changelog describes native multimodal visual understanding, folding in what the experimental V4 Flash Vision model previewed. The old vision model's API name temporarily routes to V4.1 Flash.
Verdict: The Benchmark That Matters Says Flash
Strip away the launch noise and the published record is simple: on the single benchmark where DeepSeek has scored both models, V4.1 Flash beats V4 Pro by 2.7 points while costing less, and the four remaining Flash scores are strong enough that the burden of proof has flipped — it is now on V4 Pro to justify staying in your stack, not on the new model to earn a trial. Run your own side-by-sides before migrating anything critical, keep an eye out for the V4.1 Pro successor the announcement promises, and treat every number above as DeepSeek's claim until third-party replications land. But as opening claims go, this one is unusually easy to test: change one model name and see for yourself.
If you want these releases turned into working agents and income systems the week they drop — with the testing already done for you — check out the AI Profit Boardroom → Join the operators shipping with V4.1 Flash. And if you want a personal plan for your SEO and content, book a free SEO strategy session.











