AI Mindset · Model Cheatsheets
DeepSeek

The High-Efficiency Technical Specialist

If your DeepSeek business case came off a screenshot, it’s wrong right now. You just don’t know which way yet. The price moved twice inside one month, in opposite directions. So a case built on the pre-August flat rates and a case built on the August peak rates are both wrong, and they are wrong opposite ways. Here’s what happened on September 10, 2026. DeepSeek shipped DeepSeek-V4.1-Flash as deepseek-flash, retired both V4-Flash and V4-Flash-Vision-Exp, and cut prices. The lineup is two models again: deepseek-flash and deepseek-v4-pro. Against DeepSeek's August rates the new volume model is cheaper on every billing line, so the August rise has now been partly given back. Image input is native to the default model. Weights are MIT licensed. And choosing DeepSeek still means explicit decisions about hosting, data jurisdiction and operational ownership.

Verified September 18, 2026V4.1-Flash · V4-Pro1M context · 384K outputPrices cut September 10, 2026MIT weights · dual API formats
1 / Meet DeepSeek

Strong capability per dollar, with a different operating model

Judge DeepSeek as infrastructure, not as a cheaper consumer chatbot. Its value rides on task routing, deployment choice and model evaluation. And on one human question: can your organization manage the legal and operational context that comes attached?

Personality

Technical, efficient and configurable

It’s built for teams that are comfortable choosing modes, APIs, weights and deployment architecture. Not for teams that want one polished end-user product.

  • Thinking and non-thinking modes
  • Agent and coding orientation
  • MIT-licensed weights and a hosted API
Deploy it for

High-volume technical work

Coding, tool use, document and image processing, long-context analysis, and any workload where token economics change the business case.

  • Route almost everything to `deepseek-flash`
  • Raise thinking effort before changing model
  • Exploit caching, concurrency and the off-peak clock
Choose DeepSeek when

Control and economics justify ownership

Use it when your team can evaluate quality, design safeguards and run the hosted or self-hosted path responsibly. And when processing in the People’s Republic of China under PRC law is acceptable for the data involved.

  • Benchmark on your own tasks
  • Resolve data and jurisdiction questions first
  • Budget for model operations

Is DeepSeek the right strategic choice?

Choose the requirement that is driving the decision.

API economics
Re-price against the September 10 rates, in both directions

`deepseek-flash` is the cheapest way into the lineup. It’s also got the highest listed concurrency at 2,500, and it is now the only model with image input. Those rates took effect at 04:00 UTC on September 10, 2026 and sit below DeepSeek’s August rates across the board. So anyone still costing against the pre-August flat rates is under-budgeting. Anyone costing against the August peak and off-peak rates is over-budgeting. Both are wrong, in opposite directions.
· Output: $0.60 per million tokens off-peak, $1.20 at peak.
· Cache-miss input: $0.15 off-peak, $0.30 at peak. Cache hit: $0.003 off-peak, $0.006 at peak.

  • Model quality on your data, not the headline rate
  • Cache-hit and off-peak scheduling assumptions
  • Peak is weekdays only, so weekends are off-peak all day
  • Total workflow cost, not token price alone
2 / What’s Current

A new default model, a price cut, and two retirements

Everything about the lineup changed on September 10, 2026. DeepSeek shipped DeepSeek-V4.1-Flash, retired two model IDs, and reduced prices with effect from 04:00 UTC that day. The two retired names are still accepted. But they are, in DeepSeek’s own wording, temporarily routed to V4.1-Flash and billed at the Flash price. So anything that pins them is running on a compatibility shim with no stated end date. Image input is now native to the default model rather than a separate experimental build. Pro continues with unchanged billing according to the live pricing page, although the release article says it should have stopped serving on September 14. That contradiction is unresolved. Both current models carry 1M context and 384K maximum output, and neither of the July-retired aliases appears on the pricing page.
· New: `deepseek-flash`. Retired: `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp`. Continuing: `deepseek-v4-pro`.

1M
Context length on both current models, with 384K maximum output
$0.60 / $1.20
`deepseek-flash` output per million tokens, off-peak and peak
$1.98 / $3.96
`deepseek-v4-pro` output per million tokens, off-peak and peak
04:00 UTC
The hour the current price list took effect, September 10, 2026

DeepSeek-V4.1-Flash New · September 10, 2026

The new default, and the only model in the lineup that accepts images. Everything below is a vendor architecture claim, not a measurement.
· DeepSeek describes a 552B-parameter mixture-of-experts model on a new causal encoder and decoder architecture. It splits a 40-layer Transformer into a 20-layer causal encoder and a 20-layer decoder, with just 8B active parameters for input and 16B for output.
· It claims the KV cache needs a quarter of the HBM and an eighth of the SSD storage of the previous generation, and that multimodality is native through DeepSeek-ViT, trained from scratch.
· The weights channel lists a 552B backbone and 763B total parameters under the MIT license.

  • `deepseek-flash`
  • 1M context, 384K maximum output
  • Image input supported, 2,500 listed concurrency
  • MIT-licensed weights

DeepSeek V4-Pro Continuing, with a caveat

The 0813 build reached general availability on August 13, 2026. Per the live pricing page it continues after September 14, 2026, with the billing method unchanged. What it no longer is, on DeepSeek’s own account, is the quality tier. The September 10 article says several parties measured V4.1-Flash ahead of it on quality, cost, speed and total runtime. Pro costs roughly three times Flash per token, carries 500 listed concurrency against Flash’s 2,500, and the pricing page marks image input as not supported for it. That missing vision is now its only distinguishing feature.

  • `deepseek-v4-pro`
  • No image input, 500 listed concurrency
  • Roughly three times the Flash token price
  • Read the contradiction card before you depend on it

The live V4-Pro contradiction Unresolved

Two first-party pages published the same day say opposite things. The live commercial page governs, so Pro is reported here as continuing and priced as continuing. But DeepSeek has not withdrawn, corrected or annotated the contradicting sentence in the release article. And the pricing page carries no last-updated stamp, so there is no way to establish when the reprieve was added. Read September 18, 2026.
· September 10 release article: starting at 04:00 UTC on September 14, 2026 all `deepseek-v4-pro` requests will route to V4.1-Flash at V4.1-Flash rates, and Pro is being phased out.
· Pricing page, saying the opposite: in response to user demand DeepSeek has decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. The change log agrees with the pricing page.

  • Pricing page and change log: Pro continues, billing unchanged
  • Release article, still live: Pro traffic routes away on September 14
  • Two same-day first-party pages say opposite things
  • A Pro dependency needs written confirmation, not a page reading

Retired: V4-Flash and Vision-Exp Executed

Two model IDs went out on September 10, 2026. DeepSeek states that the corresponding models have been retired, and that their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price. The change log says the two names are temporarily routed for compatibility. Temporarily is DeepSeek’s word, and it comes with no end date. The July retirement still stands separately.
· Retired September 10, 2026: `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp`.
· Dead since 2026/07/24 15:59 UTC: `deepseek-chat` and `deepseek-reasoner`. Neither appears on the pricing page.

  • Pin `deepseek-flash` or `deepseek-v4-pro`
  • A temporary route is not a supported model ID
  • Legacy `deepseek-chat` and `deepseek-reasoner` calls fail outright
  • Update tests, configuration and cost attribution

Vision is native now

The old routing decision has disappeared. You no longer send images to a separate experimental build. There is an upper bound of 1024 tokens per image after automatic resizing, and three ways to supply one: inline base64, an external http or https URL, or a file reference through the Files API. Whether the Files API is free could not be established on a September 18, 2026 read, because the Files API page would not render. So treat its cost as unconfirmed. Not zero.
· Pricing page: image input supported for `deepseek-flash`, not supported for `deepseek-v4-pro`.
· DeepSeek’s Vision guide, read September 18, 2026, for the token cap and the three input methods.

  • Upper bound of 1024 tokens per image
  • Base64, URL or Files API reference
  • Billed at the ordinary `deepseek-flash` token rates
  • Files API pricing could not be verified

Dual API formats

Two request formats, one service. Auto-mapping is convenient, and it also means a Claude model string sitting in your configuration quietly decides which DeepSeek model you pay for.
· OpenAI-compatible endpoint at `https://api.deepseek.com`. Anthropic-compatible endpoint at `https://api.deepseek.com/anthropic`.
· Anthropic guide, read September 18, 2026: model names starting with claude-haiku or claude-sonnet are mapped to `deepseek-flash`. claude-opus maps to `deepseek-v4-pro` and is billed at the V4 Pro price.

  • Lower migration friction
  • Explicit base URLs
  • claude-haiku and claude-sonnet land on `deepseek-flash`
  • claude-opus lands on Pro and bills at the Pro price

Thinking control

Thinking mode is enabled by default, and the default effort is `high`. So the expensive setting is the one you get if you say nothing. The pricing page lists no per-effort rate, so effort changes your token count rather than your rate.
· Thinking Mode guide, read September 18, 2026: four effort values, `none`, `low`, `high` and `max`.
· The parameter name differs by API format. The Anthropic-compatible format uses `reasoning.effort`. The Responses API uses `output_config.effort`.
· The aliases `minimal`, `medium`, `xhigh` and `ultra` are accepted and map onto the three active thinking levels.

  • `none` turns thinking off
  • Default is thinking enabled at `high` effort
  • Parameter name differs between API formats
  • Effort is not a separate line on the price list

Responses API coverage

Narrower than DeepSeek’s own release material implies. Standardized on the Responses API and also depend on Pro? That’s a gap to confirm with DeepSeek, not one to assume away. The August 13 release article separately claims native OpenAI Responses API support optimized for Codex with one-click setup. That’s a vendor claim to test.
· Responses API guide, read September 18, 2026: `deepseek-flash` is named as supported, and `deepseek-v4-pro` does not appear anywhere on that page.

  • Only `deepseek-flash` is named as supported
  • Pro is absent from that page, which is not the same as refused
  • Chat Completions and the Anthropic format remain the broad paths
  • Codex setup is a vendor claim, so test it

Peak and off-peak pricing Effective hour published

Peak is thirty-five hours a week, not forty-nine. So a batch job moved to a Saturday is off-peak for the full twenty-four hours. This regime is timed precisely, which the August 16 one was not: its effective hour is not stated on any current page. The pricing page itself carries no last-updated stamp, so a dated article is the only thing that can time a DeepSeek price change.
· Pricing page: "Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak)."
· September 10 release article: the new pricing takes effect at 04:00 UTC on September 10, 2026.

  • Peak is thirty-five hours a week
  • Weekends are off-peak end to end
  • The current regime has a published effective hour
  • The August 16 regime it replaced lasted twenty-five days

Open weights, MIT licensed Released

DeepSeek publishes its V4 weight collections and technical material on its own weights channel under the MIT license. That permits commercial use, modification and redistribution with attribution, and it is much more permissive than the bespoke community licenses common elsewhere. Here’s one oddity worth knowing. The retired `DeepSeek-V4-Flash-Vision-Exp` weights are still downloadable, with no retirement notice on the model card. So the weights channel is lagging the API.
· Weights channel, read September 18, 2026.

  • MIT across V4.1-Flash, V4-Pro-0813 and Vision-Exp
  • Hosting control, model operations responsibility
  • Hardware and quantization planning
  • A retired API model can still be self-hosted

Service status

DeepSeek publishes a status page with per-service uptime. Those are the vendor’s own figures, on a live page that is rewritten continuously. So they’re a snapshot, not a commitment. There is no published SLA attached to them.
· Read September 18, 2026: all systems green, no September incidents. Published uptimes of 99.90% for the V4 Pro API, 99.81% for the V4.1 Flash API, 99.87% for Chat, 100.00% for File Upload and 99.89% for Search.

  • Green across all services on September 18, 2026
  • Vendor-published uptime, not a contractual SLA
  • Subscribe to the status feed rather than reading it once
  • Design fallback for a provider with no SLA
Model and rate windowInput: cache hitInput: cache missOutputConcurrency
deepseek-flash (DeepSeek-V4.1-Flash) · off-peak$0.003 / 1M$0.15 / 1M$0.60 / 1M2,500
deepseek-flash (DeepSeek-V4.1-Flash) · peak$0.006 / 1M$0.30 / 1M$1.20 / 1M2,500
deepseek-v4-pro (DeepSeek-V4-Pro-0813) · off-peak$0.022 / 1M$0.66 / 1M$1.98 / 1M500
deepseek-v4-pro (DeepSeek-V4-Pro-0813) · peak$0.044 / 1M$1.32 / 1M$3.96 / 1M500
deepseek-v4-flash · retired, legacy aliasBilled at the Flash rate aboveBilled at the Flash rate aboveBilled at the Flash rate aboveServed by V4.1-Flash
deepseek-v4-flash-vision-exp · retired, legacy aliasBilled at the Flash rate aboveBilled at the Flash rate aboveBilled at the Flash rate aboveServed by V4.1-Flash
Model IDBuild namedContextMax outputImage inputThinkingConcurrency
deepseek-flashDeepSeek-V4.1-Flash1M384KSupportedSupported2,500
deepseek-v4-proDeepSeek-V4-Pro-08131M384KNot supportedSupported500
3 / Model Router

Use the least expensive mode that meets the task’s failure cost

The routing map got simpler on September 10. Images no longer route anywhere special, because the default model has vision. Hard work no longer routes to Pro, because the vendor no longer claims Pro is better. What is left is one model, four effort levels and a clock. Run `deepseek-flash` at the right effort, scheduled against the weekday peak windows. Keep Pro for the narrow cases that really call for it.

Route the workload

Select the dominant task property.

Everyday agent task
`deepseek-flash` · default `high` effort

The vendor default, and the right economic default too. Flash with thinking at `high` covers planning, tool calling and coding where the model needs a deliberate plan. It runs at roughly a third of the Pro token price, with five times the listed concurrency.

  • Strong economic default
  • The setting you get if you say nothing
  • Set maximum steps
  • Image input works here too, capped at 1024 tokens per image

Reasoning effort versus operational cost

DeepSeek documents four effort values: `none`, `low`, `high` and `max`. Thinking is enabled by default at `high`. The Anthropic-compatible format sets it through `reasoning.effort`, the Responses API through `output_config.effort`. The aliases `minimal`, `medium`, `xhigh` and `ultra` map onto the three active thinking levels. Raise effort only when the task actually benefits from deeper search or tool use.

ThroughputMaximum reasoning
Effort `high` (default)
Default complex work

The documented default. DeepSeek positions high effort for daily agent workflows: analysis, tool calling and coding where the model needs a deliberate plan. This is what you are billed for if you never set the parameter.

  • Vendor default, not an upgrade
  • Good balance for agents
  • Set completion criteria
  • Check whether your workload actually needs it
Routing evaluation
Run this task suite on deepseek-flash at effort none, low, high and max, and on deepseek-v4-pro at high. Compare exact-task success, latency, output tokens, retries and human correction time. Recommend a routing policy and state the cost per successful task at both peak and off-peak rates.
Agent boundary
Complete this repository task with a maximum of 12 tool steps. Stop before any external network or deployment action. Report tests, files changed and unresolved risk.
4 / Deployment

Hosted API and open weights solve different problems

Open weights improve control. They do not remove cost. They swap per-token pricing for infrastructure, model serving, security, monitoring and lifecycle work. What DeepSeek does remove is the licensing obstacle. The published V4 weights are MIT.

Choose the deployment path

Start with the organizational requirement, not the ideology.

Fastest path to production
Hosted DeepSeek API

Use the official API when speed of integration, provider-operated serving and listed token economics are the priority. Accept that content is processed and stored in the People’s Republic of China. And accept that the same Open Platform terms cover individual and enterprise developers.

  • OpenAI or Anthropic request format
  • Provider data and jurisdiction review
  • Monitor rate and pricing changes, which moved twice in a month
DimensionHosted APISelf-managed weights
Time to startFastInfrastructure project
Variable costPublished token pricingCompute, staffing and utilization
Data controlProcessed and stored in the PRC under PRC lawOrganization controls environment
LicensingOpen Platform terms, no enterprise variantMIT on the published V4 weights
ScalingProvider-managed within listed concurrencyOrganization provisions capacity
Model updatesProvider-led, and a model can retire on the day it is announcedOrganization tests and deploys
ObservabilityAPI-level signalsFull stack if implemented
Operational burdenLowerHigher

Context caching

Caching on disk is enabled by default for all users, with no code change. And the discount is not a published percentage. It’s the cache-hit column of the price list. A cache hit costs 2 percent of the cache-miss rate on `deepseek-flash` and 3.3 percent on `deepseek-v4-pro`. So prompt architecture is part of the economics, not a micro-optimization.

  • Keep stable instructions stable
  • Measure the real hit rate through `prompt_cache_hit_tokens`
  • Do not assume every request qualifies

MIT-licensed weights

Permissive licensing removes one common blocker to a self-hosted path. And the Vision-Exp repo shows a second use. Those weights are still downloadable, with no retirement notice on the card, so a model DeepSeek has pulled from the API is still available for you to run.
· MIT on V4.1-Flash, V4-Pro-0813 and the retired V4-Flash-Vision-Exp, on DeepSeek’s weights channel, read September 18, 2026.

  • Commercial use, modification and redistribution with attribution
  • A retirement escape hatch for a model you depend on
  • The weights channel lags the API, so verify against both

User isolation

The API supports a user identifier for content-safety, KV-cache and scheduling isolation.

  • Do not place personal data in the identifier
  • Use stable internal pseudonyms
  • Understand account-level limits

Concurrency planning

Published account-level limits differ sharply: 2,500 for `deepseek-flash` and 500 for `deepseek-v4-pro`. Flash is now cheaper, and on DeepSeek’s own account it measures better too. So the concurrency gap is one more reason the migration runs toward Flash rather than away from it.

  • Model peak demand
  • Handle 429 responses
  • Request capacity expansion when justified

Off-peak scheduling

The clock is a cost lever. Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, with all other hours off-peak. So batch, backfill and evaluation runs that do not need to be interactive belong in the other seventeen hours of a weekday. Or anywhere at all on a Saturday or Sunday.

  • Queue deferrable work for off-peak
  • A weekend run is off-peak for the full day
  • Track the UTC windows, not local time
  • Interactive weekday traffic will still land at peak

Service status and fallback

DeepSeek publishes per-service uptime on a live status page. Those are vendor figures on a page that is rewritten continuously, and DeepSeek publishes no SLA alongside them.
· Read September 18, 2026: all systems green, no September incidents, uptimes between 99.81% and 100.00% across the V4 Pro API, the V4.1 Flash API, Chat, File Upload and Search.

  • Subscribe to the feed rather than checking manually
  • Uptime disclosure is not an SLA
  • Design an approved fallback provider anyway
5 / Agentic Work

Thinking inside tool use is useful, and needs hard boundaries

Cheap tokens make a bad loop cheap to run, right? That’s the whole risk here. DeepSeek’s agent direction began with V3.2 and continues with V4.1-Flash, which is positioned as an agent model and ships with thinking enabled at `high` effort by default. The operational risk has not changed, and it is slightly worse now that tokens are cheaper again. An efficient model can take a lot of inexpensive wrong steps before anyone notices.

The bounded DeepSeek agent loop

Make the stop conditions part of the task.

Brief
Define outcome and maximum authority

State the finish line, the available tools, the prohibited actions, the step limit and the validation method.

  • Maximum turns or budget
  • No implicit external side effects
  • Named fallback
Tool-using analysis
Research this technical decision using only the supplied documentation and approved web domains. Maximum 10 tool calls. Cite every conclusion and stop if the sources conflict.
Agentic coding
Implement the change, run the tests and inspect the diff. Maximum 15 tool steps. Do not modify deployment, credentials or network configuration. If blocked, report the exact evidence.
Structured verification
Return the result as the required JSON schema, then run an independent validation pass that checks completeness, allowed values and source coverage.
Commit rule
After two verification passes, commit to the strongest supported answer. Do not continue rechecking unless new evidence appears.
6 / Behavior Playbook

Nine habits for extracting the economic advantage safely

DeepSeek rewards engineering discipline. Explicit routing. Structured outputs. Hard agent limits. Scheduling against the weekday peak windows. And evaluation on your organization’s real tasks, not on a price list that changed twice in one month.

01

Pin a current model ID

Send `deepseek-flash` or `deepseek-v4-pro`. `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` were retired on September 10, 2026 and are only temporarily routed to V4.1-Flash. `deepseek-chat` and `deepseek-reasoner` died at 2026/07/24 15:59 UTC and now fail outright.

02

Route by effort, not by model

Flash covers the range. Escalate `none`, `low`, `high`, `max` before you consider a different model.

03

Turn thinking off deliberately

Thinking is on by default at `high` effort. Do not pay reasoning latency for transformations a schema can validate.

04

Exploit stable prefixes

A cache hit costs 2 percent of the cache-miss rate on Flash. Design repeated context to benefit from caching without hiding changing instructions.

05

Set hard agent limits

Maximum steps, budget, tools and stop conditions belong in the task contract.

06

Evaluate cultural and political domains

Test behavior on topics that matter to your users and markets, not only on code benchmarks.

07

Separate hosted and self-managed risk

The same model can create different legal and security profiles by deployment. Hosted means PRC processing. MIT weights mean self-hosting is a real option for you.

08

Measure correction cost

Include retries and human review when you compare providers.

09

Schedule deferrable work off-peak

Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, and all other hours are off-peak. So move batch and evaluation runs out of those windows, or onto a weekend.

7 / Watch Outs

The risk profile belongs in the architecture decision

Capability answers none of the questions that actually stall a deal. Jurisdiction. Hosted-service data handling. Political behavior. Licensing. Operational responsibility. DeepSeek states the jurisdiction facts plainly in its own policies, so you can decide on them rather than guess at them.

Data residency and governing law

Stated plainly in DeepSeek’s own policies. So decide whether it is acceptable before the technical evaluation, not after.
· Privacy Policy, last updated February 10, 2026: DeepSeek directly collects, processes and stores your personal data in the People’s Republic of China.
· Open Platform Terms of Service, effective April 29, 2026: the agreement is governed by the laws of the People’s Republic of China in the mainland. Suits are filed at the location of the registered office of Hangzhou DeepSeek Artificial Intelligence Co., Ltd.

Retention is open-ended

The Privacy Policy commits to keeping personal data for as long as necessary to provide the services. No fixed deletion window is published. It also gives users the right to opt out of the use of their personal data for training DeepSeek’s models or optimizing its technologies. That is a control you have to exercise, not one that applies by default.
· The API-specific Open Platform privacy policy could not be retrieved on a September 18, 2026 read, so these facts come from the general policy and the Open Platform terms.

No separate enterprise terms

The same Open Platform Terms of Service cover individual and enterprise developers. There is no enterprise agreement, no separate data-processing addendum on any first-party page, and no published SLA behind the status page uptimes. So if your procurement process assumes a negotiated contract exists to fall back on, check that assumption early.

Political sensitivity

Evaluate behavior on politically sensitive and culturally specific topics that matter in your markets. This is an evaluation task on your own prompts. No benchmark answers it for you.

Provider concentration

Do not build a critical system without model abstraction, fallback and exit planning. A provider that retired two model IDs on the day it announced their replacement is a provider whose lineup you should be able to leave quickly.

Open-weight operations

Self-hosting transfers patching, serving, scaling, safety and monitoring to the organization. The MIT license makes that legally simple. Legally simple isn’t the same as operationally cheap.

Long-context confidence

A 1M window does not guarantee full coverage, reliable retrieval or balanced attention. DeepSeek’s efficiency claims for the new KV cache design are vendor claims about resource use, not about retrieval quality.

Agent loops

Cheap tokens can make a runaway verification or tool loop inexpensive and still operationally harmful. And tokens just got cheaper again.

A live contradiction about V4-Pro

Two DeepSeek pages published the same day say opposite things. The live commercial page governs, so Pro is reported here as continuing. But the contradicting sentence has not been withdrawn or annotated, and the pricing page carries no last-updated stamp, so the reprieve cannot be dated. Anyone with a production dependency on Pro should get written confirmation from DeepSeek rather than rely on either page.
· September 10 release article: all `deepseek-v4-pro` requests route to V4.1-Flash from 04:00 UTC on September 14, 2026, and Pro is being phased out.
· Pricing page and change log: Pro continues after that date with unchanged billing.

The price moved twice in a month

This is the one to say out loud to a client, right? The peak and off-peak policy went live on August 16, 2026 and raised every billing item. Twenty-five days later, at 04:00 UTC on September 10, the whole list was replaced again and the volume rates came down. Output about 9 percent lower, cache-miss input about 32 percent lower and cache-hit input about 57 percent lower than the August V4-Flash rates. Two full repricings in under a month, one up and one down. That’s the pattern to plan for. A vendor that can do that twice can do it again, in either direction. So budget with headroom, keep an abstraction layer, and re-read the pricing page before any commitment that rests on a published number.

Weights can move under a fixed model ID

Credit where it is due: DeepSeek names builds where buyers can see them. The pricing page identifies DeepSeek-V4.1-Flash and DeepSeek-V4-Pro-0813, matching the weights channel. The caution survives the improvement, though. A model ID is still an alias over weights that can be replaced. Naming a build on a live page is not a versioned endpoint you can pin to. And `deepseek-flash` is a generational name rather than a dated one, so the next replacement may not change the string you send at all. So re-run your evaluation suite on a schedule, not only on a version bump.

Retired in the API, still on the weights channel

The `DeepSeek-V4-Flash-Vision-Exp` weights remain downloadable with no retirement notice on the model card, days after the API retired the model. Read that two ways. It’s a real escape hatch, because an MIT-licensed model you can host does not disappear when the vendor withdraws the endpoint. It is also evidence that DeepSeek’s channels do not update together. So a live weights page is not proof that a model is still served.

No quotable numbers for the new model

The September 10 release article publishes its benchmark comparison and its price chart as images rather than text. So there are no citable V4.1-Flash figures. The only quotable DeepSeek benchmark scores from this generation belong to the now-retired V4-Flash-Vision-Exp, and a score for a retired model says nothing about the model that replaced it. So treat the vendor’s claim that V4.1-Flash outperforms V4-Pro as a claim, and measure it on your own tasks.

Risk decisionHosted API questionSelf-managed question
DataContent is processed and stored in the PRC. Is that acceptable for this data class?Who can access model inputs, logs and infrastructure?
SecurityWhat provider and network controls apply?How are serving stack and weights protected?
SafetyWhat provider policies and isolation apply?What filters, evaluation and abuse controls will you operate?
ReliabilityListed concurrency and a status page, but no published SLA. What is the fallback?How will you scale and recover?
LegalPRC governing law, Hangzhou jurisdiction, one set of terms for everyone. Who signs off?MIT license, plus your own acceptable-use and export review
LifecycleTwo models retired on the day of announcement. How fast can you migrate?Who owns model testing and rollout?
8 / Sources

First-party evidence behind this guide

Every source listed below is a dated first-party DeepSeek article or policy document. All were re-checked on September 18, 2026. Several load-bearing facts here, though, live only on pages DeepSeek rewrites in place. Those pages can’t anchor a dated claim, so they are attributed inline rather than listed below. Confirm every one of them directly against the live page before you rely on it, and keep a timestamped copy of anything you will need to cite later.
· Models and Pricing page, read September 18, 2026, and carrying no last-updated stamp. It is the source for the full price table, the peak and off-peak window text, the per-model image-input and concurrency limits, the legacy-name routing notice and the V4-Pro reprieve notice.
· Vision guide: the 1024-token image cap and the three image input methods. Thinking Mode guide: the four effort values and their per-format parameter names. Anthropic API guide: the Claude model-name mapping. Responses API guide: the `deepseek-flash`-only support statement. Rate Limit and Isolation page: the concurrency figures. All read the same day.
· Change log: the temporary routing of the two retired names, and the price reduction wording.
· DeepSeek’s weights channel on Hugging Face: the MIT license and the parameter counts. DeepSeek status page: the uptime figures.

DeepSeek-V4.1-Flash: Smarter, Faster, More EfficientDeepSeek · September 10, 2026 · the current release article · V4.1-Flash architecture and parameter claims, the retirement of V4-Flash and V4-Flash-Vision-Exp, the 04:00 UTC September 10 pricing effective hour, and the still-live sentence routing V4-Pro away on September 14 that the pricing page contradictsDeepSeek-V4-Pro GA ReleaseDeepSeek · August 13, 2026 · V4-Pro general availability and the unchanged calling method · history: its pricing announcement was superseded on September 10, 2026DeepSeek-V4-Flash-Vision-Exp Release: Multimodal API Now LiveDeepSeek · August 21, 2026 · history only · the model it announces was retired on September 10, 2026, so its benchmark figures do not describe anything DeepSeek still servesDeepSeek V4 Preview ReleaseDeepSeek · April 24, 2026 · history only · preview parameter counts and the alias retirement it still describes in the future tenseDeepSeek Open Platform Terms of ServiceDeepSeek · effective April 29, 2026 · PRC mainland governing law, suits at the registered office of Hangzhou DeepSeek Artificial Intelligence Co., Ltd., and one set of terms covering individual and enterprise developersDeepSeek Privacy PolicyDeepSeek · last updated February 10, 2026 · personal data collected, processed and stored in the People’s Republic of China, open-ended retention, and the opt-out from model training · the API-specific Open Platform privacy policy could not be retrieved on this read

AI Mindset

Explore the model cheatsheets