The short answer

DeepSeek V4 is the cost floor for SEC filing work in 2026. Its Flash tier (now V4.1-Flash) lists at $0.30/$1.20 per Mtok at peak with a 1M-token window and 384K max output, about $0.041 per 120k-token filing and roughly $408 across 10,000, $204 off-peak, falling toward $55 when the repeated schema prefix hits the cache. Pin the current model name, deepseek-flash.

DeepSeek V4 is the cost floor for SEC filing work in 2026. Its Flash tier, now served by DeepSeek-V4.1-Flash, lists at $0.30 per million input tokens and $1.20 per million output at peak, half that off-peak, with a 1M-token context window and 384K max output that holds a full 10-K without chunking. At a 120k-in / 4k-out filing shape that is about $0.041 per filing at peak, roughly $408 across 10,000 filings and $204 off-peak, dropping toward $55 when the repeated schema prefix hits the cache. Watch the model names: the legacy deepseek-chat and deepseek-reasoner aliases retired on 2026/07/24, and the Flash tier's current API name is deepseek-flash. Price your own filing volume in the Token Cost Optimizer.

TL;DR

Model Input $/Mtok Output $/Mtok Context Max output ~$/filing (120k in + 4k out)
DeepSeek V4.1-Flash $0.30 $1.20 1M 384K $0.041
DeepSeek V4.1-Flash (cache-hit prefix) $0.006 $1.20 1M 384K $0.006
DeepSeek V4-Pro $1.32 $3.96 1M 384K $0.174

Per-filing costs are arithmetic from the verified list prices (120,000 input tokens × input rate + 4,000 output tokens × output rate), not a benchmark run. Prices are peak rates, verified 2026-09-28 against DeepSeek's official pricing page; off-peak rates are half.1

Why V4 resets the filings price floor

SEC extraction is a long-input, high-volume, structured-output job: feed a 10-K or 10-Q, pull specific line items, repeat across thousands of filings. The bill is dominated by input tokens on long documents and output tokens on structured results, so per-token price and context fit decide it. V4.1-Flash pairs a 1M-token window with $0.30 / $1.20 peak rates, so a full filing fits in one call and the per-document cost lands near four cents at peak and two cents off-peak.1 Earlier coverage on this site quoted V4-Pro under a 75% promotion; that promotion is gone, and the steadier number to budget against is the Flash list rate.

The verified V4 prices

DeepSeek's pricing page lists two models, DeepSeek-V4.1-Flash and DeepSeek-V4-Pro-0813, both at a 1M context window and 384K max output. Off-peak rates are half of peak; peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday.1

V4.1-Flash (model name deepseek-flash) $0.30 / Mtok cache-miss input, $0.006 / Mtok cache-hit input, $1.20 / Mtok output at peak. This is the production default: cheaper, faster, sufficient for most extraction.

V4-Pro $1.32 / Mtok cache-miss input, $0.044 / Mtok cache-hit input, $3.96 / Mtok output at peak. The reasoning-heavier variant for harder fields, still well under frontier-tier rates.

Context caching is automatic on every request: when a prompt shares a prefix, the API bills the cache-hit rate.1 On V4.1-Flash that prefix drops from $0.30 to $0.006 per Mtok, 2% of the cache-miss rate.

The per-filing math at scale

A 10-K body commonly lands in the 100k–150k-token range. Pricing a 120k-input, 4k-output extraction on V4.1-Flash peak rates gives about $0.041 per filing.1 Across 10,000 filings that is roughly $408, or $204 if the work runs off-peak. Filings repeat heavy boilerplate, so if you pin a fixed extraction schema and instruction block ahead of the filing text, the cached prefix bills at $0.006 / Mtok and the same sweep falls toward $55. V4-Pro at $1.32 / $3.96 costs about $0.174 per filing, near $1,742 across 10,000, the price of reserving the heavier model for hard numeric fields.

What V4 does not change: the accuracy floor

Cheap tokens do not mean correct extractions. The FinanceBench study showed how hard open-book financial QA is for language models: on a 150-case sample, GPT-4-Turbo with a retrieval system incorrectly answered or refused 81% of questions, and the full benchmark runs to 10,231 questions over real filings.2 Newer models do better, but the lesson holds: a budget extractor that misreads a parenthetical "(loss)" as positive is expensive in errors. Pick the cheapest model that clears your accuracy bar on an eval of your own filings, not the cheapest model outright.

Model names: what to pin

DeepSeek shipped V4-Pro and V4-Flash on 2026-04-24 and kept the legacy aliases alive for a 90-day window.3 The names deepseek-chat and deepseek-reasoner retired on 2026/07/24 15:59 UTC, and requests using them now fail.3 The Flash tier has since moved to DeepSeek-V4.1-Flash under the model name deepseek-flash; the older deepseek-v4-flash name is still accepted, but it is served by V4.1-Flash and billed at the Flash price.1 Pin deepseek-flash or deepseek-v4-pro explicitly, re-run your extraction eval after any model-version change, and keep a fallback path.

Decision guidance

  • Absolute cheapest full-filing fit: V4.1-Flash (1M context, ~$0.041/filing at peak, ~$0.020 off-peak). Confirm data-handling terms suit your use.
  • Repeated schema across thousands of filings: turn the boilerplate into a cached prefix; the sweep cost drops toward $55 per 10,000.
  • Harder numeric fields: route the difficult subset to V4-Pro and keep V4.1-Flash on the bulk.
  • High-stakes outputs: run an eval; a budget extractor feeding a frontier verifier may be the real cheapest-correct path.
  • Still calling legacy names: deepseek-chat/deepseek-reasoner no longer work; move to deepseek-flash or deepseek-v4-pro.

Connects to

References

Footnotes

  1. DeepSeek. "Models & Pricing." api-docs.deepseek.com, verified 2026-09-28. https://api-docs.deepseek.com/quick_start/pricing ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  2. Islam, Pranab et al. "FinanceBench: A New Benchmark for Financial Question Answering." arXiv:2311.11944, accessed 2026-06-17. https://arxiv.org/abs/2311.11944 ↩

  3. DeepSeek. "Change Log." api-docs.deepseek.com/updates, verified 2026-06-16. https://api-docs.deepseek.com/updates ↩ ↩2

Frequently asked questions

How much does DeepSeek V4 cost to extract one SEC filing?
On V4.1-Flash peak rates ($0.30 input / $1.20 output per Mtok), a 120k-input, 4k-output extraction is about $0.041 per filing, roughly $408 across 10,000 filings, or $204 if the sweep runs off-peak. Pinning a fixed schema prefix that hits the cache-hit rate ($0.006 per Mtok) cuts the input cost to 2% of cache-miss, pulling a 10,000-filing sweep toward $55. Verified against DeepSeek's pricing page 2026-09-28.
Does DeepSeek V4 fit a full 10-K in context?
Yes. Both V4.1-Flash and V4-Pro carry a 1M-token context window with 384K max output, so a full 10-K, whose body commonly lands in the 100k-150k-token range, fits in one call with room for instructions and few-shot examples. That removes the overlap and stitching bugs that come with forced chunking on smaller windows.
Is the cheap price the old promotional rate?
No. Earlier coverage quoted V4-Pro under a 75% promotion that has ended. The current verified list rates are V4.1-Flash $0.30/$1.20 and V4-Pro $1.32/$3.96 per Mtok at peak, halved off-peak, with no promotional discount shown on the pricing page as of 2026-09-28. Budget against the Flash list rate as the steady cost floor, and re-check it: DeepSeek has repriced several times in 2026.
Which DeepSeek model names should I use now?
The legacy aliases deepseek-chat and deepseek-reasoner retired on 2026/07/24 15:59 UTC and no longer work. The Flash tier is now DeepSeek-V4.1-Flash under the model name deepseek-flash; deepseek-v4-flash is still accepted but is served by V4.1-Flash and billed at the Flash price. Pin deepseek-flash or deepseek-v4-pro explicitly and re-run your eval after any model-version change.
Is DeepSeek V4 accurate enough for filings?
Not automatically. FinanceBench found GPT-4-Turbo with retrieval incorrectly answered or refused 81% of a 150-case sample, showing how hard open-book financial QA is. Pick the cheapest model that clears your accuracy bar on an eval of your own filings, and consider a budget extractor feeding a frontier verifier for high-stakes numeric fields.