DeepSeek API pricing
Compare token rates, cache costs and peak-hour billing.
DeepSeek API pricing depends on the model, cached versus uncached input, output tokens and the applicable time window. At the September 28, 2026 check, the official direct API lists deepseek-flash and deepseek-v4-pro. Flash is the lower-priced route; Pro needs to justify its higher cost on your tasks.
Prices below are USD per 1 million tokens on DeepSeek's direct API, checked September 28, 2026. They are not an OmniaKey quote. Calculations are worked examples, not measured invoices or evidence of model quality. Verify the provider and route before comparing bills.
Current DeepSeek API price table
| API model | Input: cache miss, off-peak / peak | Input: cache hit, off-peak / peak | Output, off-peak / peak |
|---|---|---|---|
deepseek-flash | $0.15 / $0.30 | $0.003 / $0.006 | $0.60 / $1.20 |
deepseek-v4-pro | $0.66 / $1.32 | $0.022 / $0.044 | $1.98 / $3.96 |
Source: DeepSeek's current Models & Pricing page. Cache-hit input is a different token category, not a discount applied to the whole prompt or to output.
The direct deepseek-flash route currently serves DeepSeek-V4.1-Flash; deepseek-v4-pro serves DeepSeek-V4-Pro-0813. DeepSeek says the legacy direct names deepseek-v4-flash and deepseek-v4-flash-vision-exp remain accepted but now resolve to V4.1 Flash at Flash prices. Do not infer that every gateway maps the same names the same way.
For model behavior and migration details, see the V4.1 Flash review. The V4 Pro pricing guide covers the Pro-specific history and examples.
Peak and off-peak hours use UTC
The current official schedule defines peak hours as 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak. Weekends and Chinese public holidays are off-peak in full. Off-peak token rates are half the corresponding peak rates.
Use a timezone-aware conversion when scheduling from your country. A UTC Monday can overlap a local Sunday evening; local daylight-saving changes also matter. Do not interpret the hours as Beijing time, your browser timezone or every day of the week. Verify the provider's current calendar and your billed usage near a boundary; this guide does not guess how a request crossing that boundary will be charged.
How to calculate a DeepSeek API bill
Use the rates for the selected model and time window:
cost = cache_miss_input_tokens / 1,000,000 × input_miss_rate
+ cache_hit_input_tokens / 1,000,000 × input_hit_rate
+ billed_output_tokens / 1,000,000 × output_rate
Count billed reasoning output as well as visible text. A short final answer can still have a material output bill. Read the API's usage fields; a character count or the visible answer length is not a billing measurement.
| Same token workload | Flash off-peak | Flash peak | Pro off-peak | Pro peak |
|---|---|---|---|---|
| 100K uncached input + 20K output | $0.027 | $0.054 | $0.1056 | $0.2112 |
| 90K cache-hit + 10K uncached input + 20K output | $0.01377 | $0.02754 | $0.04818 | $0.09636 |
The first Flash result is 0.1 × $0.15 + 0.02 × $0.60 = $0.027. The cached Flash result is 0.09 × $0.003 + 0.01 × $0.15 + 0.02 × $0.60 = $0.01377. These are two alternative workloads, not two charges for the same request.
The examples exclude retries, additional tool/provider charges and currency conversion. They assume the same token counts for both models and a cache hit already reported in usage. Real task costs can differ because models consume different tokens and need different numbers of attempts.
Cache hits require a reusable prefix
DeepSeek context caching reuses matching prefixes. Keep stable instructions and reusable reference content at the front of the request, and put changing questions later when the application permits it. Repeating a topic or sending roughly similar text does not prove a cache hit.
Use prompt_cache_hit_tokens and prompt_cache_miss_tokens from usage to separate the input categories. Treat the first request as a miss unless the returned usage establishes otherwise. Cache reuse is not a promise of persistent application memory or a guaranteed future hit.
Before optimizing the prompt order, remove irrelevant input, bound output and measure retries. Caching a large prompt that the task never needed can still cost more than sending a smaller one.
Flash or Pro: compare cost per accepted task
At direct off-peak rates, Pro's uncached input rate is 4.4× Flash's and its output rate is 3.3×. Its cache-hit rate is about 7.33×. There is therefore no single multiplier that describes every mixed workload.
Start by comparing representative tasks, with identical tools, acceptance checks and time budgets. Record total billed cost, retries, latency and human corrections. Then calculate:
cost per accepted task = total billed cost / tasks meeting the acceptance criteria
A cheaper token can lose its advantage after repeated failures; a higher-priced model is not automatically more reliable for your workload. For image input, the current direct pricing table lists vision for Flash and not Pro, so capability can decide the route before price does.
DeepSeek direct billing and OmniaKey billing are separate
An OmniaKey request uses an OmniaKey key and the endpoint shown in our quick start. It does not draw down a separate DeepSeek account. A DeepSeek key belongs to DeepSeek's direct service.
Compare the exact model ID, input/output rates, cache handling and usage record on the route you actually call. DeepSeek's off-peak schedule, model aliases or granted balance do not automatically apply to an intermediary. Check the current OmniaKey pricing and your dashboard instead of multiplying the official table by an assumed gateway discount.
Frequently asked questions
Is the DeepSeek API free?
It is metered. DeepSeek's documentation says charges can be deducted from topped-up or granted balance, with granted balance used first when present. That does not promise a free allowance to every account. A free chat website, promotion and API balance are different products or conditions.
Do I pay monthly or per token?
The table here is direct API token pricing. Do not substitute a chat subscription price or a third-party “premium” plan for these rates. Check the terms of any separate plan you purchase.
Does cached input make output cheaper?
No. Input cache hits change the input part of the calculation. Billed output uses its own rate, including reasoning tokens charged as output.
Is the API always cheaper on weekends?
The checked direct schedule treats weekends as off-peak. Check UTC and the current official rules. That is not a guarantee about a gateway's schedule or your own timezone's weekend boundary.
Can I keep the old Flash model name?
DeepSeek documents compatibility aliases, but they now reach V4.1 Flash on its direct service. Prefer the documented current ID and re-evaluate behavior after a model change. Check a gateway's catalog separately.