DeepSeek API Pricing: Official vs Gateway, Before You Connect

DeepSeek API pricing through a third-party gateway: how official DeepSeek V4 rates compare with a gateway route, and when direct access is actually cheaper.

O
OurToken Team//12 min
DeepSeek API Pricing: Official vs Gateway, Before You Connect

DeepSeek API is the cheapest frontier-grade route for most developers in 2026, and its pricing has stopped being a single number you can look up once. The official DeepSeek price list has peak and off-peak tiers, a cache-hit price that is 33x cheaper than cache-miss input, and a Flash model that routes under multiple accepted names. A third-party gateway adds a second price list on top of that, usually expressed as a percentage of the official rate.

This article compares the DeepSeek API pricing you get directly from DeepSeek with the DeepSeek API pricing you get through a gateway, and shows when the gateway discount is real and when it is not. It uses a concrete example throughout: DeepSeek V4 Pro and DeepSeek V4 Flash through OurToken, an OpenAI-compatible gateway, compared against the current official DeepSeek pricing page.

Verification status: Prices in this article are dated 2026-10-08 and were verified against the live DeepSeek V4 Pro API page and DeepSeek V4 Flash API page on OurToken, and against the official DeepSeek API pricing documentation. Prices change. Confirm the current rates on both sides before committing to a budget.

Why DeepSeek API Pricing Is No Longer One Number

DeepSeek's official pricing page now lists every model with two tiers: peak and off-peak. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, excluding Chinese public holidays; all other times are off-peak, at half the peak price. That means the same request costs twice as much depending on when it runs.

The official page also distinguishes cache-hit input from cache-miss input. A request that reuses a stable prefix (a system prompt, tool schemas, or a long retrieved context) can be billed at the cache-hit rate, which is dramatically cheaper than full input.

The current official numbers for DeepSeek V4 Pro are:

Token categoryPeakOff-peak
Cache-miss input$1.32 / 1M$0.66 / 1M
Cache-hit input$0.044 / 1M$0.022 / 1M
Output$3.96 / 1M$1.98 / 1M

DeepSeek V4.1 Flash (the page lists it under deepseek-flash, and still accepts the older deepseek-v4-flash label, billing at the Flash rate) uses the same shape:

Token categoryPeakOff-peak
Cache-miss input$0.30 / 1M$0.15 / 1M
Cache-hit input$0.006 / 1M$0.003 / 1M
Output$1.20 / 1M$0.60 / 1M

A gateway does not change this. What a gateway changes is the number you multiply by.

What a Third-Party Gateway Actually Prices

A gateway such as OurToken exposes DeepSeek routes through an OpenAI-compatible endpoint, and prices them as a fraction of the official reference. It does not invent a new token category; it applies its discount to the categories the provider publishes.

The two DeepSeek routes on OurToken today are:

RouteModel IDInputOutputCache readCache write
DeepSeek V4 Prodeepseek-v4-pro$0.7920 / 1M$2.376 / 1M$0.0264 / 1M$0
DeepSeek V4 Flashdeepseek-v4-flash$0.2640 / 1M$0.792 / 1M$0.0084 / 1M$0

Both pages describe their rates as "60% of official price" against a reference rate card. Cache writes are billed at zero, which matches the official treatment of cache writes.

The API shape is identical to the official DeepSeek Chat Completions route, so the switch is a configuration change, not a rewrite:

Base URL:      https://api.ourtoken.ai/v1
Full endpoint: https://api.ourtoken.ai/v1/chat/completions
Auth header:   Authorization: Bearer YOUR_API_KEY
Model IDs:     deepseek-v4-pro | deepseek-v4-flash

This matters for the price comparison below: the gateway route is not a different product with different token economics. It is the same DeepSeek model, served through a route that bills at a lower reference rate.

Comparing the Two Price Lists

The gateway lists a single flat rate per token category, while the official list has peak and off-peak tiers. The honest comparison therefore depends on when your traffic runs.

For DeepSeek V4 Pro, comparing peak vs peak:

Token categoryOfficial peakOurToken gatewayGateway savings
Input (cache miss)$1.32$0.792040%
Cache hit$0.044$0.026440%
Output$3.96$2.37640%

And off-peak vs the flat gateway rate:

Token categoryOfficial off-peakOurToken gatewayCheaper route
Input (cache miss)$0.66$0.7920Official, 16.7% cheaper
Cache hit$0.022$0.0264Official, 16.7% cheaper
Output$1.98$2.376Official, 16.7% cheaper

The gateway discount does not beat DeepSeek's off-peak pricing for off-peak traffic. For a workload that can be scheduled into off-peak hours, the official route is about 17% cheaper on every category. For a workload that runs during peak hours, the gateway route is about 40% cheaper.

For DeepSeek V4 Flash, the comparison follows the same shape:

Token categoryOfficial peakOurToken gatewayGateway savings
Input (cache miss)$0.30$0.264012%
Cache hit$0.006$0.0084Gateway is 40% more expensive
Output$1.20$0.79234%

Flash shows why a single "X% of official" claim is misleading. The gateway's cache-read rate for Flash is not 60% of the official cache-hit price; the official cache-hit price is so low that even a discount applied to a different reference point can lose. If your Flash workload is dominated by cached prefixes, the official route at off-peak is dramatically cheaper.

The correct framing for most teams is: the gateway buys you a unified endpoint and one key across many models, and it is cheaper during peak hours. It is not cheaper on every category in every hour.

When the Gateway Route Is Worth It

The DeepSeek API gateway comparison only makes sense alongside the other costs of running a multi-model application.

You get one key and one endpoint instead of N providers. Every provider integration has a base URL, a credential, a billing cycle, and a set of provider-specific limits. A gateway collapses that into one OpenAI-compatible route. If your application calls DeepSeek, Claude, and GPT-5.x routes through the same client configuration, the OurToken model routing guide shows how to split simple requests from complex ones instead of sending everything to the most expensive model.

Retry and rate-limit behavior is centralized. A single gateway can apply a consistent retry policy, a bounded backoff, and a fallback chain across providers, which matters more than the per-token delta once your traffic is uneven.

The 60% rate is a flat simplification of a two-tier schedule. For peak-hour traffic, the gateway route is unambiguously cheaper. For unpredictable traffic that cannot be batched into off-peak windows, the flat gateway rate is cheaper on average than paying peak rates on the official list.

Prompt caching behaves the same way. The OpenAI-compatible prompt caching guide covers the accounting rule that applies to both routes: cached tokens are a subset of input tokens, and you bill them at the cache-read rate instead of the full input rate. The gateway's cache-read rates are also discounted, so caching compounds the gateway discount.

The switch does not require code changes. Because the route is OpenAI-compatible, existing openai SDK clients can point at a new base_url. The deepseek-v4-api-pricing guide covers the Python client and the cost-calculator pattern for Flash and Pro routes.

When Direct Access Is Cheaper

The official route wins in three situations, and you should not force a gateway where it does not pay.

Off-peak batch traffic. If your workload is a nightly or weekend batch job that can tolerate scheduling, the official off-peak price is roughly 16-17% cheaper than the flat gateway rate on DeepSeek V4 Pro. For a high-volume batch pipeline, that difference is real money. The LLM API pricing article walks through converting token volumes into a monthly number you can compare.

Cache-hit-dominated Flash workloads. The official Flash cache-hit price ($0.006 at peak, $0.003 off-peak) is below the gateway cache-read rate ($0.0084). If your Flash requests re-use a large stable prefix most of the time, the official route bills dramatically less on the input side.

Compliance or vendor-lock concerns. If your policy requires direct billing relationships with model providers, or you need first-party support for a specific provider, direct access is the simpler answer. A gateway is a third party in that relationship; it does not replace the provider's own account, quota, or SLA.

How to Decide With Your Own Numbers

Do not decide with the table above. Decide with your own token mix, because the split between cached and uncached input, the output share, and the peak-hour share decide which route wins.

  1. Capture a week of real usage objects from your production traffic: prompt tokens, cached tokens (under prompt_tokens_details), and completion tokens.
  2. Estimate your peak-hour share. Be honest: many "background" workloads still run during 01:00-04:00 and 06:00-10:00 UTC by default.
  3. Compute the monthly bill under the official schedule (peak vs off-peak) and under the gateway flat rate, using the current model pages as the rate source.
  4. Add the non-token costs: the gateway's unified key, centralized retries, and the option to route to other models through the same client.

A cheap way to start is to keep the gateway as the default route for peak and interactive traffic, and schedule the heavy, cacheable, off-peak batch jobs directly with the provider if the delta justifies the extra integration. The LLM model routing pattern is exactly this: route by cost and quality per task, rather than standardizing on one provider for everything.

A Worked Monthly Cost Example

To make the comparison concrete, take a realistic DeepSeek V4 Pro workload: 200 million input tokens per month, of which 80% hit the cache, plus 40 million output tokens. The exact numbers do not matter; the structure does, because the cached/uncached split and the peak-hour share decide the outcome.

First, compute the official bill assuming half the traffic falls in peak hours and half in off-peak hours. For cache-miss input, that is 40 million tokens (the uncached 20%) split evenly: 20 million at $1.32 and 20 million at $0.66. For cache-hit input, 160 million tokens split evenly: 80 million at $0.044 and 80 million at $0.022. For output, 40 million tokens split evenly: 20 million at $3.96 and 20 million at $1.98.

ComponentPeak rateOff-peak ratePeak spendOff-peak spend
Uncached input (20M each)$1.32$0.66$26.40$13.20
Cached input (80M each)$0.044$0.022$3.52$1.76
Output (20M each)$3.96$1.98$79.20$39.60
Total$109.12$54.56

The official monthly total for this mix is about $163.68.

Now the same workload on the gateway flat rate: 40 million uncached input at $0.7920 ($31.68), 160 million cached input at $0.0264 ($4.22), and 40 million output at $2.376 ($95.04), for about $130.94.

RouteMonthly spend
Official (50% peak / 50% off-peak)≈ $163.68
Gateway flat rate≈ $130.94

In this scenario the gateway saves roughly 20% — not because it beats off-peak rates (it does not), but because real production traffic is rarely cleanly schedulable, and the flat rate beats the blended peak/off-peak average. If the same workload ran 90% off-peak, the official route would pull ahead.

That is the takeaway to internalize: the gateway wins on blended, peak-heavy, or unpredictable traffic; the official route wins on disciplined off-peak batch traffic. Neither is universally cheaper, and a workload that claims to be "always off-peak" should be verified against real request timestamps before you rely on it.

A Simple Decision Checklist

  1. Do you need multiple model providers behind one key? If yes, the gateway value is the unified endpoint, not just the DeepSeek rate. If you only ever call DeepSeek, the rate comparison above is the whole decision.
  2. Can you schedule the heavy jobs off-peak? If yes, compute the official off-peak bill for those jobs and compare it with the flat gateway rate. If the delta is meaningful, run those jobs directly.
  3. Is your DeepSeek traffic cache-heavy? Pull the cached_tokens breakdown from real responses. For cache-dominant workloads, the official cache-hit rate can beat the gateway cache-read rate even at peak.
  4. Do you value centralized retries and a fallback chain? The model routing guide is the playbook: one client, one retry policy, and a cheap-model fallback for non-critical requests.
  5. Re-check the prices. Both the official pricing page and the OurToken model pages change. Re-verify them when you set a quarterly budget, and log the usage object so you can detect a billing-category change (for example, a route that stops reporting cached tokens).

Conclusion

DeepSeek API pricing is two price lists: an official schedule with peak and off-peak tiers, and a gateway flat rate expressed as a percentage of the official reference. On OurToken, DeepSeek V4 Pro lists at 60% of the official rate, which is about 40% cheaper during peak hours and about 17% more expensive than official off-peak pricing. The same shape holds for Flash, with the exception that the official cache-hit rate is so low that cache-heavy Flash workloads are cheaper on the official route.

Use a gateway when you want one OpenAI-compatible key across many models, centralized retries, and peak-hour savings. Use direct access for schedulable off-peak batch traffic, cache-heavy Flash workloads, or any policy requiring a direct provider relationship. Decide with your own token mix, not with a headline percentage.

FAQ

Is a third-party gateway DeepSeek API cheaper than official DeepSeek pricing?

It depends on the hour. The OurToken DeepSeek V4 Pro route lists at 60% of the official reference, which is about 40% cheaper during official peak hours and about 17% more expensive than official off-peak pricing.

What is the DeepSeek V4 Pro price on OurToken?

As of 2026-10-08, the live DeepSeek V4 Pro page lists $0.7920 / 1M input, $2.376 / 1M output, and $0.0264 / 1M cache read, with no cache-write charge. Confirm the page before budgeting.

What is the DeepSeek V4 Flash price on OurToken?

The live Flash page lists $0.2640 / 1M input, $0.792 / 1M output, and $0.0084 / 1M cache read. The official cache-hit rate can be lower than this, so cache-heavy Flash traffic can be cheaper on the official route.

Can I use my official DeepSeek API key through a gateway?

No. A key belongs to the provider that issued it. The gateway route uses its own key with its own base URL, even though the request shape is OpenAI-compatible.

When should I use the official DeepSeek API instead?

When your traffic can be scheduled into off-peak hours, when your Flash workload is dominated by cached prefixes, or when your policy requires a direct billing relationship with DeepSeek.