Claude Sonnet 5 Pricing: Official Rates vs a 40% Route

Claude Sonnet 5 pricing explained: $2/$10 official rates, cache and thinking costs, worked monthly scenarios, and how to run the same model at 40% of official price.

O
OurToken Team//15 min
Claude Sonnet 5 Pricing: Official Rates vs a 40% Route

Claude Sonnet 5 launched on June 30, 2026, at $2 per million input tokens and $10 per million output tokens — the same per-token pricing that made Sonnet 4 the default daily driver for thousands of teams. What sets Sonnet 5 apart is not just the price: it brings 1M-token context, 128K-token output, adaptive thinking, and agentic tool use that closes the gap with Opus on real-world coding and planning tasks.

This guide covers every rate on the official card, the levers that change your effective cost (prompt caching, batch, thinking tokens, and effort levels), worked cost scenarios you can verify yourself, and the third-party route that runs Sonnet 5 at 40 % of the official price without changing a line of code.

Official (Anthropic)OurToken (40 % route)
Input$2.00 / MTok$0.80 / MTok
Output$10.00 / MTok$4.00 / MTok
Cache read$0.20 / MTok$0.08 / MTok
Cache write (5 min)$2.50 / MTok$1.00 / MTok
Monthly savings—~60 %

Bottom line: switch your base URL from api.anthropic.com to api.ourtoken.ai and cut your Sonnet 5 bill by 60 % with zero code changes.

Claude Sonnet 5 Pricing: The Official Rate Card

The table below lists every pricing category Anthropic publishes for Sonnet 5, with the corresponding rate on OurToken (the unified API at 40 % of official pricing).

Token categoryOfficial rate ($/M tok)OurToken rate ($/M tok)Keystone
Input (standard)2.000.8040 % of official
Output (standard)10.004.0040 % of official
Prompt cache write (5 min TTL)2.501.0040 % of official
Prompt cache write (1 hour TTL)4.001.6040 % of official
Cache hits (reads)0.200.0840 % of official
Batch API input1.00n/a — OurToken standard rate is lower than Anthropic batch
Batch API output5.00n/a — same logic applies

A few notes that matter for your budget:

Cache hits cost one‑tenth of standard input. This is the single biggest lever on the card. Repeat the same system prompt or repo context and the tenth call costs you $0.20/M instead of $2.00/M.

Batch is flat 50 % off both input and output through Anthropic's own endpoint. Through a third‑party provider that does not offer batch, the standard rate may still match or beat Anthropic's batch price — run the arithmetic on your actual token mix.

There is no long‑context premium. Every token inside the 1M context window bills at the same rate. A 200k prompt costs the same per token as a 500k prompt.

Note on introductory pricing. The $2/$10 rates above were the introductory price through August 31, 2026. The current standard pricing from Anthropic is $3 per million input tokens and $15 per million output tokens. OurToken's rate adjusts proportionally: $1.20/M input and $6.00/M output — still 40 % of whatever Anthropic charges. All calculations in this guide use the introductory $2/$10 rates, which remain the most widely cited reference for Sonnet 5 pricing.

Where the rate card is misleading

The table above implies a linear model: send N tokens, get billed N × rate. In practice the final bill depends on three invisible factors that no rate card shows:

  • Thinking tokens bill as output. Sonnet 5 enables adaptive thinking by default. The reasoning tokens the model generates internally before producing its visible answer are charged at the output rate. Two identical‑looking conversations can have very different thinking depths and therefore very different output token counts.

  • Verbosity changes the token count, not the rate. Sonnet 5 is faster and more concise than Opus 5 on routine tasks, but on complex agentic traces it can still generate significant thinking overhead. Always check the usage block in the response rather than estimating from visible output.

  • Effort levels are a dial your engineers control. Sonnet 5 supports low, medium, high, and max effort levels. Flipping from medium to max effort is a single configuration flag that changes your unit economics, and no invoice line will tell you it happened.

Why Your Sonnet 5 Bill Is Not the Sticker Price

The simplest way to see the gap between sticker and bill is to compare the same task on two rate cards.

FactorOfficial standardOfficial with cache & batchOurToken standardOurToken with cache
Single 4k‑in / 800‑out chat call$0.0160—$0.0064—
30‑turn agent session (90k prefix)$5.91$1.26$2.37$0.50
Batch enrichment (30M in / 3M out)$90.00$45.00$36.00—

The numbers in each row are real (we work through the arithmetic below). The takeaway: even the most aggressive official configuration (cache + batch) can be undercut by a third‑party standard rate.

Thinking tokens are the hidden line item

When you call Sonnet 5 with the thinking parameter or when adaptive thinking is on, the model generates internal reasoning tokens that never appear in content[].text but do appear in usage.output_tokens. Those tokens bill at $10/M — the full output rate.

In practice, a 2k‑character visible answer can cost 4k–6k of billed output tokens once thinking is included. Always check the usage block in the response rather than guessing from visible output length.

Effort levels change the cost per completion

Sonnet 5 exposes an effort parameter (low, medium, high, max). Higher effort produces better answers on hard problems but uses significantly more thinking tokens.

  • low → minimal reasoning, fast, cheap, good for simple Q&A and classification
  • medium → default, balanced cost‑quality ratio
  • high → deeper reasoning, useful for coding and planning
  • max → frontier reasoning, expensive, use only for tasks that genuinely need it

The CLI argument or API parameter is one line of code; the cost difference between low and max on the same prompt can be 3–5x.

How to Cut Claude Sonnet 5 Costs

Cache the prefix

Any repeated prefix — system prompt, repo context, few‑shot examples — should be wrapped in cache_control. One cache write (at $2.50/M) saves you $1.80/M on every subsequent turn.

Practical rule: if the same text appears in three or more requests, cache it. The breakeven on a 5‑minute cache write is two reads (write $2.50 + two reads at $0.20 = $2.90 vs three standard inputs at $6.00 → $3.10 saved after three calls).

Batch what can wait

For classification, enrichment, evaluation, or any workload that tolerates asynchronous delivery, use Anthropic's Batch API. Batch is flat 50 % off and runs within 24 hours. Through a third‑party provider that does not offer batch, the standard rate may still match or beat Anthropic's batch price — run the arithmetic on your actual token mix.

Turn the effort dial down

Not every prompt needs deep reasoning. Route summarization, extraction, and simple Q&A through Sonnet 5 at low or medium effort, or better yet through a lighter model like Haiku for the cheapest possible throughput.

Route routine work to a cheaper model

Sonnet 5 is your daily driver for coding, agentic tasks, and complex knowledge work. Simple classification, fixed‑format extraction, and lightweight generation belong on Haiku. A model‑routing layer that sends cheap work to the cheap model and complex work to Sonnet 5 can cut your blended rate by 50–70 % without users noticing the difference. LLM model routing covers the implementation in detail.

Worked Claude Sonnet 5 Cost Scenarios

Each scenario below shows the arithmetic on both rate cards so you can verify every number.

Scenario 1: a single chat completion (4k in / 800 out)

A typical chat or Q&A call with 4,000 input tokens and 800 output tokens.

Official cost per call: $0.0160 → OurToken cost per call: $0.0064 (save $0.0096).

Line itemOfficialOurToken
Input: 4,000 / 1,000,000 × rate0.004 × $2.00 = $0.00800.004 × $0.80 = $0.0032
Output: 800 / 1,000,000 × rate0.0008 × $10.00 = $0.00800.0008 × $4.00 = $0.0032
Total per call$0.0160$0.0064

At 50,000 calls per month:

  • Official: $800.00
  • OurToken: $320.00
  • Savings: $480.00 (60 %)

Scenario 2: a 30‑turn coding agent session with prompt caching

A coding assistant maintains a 90,000‑token system prompt plus repo context and runs 30 turns. Each turn adds 1,000 new conversation tokens and produces 1,500 output tokens.

Official uncached: $5.91/session → official cached: $1.26/session → OurToken cached: $0.50/session (save $5.41 vs uncached).

Without cache (official): 30 turns resend the full prefix each time.

Line itemTokensRateCost
Input (90k × 30 + 30k)2,730,000$2.00$5.46
Output (1.5k × 30)45,000$10.00$0.45
Total$5.91

With 5‑minute cache (official): one cache write, then 29 cache reads.

Line itemTokensRateCost
Cache write (once)90,000$2.50$0.225
Cache reads (29 × 90k)2,610,000$0.20$0.522
New input (1k × 30)30,000$2.00$0.06
Output (1.5k × 30)45,000$10.00$0.45
Total$1.257

With cache (OurToken): same token counts at 40 % of each official rate.

Line itemTokensRateCost
Cache write (once)90,000$1.00$0.09
Cache reads (29 × 90k)2,610,000$0.08$0.2088
New input (1k × 30)30,000$0.80$0.024
Output (1.5k × 30)45,000$4.00$0.18
Total$0.5028

Per‑session comparison:

  • Official uncached: $5.91
  • Official cached: $1.26
  • OurToken cached: $0.50
  • OurToken cached is 60 % cheaper than official cached and 91 % cheaper than official uncached.

Scenario 3: a batch enrichment job (30M in / 3M out)

Process 25,000 documents, each requiring 1,200 input tokens and 120 output tokens. Total volume: 30,000,000 input tokens (30 MTok) and 3,000,000 output tokens (3 MTok).

Official batch: $45.00 → OurToken standard: $36.00 — OurToken is already cheaper than Anthropic batch, with no latency trade-off.

ConfigurationInput costOutput costTotal
Official standard30 × $2.00 = $60.003 × $10.00 = $30.00$90.00
Official batch (50 % off)30 × $1.00 = $30.003 × $5.00 = $15.00$45.00
OurToken standard (40 %)30 × $0.80 = $24.003 × $4.00 = $12.00$36.00

OurToken's standard rate ($36.00) is already 20 % cheaper than Anthropic's batch rate ($45.00), with no latency trade‑off.

Monthly rollup: the same workload on two rate cards

Combine the three scenarios into a typical team's monthly spend.

Official monthly total: $1,520.00 → OurToken monthly total: $606.00 → save $914.00 (60.1 %).

WorkloadVolumeOfficial monthlyOurToken monthly
Chat completions (scenario 1)50,000 calls$800.00$320.00
Agent sessions (scenario 2)500 sessions500 × $1.26 = $630.00500 × $0.50 = $250.00
Enrichment job (scenario 3)30M/3M tok$90.00$36.00
Total$1,520.00$606.00
Savings$914.00 (60 %)

Verification: $1,520.00 − $606.00 = $914.00. $914.00 / $1,520.00 = 0.6013 ≈ 60.1 %.

Running Sonnet 5 Through One Compatible Endpoint

You do not need separate routing code for every provider. A unified API endpoint accepts the same messages payload, the same Model ID claude-sonnet-5, and the same cache_control markers while billing at 40 % of the official rate.

┌──────────────────────┐
│ Your Application     │
│ (SDK, cURL, Python)  │
└──────┬───────────────┘
       │ POST /v1/messages
       │ Authorization: Bearer YOUR_API_KEY
       │ {"model":"claude-sonnet-5",...}
       ▼
┌──────────────────────────────────┐
│ api.ourtoken.ai/v1/messages      │
│ Base URL: https://api.ourtoken.ai/v1
│ Model ID: claude-sonnet-5         │
│ Pricing: 40 % of official         │
└──────────────────────────────────┘

cURL smoke test

curl https://api.ourtoken.ai/v1/messages \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OURTOKEN_API_KEY" \
  -d '{
  "model": "claude-sonnet-5",
  "max_tokens": 256,
  "messages": [
    {"role": "user", "content": "Estimate the cost of 4k in / 800 out tokens on Sonnet 5."}
  ]
}'

The response includes a usage object with input_tokens, output_tokens, cache_creation_input_tokens, and cache_read_input_tokens. Multiply by your rate and log every call for budget tracking.

Python: estimate cost before you ship

"""Cost estimator for Claude Sonnet 5 — official vs third‑party rates."""

RATES = {
    "official": {
        "input": 2.00, "output": 10.00,
        "cache_write": 2.50, "cache_read": 0.20,
    },
    "ourtoken": {
        "input": 0.80, "output": 4.00,
        "cache_write": 1.00, "cache_read": 0.08,
    },
}


def estimate_cost(
    input_tokens: int,
    output_tokens: int,
    cache_write_tokens: int = 0,
    cache_read_tokens: int = 0,
    calls: int = 1,
    rate_card: str = "official",
) -> float:
    r = RATES[rate_card]
    per_call = (
        input_tokens / 1_000_000 * r["input"]
        + output_tokens / 1_000_000 * r["output"]
        + cache_write_tokens / 1_000_000 * r["cache_write"]
        + cache_read_tokens / 1_000_000 * r["cache_read"]
    )
    return round(per_call * calls, 4)


if __name__ == "__main__":
    # Scenario 1 — single call
    off = estimate_cost(4000, 800, rate_card="official")
    ot = estimate_cost(4000, 800, rate_card="ourtoken")
    print(f"Single call — official: ${off:.4f}, ourtoken: ${ot:.4f}")

    # Scenario 2 — agent session with cache
    off2 = estimate_cost(30000, 45000, cache_write_tokens=90000,
                         cache_read_tokens=2610000, rate_card="official")
    ot2 = estimate_cost(30000, 45000, cache_write_tokens=90000,
                        cache_read_tokens=2610000, rate_card="ourtoken")
    print(f"Agent session — official: ${off2:.2f}, ourtoken: ${ot2:.2f}")

    # Scenario 3 — batch-equivalent
    off3 = estimate_cost(30_000_000, 3_000_000, rate_card="official")
    ot3 = estimate_cost(30_000_000, 3_000_000, rate_card="ourtoken")
    print(f"Enrichment job — official: ${off3:.2f}, ourtoken: ${ot3:.2f}")

Python: call the API with retry and usage logging

Below is a concise production-ready snippet. For the full version with configurable backoff, timeout, and batch logging, download the complete example.

import json, os, time, urllib.request

def call_sonnet_5(prompt: str, max_tokens: int = 1024) -> dict:
    api_key = os.environ["OURTOKEN_API_KEY"]
    body = json.dumps({"model": "claude-sonnet-5", "max_tokens": max_tokens,
                       "messages": [{"role": "user", "content": prompt}]}).encode()
    req = urllib.request.Request(
        "https://api.ourtoken.ai/v1/messages",
        data=body, headers={"Content-Type": "application/json",
                            "Authorization": f"Bearer {api_key}"})
    for attempt in range(3):
        try:
            with urllib.request.urlopen(req, timeout=60) as resp:
                return json.loads(resp.read().decode())
        except urllib.error.HTTPError as e:
            if e.code in (429, 502, 503) and attempt < 2:
                time.sleep(2 ** attempt)
                continue
            raise

result = call_sonnet_5("Compare the cost of 4k input tokens on Sonnet 5.")
print(json.dumps(result.get("usage", {}), indent=2))

Troubleshooting

SymptomLikely causeFix
Bill is higher than cost estimator predictedThinking tokens bill as outputReduce effort level or cap max_tokens
Cache hits never appear in the usage block5‑minute TTL expired between turnsUse 1‑hour cache write or keep session turns within 5 minutes
Output tokens dominate the billHigh verbosity on agentic tracesRoute simple work to a lighter model or reduce effort level
401 response on /v1/messagesWrong API key or missing Bearer prefixGenerate a new key at /api-keys
Model responds but discount rate is not reflectedRate card referenced an outdated dateCheck the live /models/anthropic/claude-sonnet-5 page
temperature / top_p returns a 400 errorSonnet 5 does not support non-default sampling paramsRemove temperature, top_p, and top_k from the request body

Conclusion

Claude Sonnet 5 pricing is straightforward to read and easy to mis‑budget. The official rate card says $2/$10, but the real cost per completion depends on effort level, thinking depth, cache configuration, and whether you batch. The most expensive way to use Sonnet 5 is to call each turn without caching at max effort. The cheapest official way combines prompt caching and batch — but even that can be undercut by a unified API that charges 40 % of the official rate.

If your team already writes to the messages endpoint, switching the base URL from https://api.anthropic.com/v1 to https://api.ourtoken.ai/v1 is a one‑line change that drops your Sonnet 5 bill by 60 %. The same model ID, the same cache_control markers, the same usage response structure. You keep every data integrity check you already have. Get a free key with $5 in initial credit to verify the math against your own traffic.

FAQ

What is the exact price of Claude Sonnet 5 per million tokens?

The introductory price (through August 31, 2026) is $2.00 per million input tokens and $10.00 per million output tokens. The current standard pricing is $3.00/M input and $15.00/M output. Cache reads cost $0.20/M ($0.30/M at standard), cache writes cost $2.50/M ($3.75/M at standard) for 5‑minute TTL or $4.00/M ($6.00/M at standard) for 1‑hour TTL. Batch API prices are 50 % off: $1.00/M ($1.50/M) input and $5.00/M ($7.50/M) output.

How does OurToken's Sonnet 5 price compare to Anthropic's?

OurToken charges $0.80/M input and $4.00/M output — exactly 40 % of the introductory official rate. Cache reads are $0.08/M, cache writes $1.00/M. No minimum, no commitment. If Anthropic moves to standard pricing, OurToken adjusts proportionally, maintaining the 40 % discount.

Does Sonnet 5 cost more than Sonnet 4?

Sonnet 5 maintains the same per‑token rate as Sonnet 4 ($2/$10 introductory). The key difference is that Sonnet 5's adaptive thinking may produce more output tokens per completion on complex tasks, but the rate per million tokens on the invoice is identical.

What is Sonnet 5 batch pricing?

Anthropic's Batch API cuts the standard rate in half: $1.00/M input and $5.00/M output (introductory). Through OurToken, the standard rate ($0.80/$4.00) is already below Anthropic's batch price — you save without sacrificing synchronous response time.

Does prompt caching actually reduce the bill?

Yes. A cache read costs $0.20/M instead of $2.00/M, a 90 % discount on every cached input token. In agentic sessions that repeat the same system prompt across multiple turns, prompt caching can cut the cost by 70–80 %. See the prompt caching guide for implementation patterns.

What is the cheapest way to use Claude Sonnet 5?

Combine prompt caching for repeated prefixes, batch for asynchronous jobs, set effort to medium by default (raise to high or max only on hard problems), and route through a 40 % provider. The cheapest arch is also the simplest: one endpoint, one API key, one rate card.

Can I use Sonnet 5 for real‑time user‑facing applications?

Yes. Sonnet 5 is Anthropic's fastest frontier model — it delivers the first token faster than Opus 5 and is well‑suited for chat, code review, interactive agentic sessions, and any latency‑sensitive workload. Unlike Opus, Sonnet 5 does not offer a fast mode tier because its standard speed is already competitive for real‑time use.

Does Sonnet 5 support temperature and top_p?

No. Sonnet 5 requires default values for temperature, top_p, and top_k. Setting these to non‑default values returns a 400 error. This is by design — Anthropic fixed the sampling parameters to ensure consistent adaptive thinking behavior. Remove these fields from your request body when switching to Sonnet 5.