LLM API Pricing: How to Calculate Your Real Monthly Cost

Turn LLM API rates into a real monthly bill. Calculate input, output, cached input, cache writes, retries, and routing costs with verified examples and Python.

O
OurToken Team//14 min
LLM API Pricing: How to Calculate Your Real Monthly Cost

LLM API pricing becomes predictable only when you convert the published rate card into a monthly equation: fresh input tokens, output tokens, cached input, cache writes, retries, and route choices. A $0.04 input price tells you almost nothing until you know how many tokens your product actually sends and how often failed requests are replayed. This guide builds that equation and checks it against a complete monthly workload using OurToken's live prices for GPT-5.6 Luna, GPT-5.6 Terra, Claude Sonnet 5, Claude Opus 5, and GPT-6 Astra.

The price snapshot below was verified from the OurToken model pages on September 22, 2026. LLM provider rates change, so use the formulas here with your current invoice data and recheck the pages before committing traffic.

LLM API Pricing: Where the Monthly Bill Actually Comes From

Most pricing pages lead with two numbers: input price per million tokens and output price per million tokens. Those numbers are necessary, but a monthly invoice contains several other cost drivers that never fit in a two-column rate table.

Cost driverWhy it changes the billWhere it appears
Fresh inputEvery uncached prompt and tool resultRequest usage and billing logs
OutputVisible text plus hidden reasoning or thinking tokensoutput_tokens in the usage object
Cached inputRepeated prefixes billed at a lower rateCache read fields in usage
Cache writesWriting a prefix into the cache is not freeCache creation fields in usage
RetriesFailed attempts still consume billable tokensRetry counters in application logs
Platform mechanicsCredit fees and plan features sit outside the token rateProvider pricing and account pages
Model routingA cheap route that needs four attempts can cost more than one correct expensive routeCost per successful task

The last two rows are the reason a "LLM API pricing comparison 2026" based only on published token rates can be misleading. Two providers can show the same input price and still produce very different monthly totals because their cache rules, retry behavior, and platform fees differ.

Here is the dated OurToken rate snapshot used in the examples. Prices are per million tokens.

ModelInputOutputCached inputCache writeContext
GPT-5.6 Luna$0.04$0.24$0.004$0.05250K
GPT-5.6 Terra$0.40$2.40$0.04$0.50250K
Claude Sonnet 5$0.80$4.00$0.08$1.001M
Claude Opus 5$2.00$10.00$0.20$2.501M
GPT-6 Astra$2.00$10.00$0.20$2.501.05M
GPT-5.6 Sol$1.00$6.00$0.10$1.25250K

GPT 5.6 Luna pricing is the lowest unit-cost row and is usually the first route to evaluate for high-volume classification and extraction. GPT-6 Astra is a frontier route with the same OurToken unit rate as Claude Opus 5 but a larger OpenAI context window. Claude Sonnet 5 and Claude Opus 5 API pricing are often compared for coding agents: Sonnet 5 is usually the first choice for most agentic work, while Opus 5 is reserved for tasks that need more sustained reasoning.

A rate card alone cannot choose between those rows. You need a monthly volume model.

The Real Monthly Cost Formula

The monthly cost formula is a sum over token categories. Every volume is multiplied by its rate, then divided by one million because the rates are quoted per million tokens.

monthly_cost = (
    fresh_input_tokens   / 1_000_000 * input_rate
  + output_tokens        / 1_000_000 * output_rate
  + cache_read_tokens    / 1_000_000 * cache_read_rate
  + cache_write_tokens   / 1_000_000 * cache_write_rate
)

For a single request, use the same formula without the monthly multiplier. Then multiply the per-request cost by the expected number of successful calls and add the cost of every retried attempt.

Read the usage object, not the text length

Output cost is the easiest line to undercount. A visible answer may be 400 words, but the billed output may be several thousand tokens because the model generated reasoning tokens before the visible answer. Reasoning effort also changes output volume without changing the rate.

OpenAI-style Responses usage reports input_tokens, output_tokens, and cached tokens inside input_tokens_details. Claude Messages usage reports input_tokens, output_tokens, cache_creation_input_tokens, and cache_read_input_tokens. The two formats are similar but not interchangeable.

The critical rule is not to double count cached input. In OpenAI-style usage, cached tokens are usually a breakdown inside total input, not an additional input category. In Claude-style usage, cache creation and cache read are separate fields. When you normalize both formats into one ledger, map every token to exactly one category: fresh input, cached input, cache write, or output.

Platform fees and credit mechanics

Token rates are not always the entire platform cost. OpenRouter credits are prepaid balance purchases, and OpenRouter fees are added when credits are purchased rather than per token. An OpenRouter subscription changes account features such as support and limits, but it is not a simple substitute for understanding the underlying model rate. The useful question for a budget is whether the effective cost after purchase fees, minimums, and plan differences matches the token-only calculation.

AI gateway pricing has the same requirement. A gateway can reduce integration work and improve observability, but you still need to know which line items are token charges, which are gateway fees, and which are retries or routing overhead. Include all of them in the same ledger instead of comparing only the published per-token numbers.

Build a Cost Model with Python and Usage Logs

Build the usage ledger first

Before trusting any forecast, store one log row for every model attempt: request ID, route, fresh input, cached input, cache write, output, status, and retry count. The row should also include the estimated cost. That design makes the monthly total reproducible and makes a surprise line item easy to find.

The calculator below encodes the September 22, 2026 OurToken rates and the five workloads used in the next section. It computes every row and checks the monthly total.

from dataclasses import dataclass


@dataclass(frozen=True)
class Route:
    input_rate: float
    output_rate: float
    cache_read_rate: float
    cache_write_rate: float


@dataclass(frozen=True)
class Workload:
    route: str
    calls: int
    fresh_input_per_call: int
    output_per_call: int
    cached_input_per_call: int = 0
    cache_writes_per_month: int = 0
    prefix_tokens_per_write: int = 0


ROUTES = {
    "gpt-5.6-luna": Route(0.04, 0.24, 0.004, 0.05),
    "gpt-5.6-terra": Route(0.40, 2.40, 0.04, 0.50),
    "claude-sonnet-5": Route(0.80, 4.00, 0.08, 1.00),
    "claude-opus-5": Route(2.00, 10.00, 0.20, 2.50),
    "gpt-6-astra": Route(2.00, 10.00, 0.20, 2.50),
}


def cost_usd(tokens: int, rate: float) -> float:
    return tokens / 1_000_000 * rate


def monthly_cost(workload: Workload) -> float:
    route = ROUTES[workload.route]
    return (
        cost_usd(workload.calls * workload.fresh_input_per_call, route.input_rate)
        + cost_usd(workload.calls * workload.output_per_call, route.output_rate)
        + cost_usd(workload.calls * workload.cached_input_per_call, route.cache_read_rate)
        + cost_usd(
            workload.cache_writes_per_month * workload.prefix_tokens_per_write,
            route.cache_write_rate,
        )
    )


WORKLOADS = [
    Workload("gpt-5.6-luna", 200_000, 1_200, 120),
    Workload("gpt-5.6-terra", 60_000, 5_000, 800),
    Workload(
        "claude-sonnet-5",
        20_000,
        4_000,
        1_200,
        cached_input_per_call=16_000,
        cache_writes_per_month=1_000,
        prefix_tokens_per_write=16_000,
    ),
    Workload("claude-opus-5", 2_000, 20_000, 3_000),
    Workload("gpt-6-astra", 1_000, 50_000, 10_000),
]


rows = []
for workload in WORKLOADS:
    total = monthly_cost(workload)
    rows.append((workload.route, total))
    print(f"{workload.route:18s} ${total:,.2f}")

combined = sum(total for _, total in rows)
print(f"combined monthly cost: ${combined:,.2f}")
assert abs(combined - 792.16) < 0.01

The Sonnet workload also calculates cache writes. The prefix is written 1,000 times during the month, each write is 16,000 tokens, and those writes are amortized across 20,000 calls.

Once the forecast is in place, record actual usage with a small client wrapper. The example below uses the OpenAI-compatible Responses endpoint for a Luna request, but the same pattern works for any model ID with the matching endpoint.

import json
import os
import time
import urllib.error
import urllib.request


MAX_RETRIES = 3
usage_ledger = []


def call_model(prompt: str, model: str = "gpt-5.6-luna") -> dict:
    api_key = os.environ.get("OURTOKEN_API_KEY")
    if not api_key:
        raise RuntimeError("Set the OURTOKEN_API_KEY environment variable")

    request_body = json.dumps({
        "model": model,
        "input": prompt,
        "max_output_tokens": 256,
    }).encode()

    for attempt in range(MAX_RETRIES):
        request = urllib.request.Request(
            "https://api.ourtoken.ai/v1/responses",
            data=request_body,
            headers={
                "Authorization": f"Bearer {api_key}",
                "Content-Type": "application/json",
            },
            method="POST",
        )
        try:
            with urllib.request.urlopen(request, timeout=60) as response:
                payload = json.loads(response.read().decode())
                usage = payload.get("usage", {})
                details = usage.get("input_tokens_details", {})
                usage_ledger.append({
                    "model": model,
                    "attempt": attempt + 1,
                    "input_tokens": usage.get("input_tokens", 0),
                    "cached_tokens": details.get("cached_tokens", 0),
                    "output_tokens": usage.get("output_tokens", 0),
                })
                return payload
        except urllib.error.HTTPError as exc:
            retryable = exc.code in {408, 409, 429, 500, 502, 503, 529}
            if retryable and attempt < MAX_RETRIES - 1:
                time.sleep(min(2 ** attempt, 8))
                continue
            raise
        except urllib.error.URLError:
            if attempt < MAX_RETRIES - 1:
                time.sleep(min(2 ** attempt, 8))
                continue
            raise

    raise RuntimeError("Maximum retries exceeded")


if __name__ == "__main__":
    result = call_model("Return the ticket priority as LOW, MEDIUM, or HIGH.")
    print(json.dumps(usage_ledger, indent=2))

A cURL smoke test uses the same endpoint. The retry flags are included so a network failure does not silently change the monthly count.

curl -sS --retry 3 --retry-delay 1 --max-time 60 \
  https://api.ourtoken.ai/v1/responses \
  -H "Authorization: Bearer $OURTOKEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-luna",
    "input": "Return the request category as JSON.",
    "max_output_tokens": 128
  }'

Worked Monthly Scenario: One Product, Five Model Routes

Consider a product that combines support triage, document drafting, a coding agent, hard review tasks, and long-document synthesis. Each job uses a different model route because the tasks have different quality requirements and output volumes.

WorkloadModelCallsFresh input / callCached input / callOutput / callCache writesCost / requestMonthly
Support triageGPT-5.6 Luna200,0001,20001200$0.0000768$15.36
Draft assistantGPT-5.6 Terra60,0005,00008000$0.00392$235.20
Coding agentClaude Sonnet 520,0004,00016,0001,2001,000 x 16K$0.01008$201.60
Hard reviewClaude Opus 52,00020,00003,0000$0.07$140.00
Document synthesisGPT-6 Astra1,00050,000010,0000$0.20$200.00
Combined283,000$792.16

The per-request numbers are all derivable. Luna is 1,200 / 1,000,000 x $0.04 = $0.000048 for input plus 120 / 1,000,000 x $0.24 = $0.0000288 for output, totaling $0.0000768. Terra is 5,000 x $0.40 + 800 x $2.40, all divided by one million, totaling $0.00392.

Sonnet includes cache. Each request costs $0.0032 for fresh input, $0.00128 for cached input, and $0.0048 for output. The 1,000 cache writes add 16,000,000 x $1.00 / 1,000,000 = $16.00 for the month, or $0.0008 per request when amortized across 20,000 calls. The row total is therefore $0.01008 per request and $201.60 per month.

When the same numbers are grouped by cost category, output and fresh input are the two dominant line items.

Cost componentMonthly tokensMonthly cost
Fresh input across routes710,000,000$373.60
Output across routes112,000,000$376.96
Cached input320,000,000$25.60
Cache writes16,000,000$16.00
Total$792.16

The same traffic against the published list rates shown on the model pages would total $3,106.80. The monthly total on the verified discounted routes is $792.16, a difference of $2,314.64, or about 74.5%.

RouteOurToken monthlyPublished list-rate equivalent
GPT-5.6 Luna$15.36$76.80
GPT-5.6 Terra$235.20$1,176.00
Claude Sonnet 5$201.60$504.00
Claude Opus 5$140.00$350.00
GPT-6 Astra$200.00$1,000.00
Combined$792.16$3,106.80

Those numbers are a dated illustration, not a promise about every workload. The important engineering habit is to build the same table from your own token logs, model IDs, cache hit rate, and retry rate.

Cost Levers That Change the Same Traffic

Cache stable prefixes

Cache reads are much cheaper than fresh input on the routes in this snapshot. Luna cached input is $0.004 per million tokens versus $0.04 for standard input. Sonnet cached input is $0.08 versus $0.80. The discount is large, but the prefix must be stable and positioned consistently.

Cache writes have a cost too. A one-off prompt should not be cached merely to get a lower read rate. Cache a system prompt, repository prefix, tool schema, or long shared document only when the same bytes will be reused enough times to amortize the write. Log cache creation and cache reads separately so the monthly report shows whether the optimization is paying for itself.

Route by task difficulty

There is no universally cheapest LLM API. A low-cost model that cannot complete the task may be more expensive after retries, human review, and failed user sessions. The practical routing policy is task-based: send classification and extraction to a low-cost route such as Luna, medium drafting to Terra, most agentic coding to Sonnet 5, and only hard reviews or long-horizon work to Opus 5 or Astra.

The LLM model routing guide covers policy and fallback design. For cost reporting, add one more requirement: measure cost per successful task, not cost per request. A route that reaches the correct answer on the first attempt can win even when its token rate is higher.

Put a budget on retries

Retries are legitimate for rate limits, transient network errors, and selected 5xx responses. Unlimited retries are not a reliability strategy. Set a maximum attempt count, use exponential backoff with a cap, and do not retry validation errors or authentication failures. Record every attempt in the ledger because a retried request that eventually succeeds still contributes tokens to the monthly bill.

The same rule applies to model fallback. If the primary route fails after consuming a long prompt, the fallback request may send the prompt again. That is a second billable request. Measure the full chain from first attempt to final success rather than reporting only the successful attempt.

Troubleshooting your monthly bill

SymptomLikely causeFix
Output dominates the billReasoning tokens or longer completions than expectedInspect output_tokens, lower reasoning effort where possible, and cap output length
Cached input stays at zeroPrefix changed, moved, or expiredKeep the shared prefix exact and first, then confirm cache fields in usage
Small batch costs more than forecastCache writes were not amortized across enough requestsCache only repeated prefixes with a clear reuse period
Invoice exceeds token-only estimateCredit purchase fee, plan cost, or another platform line itemReconcile platform fees separately from model token charges
Bill rises without a traffic increaseHigher reasoning effort, longer outputs, or more retriesVersion the request configuration and compare cost per successful task by release
401 on an OurToken requestMissing or malformed Bearer keyRead the key from an environment variable and send Authorization: Bearer ...
404 or model not foundWrong endpoint for the model family or an incorrect model IDUse /v1/responses for OpenAI models and /v1/messages for Claude models

Conclusion

LLM API pricing is a monthly modeling problem, not a two-column rate-card problem. Multiply every token category by its current rate, include cache writes and retried attempts, and group the result by model route. The dated example in this guide produces $792.16 for a mixed workload that would total $3,106.80 at the published list-rate equivalents. Your own traffic should replace those volumes before any routing decision is made.

Start with one high-volume workload, log usage on every attempt, and compare the modeled total with the real invoice after the first full month. Then move the next workload behind the same reporting layer. A small test key on the API key page is enough to verify the endpoint and usage fields before scaling production traffic.

FAQ

How do I calculate LLM API monthly cost?

Multiply monthly fresh input, output, cached input, and cache-write token volumes by their per-million-token rates, add platform fees, and include every retried attempt. Do not estimate from the number of user messages alone.

Why is my invoice higher than the token-rate estimate?

The most common causes are hidden reasoning tokens billed as output, longer completions than expected, cache writes, retries, and platform fees that sit outside the token rate. Use the provider usage object rather than visible word count.

Does prompt caching always reduce cost?

No. Cached reads are cheaper, but cache writes also cost money. Caching wins when a stable prefix is reused enough times during the cache lifetime. One-off prompts usually do not benefit.

What is the cheapest LLM API?

There is no permanent answer because model prices and task requirements change. In this snapshot, GPT-5.6 Luna has the lowest unit cost of the examples, but the cheapest successful route depends on accuracy, retries, output length, and the work being performed.

Are OpenRouter fees charged per token?

OpenRouter charges a fee when credits are purchased on its standard payment path. That is a platform-level cost, not a per-token surcharge, so forecast it separately from model token rates. Recheck the current OpenRouter pricing page before relying on a fee percentage.

Do I need an OpenRouter subscription or monthly platform plan?

Only if the plan's limits or features match your workload. A subscription does not replace the need to calculate effective token cost. For variable workloads, compare the plan fee plus usage against a pay-as-you-go route with no monthly platform commitment.

How should I compare LLM API pricing across providers?

Compare the exact model ID, all token categories, context limits, cache rules, platform fees, retry behavior, and quality on your own replay set. Rank providers by cost per successful task, not by the lowest advertised input price.