Groq API Pricing vs OurToken: Cheap LLM Routes Compared

Compare Groq API pricing with OurToken for GPT-OSS, Llama, Claude 5, GPT-6 Astra, DeepSeek, and GLM. Includes plans, prompt cache math, cURL, Python, and routing guidance.

O
OurToken Team//14 min
Groq API Pricing vs OurToken: Cheap LLM Routes Compared

Groq API pricing is easy to underread if you stop at the per-token model table. The platform has a free tier, a pay-per-token Developer plan, enterprise pricing, automatic prompt-cache discounts on supported GPT-OSS routes, and rate limits that vary by account. OurToken, by contrast, is a unified LLM API gateway focused on lower-cost OpenAI and Claude routes. The two platforms are not direct substitutes; they are useful on opposite sides of a routing decision.

This guide compares the public Groq rate card with OurToken's live model pages as of September 21, 2026. It includes plan mechanics, cached-input math, cURL and Python examples, a hybrid routing design, and a migration checklist. Treat every price as a dated snapshot and recheck both provider pages before moving production traffic.

Groq API Pricing: Plans and Model Rates

Groq's billing plans divide into three levels. The Free plan has no recurring cost and gives access for experimentation. The Developer plan is pay-per-token and raises usage limits; Groq does not describe it as a fixed monthly model subscription in the public plan pages checked for this article. Enterprise accounts use custom pricing and contracts. The correct way to forecast a production bill is to start from the model rate, add your actual output volume, and inspect the limit class shown in your Groq dashboard.

The model table is small enough to read carefully. The production text-generation routes that matter for cost-sensitive API builders are the GPT-OSS family and the Llama routes. Prices are per million tokens and were checked against Groq's public model documentation on September 21, 2026.

| Model | Input | Output | Context | Max completion | Documented throughput | |---|---|---:|---:|---:|---:|---| | openai/gpt-oss-20b | $0.075 | $0.30 | 131,072 | 65,536 | About 1,000 tokens/s | | openai/gpt-oss-120b | $0.15 | $0.60 | 131,072 | 65,536 | About 500 tokens/s | | llama-3.1-8b-instant | Contact sales | Contact sales | Contact sales | Contact sales | About 560 tokens/s | | llama-3.3-70b-versatile | Contact sales | Contact sales | Contact sales | Contact sales | About 280 tokens/s |

The Llama routes are not clean public pay-per-token lines in this snapshot; Groq directs those production routes to enterprise/contact-sales pricing. For a self-serve evaluation, the GPT-OSS models are the easiest place to measure real cost and latency.

Preview models can look deceptively attractive. For example, qwen/qwen3.8-27b was listed at $0.80 input and $4.00 output in Groq's preview documentation. Preview pricing can change or the model can be removed, so a production cost plan should not depend on a preview ID. Pin the model version, test it on your replay set, and monitor its status before scaling it.

Groq's official model reference is https://console.groq.com/docs/models.md. The billing-plan reference is https://console.groq.com/settings/billing/plans. Both pages should be checked at the same time as the rate card because available models and plan limits change together.

Automatic prompt caching on Groq GPT-OSS routes

Groq prompt caching is automatic on supported GPT-OSS routes. Cached input gets a 50% discount, there is no separate cache-write fee in the documented behavior checked here, and unused cache expires after two hours. A stable prompt prefix is required: the cached portion must be exact and at the start of the request. Reordering system instructions or inserting a changing tenant header before the stable prefix can invalidate the cache.

This changes the effective input price more than the headline row suggests. openai/gpt-oss-20b has a standard input rate of $0.075 per million tokens and an effective cached-input rate of $0.0375. openai/gpt-oss-120b moves from $0.15 to $0.075. Cached tokens also do not count against the cached request's token-rate limit according to Groq's prompt-caching documentation. The practical result is that a stable agent prefix or repository snapshot is cheaper than fresh context on every repeated turn.

Groq's caching reference is https://console.groq.com/docs/prompt-caching.md. Use it to confirm which model routes currently qualify before assuming your preferred model gets the discount.

What the Groq rate card does not show

Four variables outside the model table usually explain the gap between a prototype cost and a production invoice.

  1. Output volume. A model that reasons longer can cost more than a higher-input-price model with shorter output.
  2. Cache hit rate. Prefix stability determines whether the automatic GPT-OSS discount applies at all.
  3. Rate limits. Limits vary by plan, account age, model, and current capacity. Groq's public reference pages describe the mechanics but direct teams to their dashboard for account-specific values.
  4. API compatibility. Groq's OpenAI-compatible endpoint accepts many common fields but rejects or ignores some parameters, including logprobs, logit_bias, top_logprobs, and messages[].name. N must be 1. A temperature value of 0 becomes 1e-8; use a small nonzero value such as 0.1 when you need deterministic-but-valid behavior under those semantics.

Rate-limit details are account-specific, so the production-safe rule is to load-test with your own key and inspect the x-ratelimit-* response headers. Groq's general rate-limit documentation is at https://console.groq.com/docs/rate-limits.md.

Groq vs OurToken: What the Rate Cards Actually Compare

Groq is optimized for fast inference on open-weight models. OurToken is a lower-cost unified route for OpenAI and Claude models, with some open-family routes such as DeepSeek and GLM. A direct "Groq vs OurToken" model table is only useful when the workload can actually run on either platform. For GPT-OSS-class work, both sides have relevant routes. For Claude Sonnet 5, Claude Opus 5, and GPT-6 Astra, OurToken provides a route Groq does not offer.

The table below uses OurToken's live model pages on September 21, 2026. Prices are per million tokens.

RouteTypical roleInputOutputCached inputContext
Groq openai/gpt-oss-20bFast open-weight extraction and chat$0.075$0.30$0.0375131K
Groq openai/gpt-oss-120bLarger open-weight reasoning$0.15$0.60$0.075131K
OurToken GPT-5.6 LunaHigh-volume classification and routing$0.04$0.24$0.004250K
OurToken GPT-6 AstraFrontier OpenAI work and long context$2.00$10.00$0.201,050K
OurToken Claude Sonnet 5Agentic and coding work$0.80$4.00$0.081M
OurToken Claude Opus 5Hard frontier tasks$2.00$10.00$0.201M
OurToken deepseek-v4-flashLong-context open-family route$0.264$0.792$0.00841M
OurToken glm-5.3-flashMultilingual and long-context route$0.09$0.30$0.0181.25M

The two cheapest rows illustrate the decision. Groq's gpt-oss-20b is cheaper per input token than OurToken's deepseek-v4-flash, but OurToken's gpt-5.6-luna has a lower headline input and output rate than Groq's smallest GPT-OSS model in this snapshot. If the Luna model passes your task evaluation, it can reduce both standard and cached-input cost on high-volume triage. If your task requires a specific open-weight model, Groq may be the only suitable route.

OpenRouter, DeepInfra, and Together AI

Groq vs OpenRouter is less about model overlap than about how a platform charges for access. OpenRouter aggregates many providers, passes through upstream token prices, and adds a platform fee when you purchase credits; its public pricing page has shown a 5.5% standard credit fee and an 8% crypto fee. For a detailed rate comparison, see OpenRouter Pricing vs OurToken.

DeepInfra pricing and Together AI pricing are also commonly searched by developers assembling a cheap API stack. Both platforms emphasize self-serve open-model inference, dedicated capacity, and different latency/network tradeoffs. Do not compare only the headline token price. Compare the exact model ID, cache behavior, context limit, rate limits, platform fee, network location, and output quality on your own replay set. A $0.05 input rate is not cheaper if the model needs three times as many calls to pass acceptance.

The durable answer is not a single cheapest AI API. It is a small routing surface with two or three evaluated routes. That is why an LLM API gateway is valuable only if it has clear model IDs, usage fields, and per-route monitoring. A large catalog without observable per-model behavior makes cost control harder, not easier.

Monthly Cost Scenarios with Groq and OurToken

Price differences become decisions when the same workload is multiplied across a month. Start with a high-volume support triage job: 120,000 requests, 1,000 fresh input tokens and 120 output tokens per request.

RouteMonthly inputMonthly outputMonthly total
Groq openai/gpt-oss-20b120.0M tokens14.4M tokens$13.32
Groq openai/gpt-oss-120b120.0M tokens14.4M tokens$26.64
OurToken gpt-5.6-luna120.0M tokens14.4M tokens$8.26
OurToken glm-5.3-flash120.0M tokens14.4M tokens$15.12
OurToken deepseek-v4-flash120.0M tokens14.4M tokens$43.08

The math follows one formula:

monthly_cost =
    (input_tokens  / 1,000,000 * input_rate)
  + (output_tokens / 1,000,000 * output_rate)

Groq's 20B route is 120,000,000 / 1,000,000 x $0.075 = $9.00 for input, plus 14,400,000 / 1,000,000 x $0.30 = $4.32 for output, totaling $13.32. OurToken's Luna route is $4.80 plus $3.456, totaling $8.256. In this workload, the price gap is real, but it only matters if Luna meets the quality threshold.

Cache-aware support triage

Now make 85.3% of each prompt stable. Every request sends 1,024 cached tokens and 176 fresh input tokens, plus 120 output tokens. For Groq's GPT-OSS-20B, cached input costs half the standard rate.

RouteMonthly cached inputMonthly fresh inputMonthly outputTotal
Groq openai/gpt-oss-20b122.88M tokens21.12M tokens14.4M tokens$10.51
OurToken gpt-5.6-luna122.88M tokens21.12M tokens14.4M tokens$4.79

Groq's total is 122,880,000 x $0.0000000375 = $4.608 for cached input, 21,120,000 x $0.000000075 = $1.584 for fresh input, and $4.32 for output. OurToken's total is $0.49152 for cached input, $0.8448 for fresh input, and $3.456 for output. The Luna route remains lower, but the important engineering point is that cache structure changes the cost ranking and the cash saved.

Frontier Claude work is a different comparison

Consider 20,000 requests per month with 12,000 input tokens and 800 output tokens. Groq does not serve Claude Sonnet 5, so the practical choices are Anthropic's standard route, OurToken's Claude route, or a different model such as Groq's openai/gpt-oss-120b after a quality evaluation.

RouteInput costOutput costMonthly total
Claude Sonnet 5 standard list rate240M x $2.00 = $480.0016M x $10.00 = $160.00$640.00
OurToken claude-sonnet-5240M x $0.80 = $192.0016M x $4.00 = $64.00$256.00
Groq substitute openai/gpt-oss-120b240M x $0.15 = $36.0016M x $0.60 = $9.60$45.60

Do not read the last row as proof that GPT-OSS always beats Claude. It says the open model is far cheaper if, and only if, it passes the same acceptance suite. Code generation, tool use, long-context retrieval, and instruction-following can produce very different failure costs from classification. If the Claude route is required, the verified OurToken route is the lower-cost Claude option in this comparison. If the open model passes your tests, the cheapest route may be Groq.

Reusable Python cost calculator

Keep the same arithmetic in a small function so a rate-card change does not create a spreadsheet error.

def request_cost(
    fresh_prompt_tokens,
    completion_tokens,
    input_rate,
    output_rate,
    cached_tokens=0,
    cached_rate=None,
):
    """Return a request cost in USD for per-million-token rates."""
    effective_cache_rate = input_rate / 2 if cached_rate is None else cached_rate
    input_cost = fresh_prompt_tokens / 1_000_000 * input_rate
    output_cost = completion_tokens / 1_000_000 * output_rate
    cache_cost = cached_tokens / 1_000_000 * effective_cache_rate
    return input_cost + output_cost + cache_cost


groq_20b = request_cost(
    fresh_prompt_tokens=176,
    completion_tokens=120,
    input_rate=0.075,
    output_rate=0.30,
    cached_tokens=1_024,
)
ourtoken_luna = request_cost(
    fresh_prompt_tokens=176,
    completion_tokens=120,
    input_rate=0.04,
    output_rate=0.24,
    cached_tokens=1_024,
    cached_rate=0.004,
)

print(f"groq_20b={groq_20b:.6f}")
print(f"ourtoken_luna={ourtoken_luna:.6f}")

Call both APIs with cURL

Groq's OpenAI-compatible base URL is https://api.groq.com/openai/v1. A minimal chat-completion call looks like this:

curl -sS https://api.groq.com/openai/v1/chat/completions \
  -H "Authorization: Bearer $GROQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-20b",
    "messages": [
      {"role": "system", "content": "Return only LOW, MEDIUM, or HIGH."},
      {"role": "user", "content": "Summarize this support ticket risk."}
    ],
    "temperature": 0.1,
    "max_completion_tokens": 100
  }'

OurToken uses the same base URL for its OpenAI-compatible and Claude-compatible surfaces: https://api.ourtoken.ai/v1. OpenAI-family requests use /v1/responses, and Claude-family requests use /v1/messages. The model ID changes with the route; there is no Groq-style provider prefix in OurToken model IDs.

curl -sS https://api.ourtoken.ai/v1/responses \
  -H "Authorization: Bearer $OURTOKEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-luna",
    "input": "Return a JSON object with severity and category.",
    "max_output_tokens": 128
  }'

Both calls authenticate with a Bearer key, but compatibility is not the same as identical behavior. Test usage, streaming fields, tool calls, and malformed-request errors for the exact route you intend to use. Groq's OpenAI-compatibility reference is https://console.groq.com/docs/openai.md.

Hybrid Routing: Keep Groq Where It Wins

The most useful comparison between Groq and OurToken is not one permanent winner. It is a route table that sends each task to the platform where the task meets its latency, quality, and cost requirements.

Application request
  |
  +-- task classifier / routing rule
        |
        +-- open-weight + latency-sensitive --> Groq /openai/v1/chat/completions
        |                                      openai/gpt-oss-20b or gpt-oss-120b
        |
        +-- OpenAI frontier / long context --> OurToken /v1/responses
        |                                      gpt-5.6-luna or gpt-6-astra
        |
        +-- Claude agentic / coding         --> OurToken /v1/messages
                                               claude-sonnet-5 or claude-opus-5

The routing rules should be evaluated, not written from model-brand assumptions. Run a shadow period with production-shaped prompts and compare each route on acceptance, P50/P95 latency, error rate, and effective token cost. Add a model alias or provider adapter so a failed route can be moved back without rewriting the application. The LLM model routing guide shows the fallback and policy pattern in more depth.

Use Groq first when a specific open-weight model is already validated, output length is small, latency is user-visible, or the cached prefix can be kept exact. Use OurToken first when the task requires Claude Sonnet 5, Claude Opus 5, GPT-6 Astra, a 250K-1.25M context window, or a single platform that needs both OpenAI Responses and Claude Messages shapes.

Migration and Troubleshooting

Treat a Groq-to-OurToken migration as a configuration change, not a data migration.

  1. Record your current model IDs, token counts, and monthly spend by workload.
  2. Create a small OurToken key and run a shadow test against gpt-5.6-luna for extraction-style traffic.
  3. Compare Claude work on OurToken's claude-sonnet-5 against your current Anthropic invoice.
  4. Keep Groq in the route table for the open-weight tasks that pass its acceptance suite.
  5. Move traffic in small percentages and watch cost, quality, latency, and rate-limit headers together.
SymptomLikely causeFix
Groq 400 on logprobs, logit_bias, top_logprobs, or messages[].nameUnsupported compatibility fieldRemove the field or use a Groq-supported parameter set
Groq returns an error with N != 1Multiple completion requestSet N=1
Groq cache discount does not applyPrefix changed before the cached portionKeep system and shared context exact and first
Groq 429 during load testAccount/model-specific rate limitRead x-ratelimit-* headers and check the dashboard for the current plan
OurToken 404Wrong endpoint for the model familyUse /v1/responses for OpenAI models and /v1/messages for Claude models
OurToken 401Missing or malformed Bearer keyGenerate a key and send Authorization: Bearer YOUR_API_KEY
Model ID rejectedProvider-prefixed Groq ID sent to OurTokenUse the OurToken model ID exactly as shown on its model page

The two platforms also differ in API shape. Groq's compatibility layer is chat-completions oriented. OurToken separates the OpenAI-compatible Responses path from the Claude-compatible Messages path. A one-line base-URL change is useful for the first smoke test, but the production check must cover the exact request and response fields your application depends on.

Conclusion

Groq is a strong route when the job is already served by GPT-OSS or a supported open-weight model and speed is part of the service-level objective. Its automatic cache discount can cut the effective input cost in half for stable prefixes, and its OpenAI-compatible endpoint is easy to start with. The limiting factors are model coverage, account-specific rate limits, preview churn, and unsupported compatibility fields.

OurToken is the lower-cost route when the job needs Claude Sonnet 5, Claude Opus 5, GPT-6 Astra, or a large unified context across OpenAI and Claude APIs. In the dated examples above, Luna cuts the high-volume triage bill below the cheapest GPT-OSS route, and the Claude Sonnet 5 route is 60% below the standard list comparison. The best architecture is usually both: keep Groq for fast open-weight tasks and use OurToken for the frontier and long-context traffic.

FAQ

Is Groq cheaper than OurToken?

Not for every model. In this snapshot, Groq's gpt-oss-20b is cheaper per input token than OurToken's deepseek-v4-flash, while OurToken's gpt-5.6-luna is lower on input and output than Groq's GPT-OSS-20B. The useful answer depends on the exact model ID and whether the model passes your task tests.

What does Groq charge for its Developer plan?

Groq describes Developer as a pay-per-token plan with higher limits. The public pages checked for this article do not present a fixed monthly subscription price, so quote the plan against your own account rather than assuming a recurring figure.

Does Groq serve Claude or GPT-6 Astra?

No. For those models, OurToken's Claude and OpenAI routes are the relevant comparison in this article.

How do I make Groq prompt caching apply?

Keep the cached prefix exact and first. On supported GPT-OSS routes, cached input is billed at 50% of the standard input rate, has no separate cache-write fee in the documented behavior checked here, and expires after two hours without use.

Should I replace Groq or run both?

Run both if either route serves a validated part of your workload. Keep Groq for latency-sensitive open-weight tasks and add OurToken for Claude, GPT-6 Astra, long context, and the OpenAI Responses/Messages surface. Move traffic gradually behind model aliases and compare cost, latency, and quality on the same logs.

Where can I verify the latest rates?

Check Groq's official model, billing, prompt-caching, and rate-limit pages and OurToken's live model pages on the same day. The prices in this article are snapshots from September 21, 2026, and provider pages are the authoritative source.