Skip to content
User manuals
ManualsBuy inference
Buy inference

Understanding token prices

Input, output, cache, and input-equivalent allowances.

4 min read·User manual

Read prices in the same unit

Public model comparisons use USD per one million tokens. The original source may publish a price per token; the detail view preserves the source units and shows comparable token rates separately.

Input and output have separate prices. Cache reads, cache writes, batch processing, long-context tiers, and non-text operations may have their own conditions or units. Your accepted rate card controls your agreement.

What 100 MT means

100 MT means 100,000,000 input-equivalent tokens. This is the minimum for a new agreement.

Input-equivalent allowance describes consumption relative to the agreement’s input base rate. It is not a promise of 100 million output tokens or 100 million requests. A more expensive usage class consumes more equivalent allowance.

A worked example

The following rates are illustrative, not a live offer:

Usage classAgreed price per millionEquivalent multiplier
Uncached input$1.00
Output$4.00
Cache read$0.100.1×
Cache write$1.251.25×

At these example rates, 100 MT represents $100 of consumption. Suppose a request group uses one million uncached input tokens, 250,000 output tokens, and 500,000 cache-read tokens:

UsageCharge
1,000,000 input tokens$1.00
250,000 output tokens$1.00
500,000 cache-read tokens$0.05
Total$2.05 — 2.05 million input-equivalent tokens

The service uses the exact agreed prices and actual token quantities for accounting. Rounded on-screen multipliers help you understand the result; they do not control settlement.

Cache and batch conditions

A published cache price does not guarantee a cache hit. Cache reads and writes are separate usage classes, and a provider may have different retention-duration or tier rules. Cached input must not also be charged as ordinary uncached input.

Batch is a distinct execution mode that needs a supported adapter and an agreed delivery deadline. It is not a universal discount. Do not combine batch and cache reductions unless the provider’s accepted schedule explicitly supports that combination.

Zero prices and other units

Some public models have a zero input reference price. A zero price cannot define a division-based input-equivalent multiplier. A standard wholesale allowance needs a separately defined positive input rate card before it can be traded.

Images, audio duration, video duration, search calls, and per-request fees are not ordinary text tokens. They need explicit units, supported capabilities, and applicable quoted terms. An unsupported billable operation is rejected rather than assigned an invented token charge.

Prices in a multi-provider agreement

An RFQ can be fulfilled through multiple accepted provider allocations. Each allocation keeps its own rate card, allowance, remaining value, and delivery terms. Use Provider allocation rate cards in the agreement to inspect these prices; the agreement overview does not replace the individual cards.

For an illustrative 100 MT request, a 60 MT allocation at $1 per million input represents $60, while a 40 MT allocation at $2 per million input represents $80. The combined agreed value is $140. These are example rates, not a live offer.

Output and cache consumption use the rate card of the allocation that actually serves the request. Do not apply one provider’s output multiplier or cache price to every provider in the agreement.

Before accepting

Check input and output prices, cache conditions, supported execution modes, tier thresholds, currency, expiry, and delivery terms. Ask for clarification through your RFQ before accepting a rate card you do not understand.

Need a hand with your next step?
Tell the team what you’re trying to do. Include a record reference if you have one.
Megatron Wholesale · Actual inference, clear delivery terms.