Understanding token prices
Input, output, cache, and input-equivalent allowances.
Read prices in the same unit
Public model comparisons use USD per one million tokens. The original source may publish a price per token; the detail view preserves the source units and shows comparable token rates separately.
Input and output have separate prices. Cache reads, cache writes, batch processing, long-context tiers, and non-text operations may have their own conditions or units. Your accepted rate card controls your agreement.
What 100 MT means
100 MT means 100,000,000 input-equivalent tokens. This is the minimum for a new agreement.
Input-equivalent allowance describes consumption relative to the agreement’s input base rate. It is not a promise of 100 million output tokens or 100 million requests. A more expensive usage class consumes more equivalent allowance.
A worked example
The following rates are illustrative, not a live offer:
| Usage class | Agreed price per million | Equivalent multiplier |
|---|---|---|
| Uncached input | $1.00 | 1× |
| Output | $4.00 | 4× |
| Cache read | $0.10 | 0.1× |
| Cache write | $1.25 | 1.25× |
At these example rates, 100 MT represents $100 of consumption. Suppose a request group uses one million uncached input tokens, 250,000 output tokens, and 500,000 cache-read tokens:
| Usage | Charge |
|---|---|
| 1,000,000 input tokens | $1.00 |
| 250,000 output tokens | $1.00 |
| 500,000 cache-read tokens | $0.05 |
| Total | $2.05 — 2.05 million input-equivalent tokens |
The service uses the exact agreed prices and actual token quantities for accounting. Rounded on-screen multipliers help you understand the result; they do not control settlement.
Cache and batch conditions
A published cache price does not guarantee a cache hit. Cache reads and writes are separate usage classes, and a provider may have different retention-duration or tier rules. Cached input must not also be charged as ordinary uncached input.
Batch is a distinct execution mode that needs a supported adapter and an agreed delivery deadline. It is not a universal discount. Do not combine batch and cache reductions unless the provider’s accepted schedule explicitly supports that combination.
Zero prices and other units
Some public models have a zero input reference price. A zero price cannot define a division-based input-equivalent multiplier. A standard wholesale allowance needs a separately defined positive input rate card before it can be traded.
Images, audio duration, video duration, search calls, and per-request fees are not ordinary text tokens. They need explicit units, supported capabilities, and applicable quoted terms. An unsupported billable operation is rejected rather than assigned an invented token charge.
Prices in a multi-provider agreement
An RFQ can be fulfilled through multiple accepted provider allocations. Each allocation keeps its own rate card, allowance, remaining value, and delivery terms. Use Provider allocation rate cards in the agreement to inspect these prices; the agreement overview does not replace the individual cards.
For an illustrative 100 MT request, a 60 MT allocation at $1 per million input represents $60, while a 40 MT allocation at $2 per million input represents $80. The combined agreed value is $140. These are example rates, not a live offer.
Output and cache consumption use the rate card of the allocation that actually serves the request. Do not apply one provider’s output multiplier or cache price to every provider in the agreement.
Before accepting
Check input and output prices, cache conditions, supported execution modes, tier thresholds, currency, expiry, and delivery terms. Ask for clarification through your RFQ before accepting a rate card you do not understand.