Skip to content
Token Perks

Menu

Compare the tracked offers

Verified Sep 6–7 2026

Version 2 · dated Sep 7 2026

Methodology

We publish point-in-time research snapshots, not live prices. Every figure carries a verification date; every page tells you to re-verify at official terms before paying. Version 2 extends the v0.1 verification rules with the cost-ledger data model and a strict policy for quoting third-party benchmark scores.

The rules (retained from v0.1)

  1. Official sources only. Provider pages and in-product listings. No forum screenshots, no second-hand price tweets.
  2. Median-window aggregation. Per-task references use the median over a trailing 7-day example window — stated cadence, stated window, no silent averaging.
  3. Dual display. Unit price (what the provider publishes) alongside cost-per-task (what you feel).
  4. Uncertainty labels. Every evidence item is marked. We never invent verification dates.

Evidence labels

Every row in the full price table carries one of three labels:

DIRECT
Read directly on the provider's own page during this pass.
EXCERPT
Official copy obtained via a site snapshot or a search-indexed rendering of the official page — the numbers are the provider's, the fetch was indirect.
UNCERTAIN
Not verified this pass. The row appears with the price field marked “not fetched” — never a figure from memory. UNCERTAIN rows are included in the full price table on purpose: gaps are data.

Cost arithmetic

The leaderboard ranks pay-per-token routes on blended effective cost per million tokens:

blended $/M = (3 × input + 1 × output) / 4

The 3:1 input-to-output weighting mirrors a write-heavy agent workload; it is our convention, stated everywhere the number appears, and never attributed to providers. List prices only: cache discounts, off-peak windows, and long-context surcharges are recorded in each row's notes and caveats, not averaged in. Estimated $/task multiplies the blend by a task-size preset (12k, 60k, 120k, or 200k tokens — our estimates). Subscription rows without published quotas get no per-task figure at all rather than a guess.

batch $/M = blended $/M × (1 − published batch discount)

Batch pricing is a separate computed column, not part of the blend: each row's stated batch modifier (for example −50%) scales that row's blended figure, with batch $/task at the same task-size preset. It appears only where the provider publishes a batch rate — a dash means no published batch modifier, never an assumed one. OpenAI's Batch −50% is a cross-cutting modifier documented on the provider pricing page and applied to that page's per-token rows; every other batch figure comes from the row's own notes.

Overage and cache terms

Two optional field groups extend the ledger rows beyond list prices. Both are quoted from official provider docs — never inferred, never filled from memory:

  • Overage. Subscription, credit, and coding-tool rows carry an overage group: the per-unit excess or top-up rate (per extra seat, per extra request, per credit pack), the official page it was read on, and the access date. Cap behavior alone is not a rate — where the docs describe caps but publish no dollar figure, the rate reads “not published”.
  • Cache terms. API per-token rows carry a cacheTerms group: TTL, minimum cacheable or billable tokens, cache-write fee, and cached-input read discount, each with its official source and access date. Any term the docs do not state reads “not published”.

“Not published” is a finding, not a gap in our process: it means the official docs were checked on the stated date and carried no figure. A missing group means the terms do not apply to that row (for example, cache terms on a flat-rate subscription). Neither group enters the blended $/M blend, which stays list-price-only so rows remain comparable.

Quoted intelligence scores (AA policy)

Intelligence values on this site come from the Artificial Analysis Intelligence Index v4.3, accessed 2026-09-07. Our posture, per AA's published terms:

  • Per-datum citation. Every score travels with its source: “AA Intelligence Index v4.3 — Source: Artificial Analysis, accessed 2026-09-07,” linked to the exact AA model page or leaderboard row. Footer attribution does not substitute.
  • No reproduction. AA tables and feeds are not reproduced. Scores appear only as single quoted reference values inside cost-side tables. Machine feeds (leaderboard.json, llms.txt, llms-full.txt) contain cost data only — no AA values.
  • Estimates carried, not laundered. Where AA marks a score as an estimate, the asterisk follows it here too.
  • Coverage limits stated. Only AA's Overall index is public per model. Coding, math, and agentic sub-indices are not publicly published, so those tabs do not exist here — absence is documented, not zero-filled.

Benchmark provenance

Which index version we quote, what the number is made of, and where the detail lives (on AA's pages, linked — not reproduced here). All AA pages below were read on 2026-09-07.

  • Index version: v4.3. Confirmed the same day on the Intelligence Index methodology page, on individual AA model pages, and in AA's data-API docs (the intelligence_index_version field reports 4.3). Some AA page furniture still says v4.2 — stale copy; v4.3 is current.
  • The suite behind the number. v4.3 is a weighted average over 10 evaluations in four categories — Agents 30% (AA-Briefcase 15%, GDPval-AA v2 10%, AutomationBench-AA 5%), Coding 20% (Terminal-Bench v4.0 10%, SciCode 10%), General knowledge 30% (AA-Omniscience 15%, GDP.pdf 10%, AA-LCR v1.1 5%), Scientific reasoning 20% (Humanity's Last Exam 10%, CritPt 10%). v4.3 replaces τ3-Banking with AutomationBench-AA and upgrades Terminal-Bench to v4.0. The weights, harness conditions, and normalization are AA's methodology detail — read them at the source, not here.
  • One score per effort variant. AA scores each model endpoint and reasoning-effort setting as its own leaderboard row (e.g. “GPT-5.6 Sol (max)” vs “(medium)” vs “(Non-reasoning)”). Scores move with the effort setting, so every citation on this site names the variant quoted and links that variant's exact source row — a bare model name without its qualifier is not a complete citation.
  • What we cite vs what we link. Cited here: single index values and AA's own estimate flags, each inline with its source link (provider pages carry a full Sources & methods table). Linked, never copied: the AA leaderboard, per-model pages, and the methodology page above. No AA table, CSV, or score feed is reproduced anywhere on this site, and machine feeds carry cost data only.
  • What does not exist publicly. AA publishes no per-model Coding Index composite on its website, and there is no Math Index composite (math ability is carried inside the index by HLE and CritPt) — so this site quotes only the Overall index and documents the absence rather than filling it.
  • Consent gate. Our intelligence-x-cost weighting (TPVS, next section) is implemented and documented but renders nowhere: a merged ranking over AA's scores is a derivative work, and it ships only with AA's written consent.

Our score — documented, withheld

The Token Perks Value Score (TPVS) is our experimental weighting of intelligence against effective cost:

TPVS = I × (Cref / Ceff)α

where I is the cited AA score, Ceff the blended cost for the selected task size, Cref the median cost of the ranked set (so the formula is ordering-invariant), and α a weighting preset: Performance 0.15, Balanced 0.5, Budget 1.0. No TPVS value is rendered anywhere on the site. A merged intelligence-x-cost ranking is a derivative work over AA's scores; it ships only with AA's written consent. Until then the leaderboard ranks on cost, quotes intelligence as a cited reference, and the frontier chart shows the cost-intelligence staircase — a visualization, not a score.

Staleness gates

Cost rows re-verify weekly; active promos daily while live. Score-dependent blocks refuse to render if the intelligence snapshot is older than 35 days — AA indexes move, and an old score presented as current is worse than none; the page then shows a “score paused” banner. Every pass, including passes that find nothing, is recorded in the verification log.

Data schema

The site renders from two snapshot files, both dated:

content/leaderboard/universe.json
  { snapshot: "2026-09-07",
    rows: [ { id, provider, category: a|b|c|d|e,
              plan, listPrice, priceMonthly,
              apiIn, apiOut,          // USD per 1M tokens, null when absent
              unit, notes, caveats[],
              sourceUrl, label: DIRECT|EXCERPT|UNCERTAIN,
              offer, modelId,
              batchDiscount?,        // fraction off list (0.5 = -50%), absent when unpublished
              batchApprox?,           // true when the published wording is approximate
              overage?:               // subs / credits / tools only
                { rate,               // per-unit excess, or "not published"
                  sourceUrl, accessed },
              cacheTerms?:            // API per-token rows only
                { ttl, minTokens, writeFee, readDiscount,
                  sourceUrl, accessed } } ] }

content/intelligence/2026-09-07.json
  { accessed: "2026-09-07", source: "artificialanalysis.ai",
    indexVersion: "4.3",
    models: { <modelId>: { intelligenceIndex, estimate,
                            sourceUrl, aaName, aaVariant } } }

Machine contract: /api/leaderboard.json (cost side only), /api/offers.json (tracked offers). Snapshot datasets are published under CC-BY-4.0 with attribution.

FAQ

How often do you re-verify prices?

Full re-verification weekly; active promos are checked daily while live. Every figure carries its verification date, and every page states it is a snapshot, not a live feed.

What does 'blended $/M' mean?

A single per-million-token figure computed as (3 x input price + 1 x output price) / 4. The 3:1 ratio mirrors a write-heavy workload. It is Token Perks arithmetic, not a provider figure, and it prices each input and output at the route's current published rate — normally list, but a launch-promo price while a promo is live (rows priced at a promo caption their list-price blend in their caveats, e.g. Z.ai GLM-5.3-Flash: $0.119/M at the promo, $0.2375/M at list). Cache discounts, off-peak windows, and long-context surcharges are documented per row but not baked into the blend. Batch discounts get their own computed column instead: batch $/M = blended $/M x (1 - published batch discount), shown only where the provider publishes a batch rate; rows without one show a dash, never a guess.

Can AI subscriptions be gifted?

It depends on the offer, and we state it per offer rather than generically. Consumer AI subscriptions are tied to the account that buys them — none of our tracked offers sells a gift card or a transferable subscription — so a gift means either a $0 route the recipient activates themselves, a credit/top-up pack where the provider sells one, or paying for (or starting) the account with the recipient's consent. Refund and cancellation windows decide how safe a prepaid gift is, and they differ by offer: some refund recent unused charges, others end access at the billing-cycle close with no proration. Each offer page carries its verified refund row; the gift guide reads them side by side.

Are the intelligence scores yours?

No. Scores come from the Artificial Analysis Intelligence Index and are quoted per datum with a link to the exact source row, accessed 2026-09-07. We do not re-measure models, and we do not republish AA tables or feeds. Where AA flags a score as an estimate, we carry the flag.

Why is there no overall value ranking on the site?

A ranking that multiplies intelligence by cost is a derived work over AA's scores. Until we have written consent from AA to render that derivation, the formula ships in code and is documented here, but the leaderboard ranks on cost alone and quotes intelligence as a cited reference column.

What is the trailing-7-day median window?

Per-task cost references (like $0.80/task) use the median tokens-per-task over a trailing 7-day example window at an illustrative blended rate — medians, so one outlier session cannot skew the number.

Do you republish benchmark leaderboards?

No. AA scores appear only as individual quoted values inside our cost tables, each with its own citation and link. Machine feeds contain cost data only. Full AA tables, CSVs, and score feeds are not reproduced anywhere on this site.

Cite this page

Token Perks. "Methodology v2." Dated Sep 7 2026. https://token-perks.com/methodology/

CC-BY-4.0 with attribution. Includes the snapshot date so readers know how fresh the numbers are.