Solutions · two units, one balance

Buy the output, or buy the machine.

An AI credit buys a unit of finished work — a million tokens, a video second, a scanned page. A GPU credit buys an hour of a named accelerator. Both settle against the same balance, both are transferable, and both clear T+0. Most customers hold some of each, because the right answer changes by workload and by quarter.

12AI modalities
6GPU classes
USDsettlement currency
T+0settlement

AI credits

Twelve modalities, each with its own unit of account

The mistake most compute contracts make is pricing everything in GPU-hours when the buyer thinks in pages, minutes or agent-hours. Each modality below bills in the unit its buyer already plans in, and settles in dollars underneath.

Text & reasoning

per 1M tokens

Chat, extraction, summarisation, classification, long-horizon reasoning.

Cost driver
Context length and output length. Reasoning models bill thinking tokens as output.
Who buys it
Every software company. The default first workload.

Code

per 1M tokens

Completion, review, migration, agentic repair against a repository.

Cost driver
Repo context dominates. Caching the tree is the single biggest lever on unit cost.
Who buys it
Engineering orgs, dev-tool vendors.

Image

per image

Generation, editing, upscaling, background and product composition.

Cost driver
Resolution and step count, not prompt length.
Who buys it
Commerce, marketing, gaming, design tools.

Video

per second of output

Text-to-video, image-to-video, interpolation, restoration.

Cost driver
Seconds × resolution × frame rate. The most compute-hungry unit on this page.
Who buys it
Studios, advertising, social platforms.

Speech

per audio minute

Transcription, diarisation, translation, synthesis, voice conversion.

Cost driver
Audio duration. Realtime bidirectional costs more than batch by a wide margin.
Who buys it
Contact centres, media, accessibility, healthcare scribing.

Search & retrieval

per 1K queries

Embeddings, vector and hybrid retrieval, reranking, grounded answers.

Cost driver
Index size and recall target. Reranking is usually the expensive half.
Who buys it
Anyone with a corpus. The quiet workhorse of enterprise AI.

Documents & OCR

per page

Layout parsing, table extraction, handwriting, signature and stamp detection.

Cost driver
Page count and whether structure must be preserved for downstream systems.
Who buys it
Banks, insurers, logistics, public sector, legal.

Tabular & forecasting

per 1M rows scored

Scoring, imputation, anomaly detection, demand and risk forecasting.

Cost driver
Row count and feature width. Cheap per unit, enormous in volume.
Who buys it
Retail, energy, insurance, credit, supply chain.

Realtime vision

per stream-hour

Continuous inference on live camera and sensor feeds at the edge or in region.

Cost driver
Streams × frame rate × model size. Latency budget sets the floor on where it runs.
Who buys it
Manufacturing, logistics, physical security, transport, agriculture.

Agents

per agent-hour

Long-running tool use, planning, and multi-step execution against real systems.

Cost driver
Wall-clock time and tool calls, not a single prompt. Costs behave like a job, not a request.
Who buys it
Operations, finance, support, engineering. The fastest-growing line.

3D & simulation

per asset / sim-hour

Mesh and scene generation, physics simulation, synthetic data for robotics.

Cost driver
Asset fidelity and simulation wall-clock.
Who buys it
Robotics, industrial digital twins, games, AV programmes.

World models

per rollout-hour

Learned environment models for planning, policy training and evaluation.

Cost driver
Rollout horizon and parallel environments. Frontier-priced and frontier-scarce.
Who buys it
Robotics and AV labs, a small number of frontier research groups.

NoteRates for each modality live on the pricing page. For calibration on the text line, published frontier-model list prices currently run from $1 / $5 per million input / output tokens at the small end4 to $3 / $15 for mid-tier frontier models5. Exascale prices against the market, not against a single vendor.

GPU credits

Six classes, from the reference unit to the current ceiling

A GPU credit is an hour of a specific accelerator in a specific region, not a vague claim on “capacity”. That specificity is what makes it tradable: two H100-hours in the same region are fungible, and an H100-hour and a GB300-hour are not.

ClassMemoryBandwidthWhat it is forNotes
H100 SXM80 GB HBM33.35 TB/sThe market's reference unit. Fine-tuning, mid-size training, dense inference.The most liquid instrument on the venue — and the one whose list price varies most between providers, from roughly $1.49 to $6.98 per GPU-hour for identical silicon6.
H200 SXM141 GB HBM3e4.8 TB/sLong-context inference and memory-bound serving where H100 spills.NVIDIA describes it as the first GPU to offer 141 GB of HBM3e at 4.8 TB/s1. The memory step, not the FLOPs, is what moves serving economics.
B200192 GB HBM3e8 TB/sBlackwell-generation training and high-throughput inference.The generational step from Hopper. 192 GB per GPU3 changes what fits on a single device and therefore how many devices a run needs.
B300 (Blackwell Ultra)288 GB HBM3e8 TB/sReasoning-model inference where FP4 throughput and memory both bind.288 GB via 12-high stacks — 50% more than B200 — with roughly 1.5× B200’s dense FP4 throughput3. Confirm per-GPU figures against an NVIDIA datasheet before contracting on them.
GB200 NVL72Rack-scale, 72 GPUNVLink domainFrontier training runs that need one coherent memory domain, not a network of them.Sold as a rack, priced as a rack, and scheduled as a rack. Fragmenting an NVL72 destroys the thing you are paying for.
GB300 NVL7220.7 TB HBM3e / rack130 TB/s NVLinkThe current ceiling. Frontier training and the largest reasoning deployments.72 Blackwell Ultra GPUs and 36 Grace CPUs in one liquid-cooled rack-scale domain2. Availability, not price, is the binding constraint through 2026.

Memory and bandwidth figures are vendor specifications, cited where quoted. Regional availability and contracted rates are on the pricing page.

Why a credit and not a contract

A reserved instance is a bet you cannot unwind

The problem with the current market is not price, it is shape. Capacity is sold in year-long blocks to buyers whose demand moves weekly, which is why so much reserved capacity sits idle while spot buyers queue.

Transferable

A credit you no longer need can be sold back to the book at the clearing price. A reservation you no longer need is a sunk cost you explain to your CFO.

Priced continuously

One venue, one price, visible to both sides. Today the same H100-hour trades across roughly a 4.7× range depending on who you ask6, which is what a market without a reference price looks like.

Hedgeable

Forwards let a lab fix the cost of a training run before committing to it, and let an operator fix revenue before energising a hall. Both sides of that trade exist today and have no venue to meet on.

Power-aware

Data-centre electricity demand is projected to more than double by 2030, to around 945 TWh7. Capacity that follows power is worth more than capacity that sits where power is scarce — a market prices that, a fixed contract cannot.

Sources

Every figure above is numbered to an entry here. Links last read 27 July 2026.

  1. 1

    NVIDIA H200 Tensor Core GPU

    NVIDIA · Product page · Primary

    The NVIDIA H200 is the first GPU to offer 141 gigabytes (GB) of HBM3e memory at 4.8 terabytes per second (TB/s).
  2. 2

    NVIDIA GB300 NVL72

    NVIDIA · Product page · Primary

    ParaphraseGB300 NVL72 connects 72 Blackwell Ultra GPUs and 36 Grace CPUs in a rack-scale, liquid-cooled design with 20.7 TB of HBM3e and 130 TB/s of NVLink bandwidth.
  3. 3

    NVIDIA Blackwell Ultra B300: Full Specs, 288GB HBM3e Memory, 15 PFLOPS FP4

    server-parts.eu · 2026 · Third-party estimate

    ParaphraseB300 features 288 GB of HBM3e per GPU — 50% more than B200's 192 GB — via 12-high stacks, and delivers roughly 1.5× the dense FP4 throughput of B200.

    Third-party compilation of vendor material. Per-GPU B300 figures should be confirmed against an NVIDIA datasheet before being used commercially.

  4. 4

    Pricing — Claude Platform Docs

    Anthropic · Accessed July 2026 · Primary

    Claude Haiku 4.5 is priced at $1 per million input tokens and $5 per million output tokens.
  5. 5

    Introducing Claude Sonnet 5

    Anthropic · 2026 · Primary

    ParaphraseAvailable at an introductory price of $2 per million input tokens and $10 per million output tokens through August 31, 2026, moving to $3 / $15 thereafter.
  6. 6

    H100 Rental Prices Compared: $1.49–$6.98/hr Across 15+ Cloud Providers

    IntuitionLabs · 2026 · Third-party estimate

    ParaphraseOn-demand H100 list prices span $1.49 to $6.98 per GPU-hour across more than fifteen providers, a spread of roughly 4.7× for the same silicon.

    Survey of published list prices. Committed and reserved rates are negotiated and are not represented here.

  7. 7

    Energy and AI — Executive summary

    International Energy Agency · April 2025 · Primary

    Electricity demand from data centres worldwide is set to more than double by 2030 to around 945 terawatt-hours (TWh) … slightly more than the entire electricity consumption of Japan today.

Where a claim rests on a third-party estimate rather than the party that owns the number, the entry says so. Figures that are Exascale’s own — our rate card, our fee schedule — carry no citation, because they are ours to set rather than facts about the world.