Products · one ledger, three surfaces

A marketplace is two consoles and a price between them.

Most compute companies build one side. Buy-side clouds resell capacity they do not own; supply-side tools monitor fleets they cannot sell from. Neither produces a price. Exascale runs both consoles against one settlement ledger and puts an order book in the middle, which is the only arrangement that produces a number both sides can plan against.

3surfaces
1settlement ledger
4instrument types
T+0clearing

The three surfaces

Each one is a product on its own. Together they are a venue.

Demand

Cloud & API

AI enterprises, frontier labs, product teams

Buying compute means negotiating separately with every operator, on annual terms, at prices you cannot benchmark.

  • Provision across facilities from one catalog, priced at the clearing rate
  • Spend AI credits or GPU credits from a single balance
  • Managed control plane — you submit work, we operate the scheduler
  • Per-second metering on the same ledger the seller is paid from
Live at /cloud and /platform
Supply

Substrate

Data centers, neocloud operators, enterprises with idle fleets

Idle capacity is invisible and unsellable. Utilization below the contracted floor is pure loss, and there is nowhere to sell the gap.

  • List spare capacity by SKU, region and window; the market prices it
  • Per-node telemetry — utilization, thermals, power draw, what is running
  • Proof-of-capacity attestation before a node can be sold against
  • Per-second metering and payout against escrowed settlement
Live at /supply
Exchange

The book

Traders, desks, market makers, and both sides above

Without a venue there is no reference price, so neither side can hedge, benchmark, or plan against anything.

  • Central limit order book, price-time priority, per-SKU per-region instruments
  • Credits, spot, forwards and perpetual futures
  • T+0 clearing against segregated escrow
  • FIX 4.4, REST and WebSocket access
Live at /trade

Why the middle matters

The same GPU-hour trades across a 4.7× range

This is the whole argument for a venue, and it is checkable today. Published on-demand list prices for an H100-hour run from roughly $1.49 to $6.98 across more than fifteen providers1 — specialized neoclouds cluster at $1.50–$2.50 while the large general-purpose clouds sit in the mid-single digits2.

$1.49 · marketplace floor$1.50–2.50 · neoclouds$6.98 · general-purpose cloud

A 4.7× spread on an identical, fungible, commodity input is not a pricing strategy — it is the absence of a price. Every mature commodity resolved this the same way: a venue, a reference rate, and instruments written against it. Compute has the volume and the volatility, and does not yet have the venue.

Instruments

What actually trades

Four instruments, each answering a question one side of the market is already asking.

InstrumentAnswersBought bySettles
CreditsI want a balance I can hold, transfer, and spend across operators.Both sides. The settlement asset the other three clear intoEscrowed USD
SpotI need capacity now, at whatever it costs now.Product teams, burst inference, anyone with a queueT+0
ForwardsI need to know what a Q4 training run costs before I commit to it.Labs and enterprises with planned runs; operators fixing revenuePhysical delivery
Perpetual futuresI want to hedge or express a view on the price without taking capacity.Desks, market makers, anyone managing compute cost exposureCash · no delivery

NoteCredits, forwards and perpetual futures are live on the terminal as instruments. Regulatory permissions differ by venue and jurisdiction and are covered under offices and entities.

Timing

Why this market exists now and did not exist in 2022

The volume arrived

The four largest US cloud providers have guided to roughly $725B of 2026 capital expenditure, up about 77% year on year3. A market needs something to trade, and there was not enough of it three years ago.

Power became the constraint

Data-center electricity demand is set to more than double by 2030 to around 945 TWh4. When the binding input is power rather than silicon, capacity stops being uniform and starts having a location-dependent price.

Supply fragmented geographically

The US accounted for 45% of data-center electricity consumption in 2024, China 25%, Europe 15%5. Regional dispersion plus mobile demand is the precondition for arbitrage, and arbitrage is what makes a book liquid.

The buyers became sophisticated

Teams spending nine figures on compute already think in forward curves and unit economics. They have been asking for instruments their vendors cannot write.

Exchange

The venue

Spot, forward and cash-settled perpetual futures on GPU-hours, on one book with one clearing and settlement stack behind it.

Tenors

Spot, Forward, Futures

Three ways to hold the same underlying. Spot is capacity now, delivered and metered. A forward fixes a rate and a delivery window ahead of time. Perpetual futures are cash-settled and never deliver — they exist to move the price risk, not the hours.

One book, one margin engine, one settlement. A participant long a forward and short futures is margined on the net, not on each leg separately.

Spot
Delivered capacity, metered per second, settled T+0
Forward
Fixed rate, fixed window, physically delivered
Perpetual futures
Cash-settled, funded continuously, no delivery leg
Margin
Cross by default, isolated on request
Underlyings

Asset Classes

Each accelerator is its own instrument, because each has its own supply curve and its own retirement date. Quoting them as one number would average a part that is still being installed against a part that is being written down.

Hopper
H100 SXM, H100 PCIe, H200
Blackwell
B200, B300
Grace Blackwell
GB200 NVL72, GB300 NVL72
Also listed
MI300X, MI325X, GH200, L40S, RTX 6000 Ada
Derived

AI Futures by Modality

Contracts written on tokens rather than on hardware. A buyer whose exposure is “a million image generations a month” is not naturally hedged by a GPU-hour: the mapping between the two moves every time a model or a serving stack improves.

Modality futures put the contract where the exposure actually sits — text, reasoning, vision, speech, image, video, embedding — and leave the hardware basis to the participants who want it.

Written on
Tokens or generations, by modality
Settlement
Cash, against a published modality index
Why
A buyer's exposure is to output, not to silicon
Status
Listed alongside the hardware curve
Form factor

By GPU Configuration

The same die in a different chassis is a different product. Interconnect topology, memory domain and power envelope all change what an hour can actually run, so SXM, DGX, HGX and MGX are quoted separately rather than normalized into one line.

SXM
Board-level module, NVLink domain within the node
HGX
Baseboard, eight accelerators, OEM chassis
DGX
Integrated system, vendor-validated software stack
MGX
Modular reference design, mixed accelerator support

AI Platform

The router

One API in front of the whole catalog, billed from the same balance as capacity.

Inference

AI Models Router

A single endpoint across frontier and open-weight models, routed on published price and measured latency. Every model is addressed by name and version — a routing change can never silently swap which weights answered a request.

Catalog
500+ models, eighteen modalities
Billing
Per million tokens, from the credit balance
Keys
Scoped per environment, revocable

GPU Cloud

The capacity

Dedicated accelerators, three ways to take delivery, one meter behind all of them.

Metered

GPU Hours

The base unit. Metered per second, settled against the same index the book quotes, so what an hour costs on the cloud and what it costs on the exchange are the same number rather than two prices for one thing.

Unit
One GPU-hour, metered per second
Rate
The prevailing venue rate at provisioning
Minimum
None
Scheduled

Kubernetes & Slurm

Managed schedulers on dedicated capacity. Nodes are not shared between tenants, so a neighbor cannot take the memory bandwidth a job was benchmarked with — which is the usual reason a cluster performs differently in production than in a trial.

Kubernetes
Managed control plane, GPU operator preinstalled
Slurm
Managed controller, partitioned by accelerator
Tenancy
Single-tenant nodes
Unmanaged

Bare Metal

Whole nodes with root access and no hypervisor in the path. For teams whose stack already assumes the machine, virtualisation is not a feature — it is a tax on memory bandwidth and a source of variance nobody can profile.

Access
Root, out-of-band management included
Hypervisor
None
Network
RDMA fabric, non-blocking within the pod

Data Center Network

The substrate

The layer that turns an operator's idle hours into inventory, and delivered hours into a settlement record.

Onboarding

Substrate

An operator connects a fleet once and its idle hours become listable inventory. Attestation runs at onboarding and continuously after it, so what is offered on the book is capacity that has been proved to exist rather than capacity that was typed into a form.

Onboards
Data centers, neoclouds, enterprises with idle fleet
Attestation
At connection, then continuously
Proceeds
Settled to the operator on the venue's cycle
Record

Telemetry

The metering record settlement is computed from: per-node health, power draw and utilization, recorded continuously and retained through the dispute window. Both sides of a trade read the same series.

Granularity
Per node, per second
Retention
Through the dispute window
Access
Operator sees its fleet; a buyer sees its own hours

Sources

Every figure above is numbered to an entry here. Links last read 27 July 2026.

  1. 1

    H100 Rental Prices Compared: $1.49–$6.98/hr Across 15+ Cloud Providers

    IntuitionLabs · 2026 · Third-party estimate

    ParaphraseOn-demand H100 list prices span $1.49 to $6.98 per GPU-hour across more than fifteen providers, a spread of roughly 4.7× for the same silicon.

    Survey of published list prices. Committed and reserved rates are negotiated and are not represented here.

  2. 2

    GPU Cloud Pricing Comparison 2026

    Spheron · 2026 · Third-party estimate

    ParaphraseSpecialised neoclouds cluster at $1.50–$2.50 per H100-hour while the large general-purpose clouds remain in the mid-single digits; spot capacity has traded near $1.03.
  3. 3

    Google, Microsoft, Meta, and Amazon capex spending to hit $725 billion in 2026, up 77% from last year

    Tom's Hardware · February 2026 · Reporting

    Google, Amazon, Microsoft, and Meta collectively plan to allocate $725 billion to capital expenditures in 2026 — up 77% from last year's $410 billion.

    A sum of separate company guidance ranges, not a reported figure. Individual guidance: Amazon ~$200B, Google $175–185B, Meta $115–135B, Microsoft $110–120B.

  4. 4

    Energy and AI — Executive summary

    International Energy Agency · April 2025 · Primary

    Electricity demand from data centres worldwide is set to more than double by 2030 to around 945 terawatt-hours (TWh) … slightly more than the entire electricity consumption of Japan today.
  5. 5

    Energy and AI — Energy demand from AI

    International Energy Agency · April 2025 · Primary

    The United States accounted for the largest share of global data centre electricity consumption in 2024 (45%), followed by China (25%) and Europe (15%).

Where a claim rests on a third-party estimate rather than the party that owns the number, the entry says so. Figures that are Exascale’s own — our rate card, our fee schedule — carry no citation, because they are ours to set rather than facts about the world.