Exascale API
$0.00 credits
510 models
Kimi K3Text
kimi-k3

Moonshot's flagship MoE. Long-context agentic work with native tool use.

MoEflagshipagentic
Context
256k
Speed
48 tok/s
Input
$0.58 / 1M tokens
Output
$2.2 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
Kimi K2 InstructText
kimi-k2

1T-parameter MoE with 32B active. Strong tool-calling and code agents.

MoEtools
Context
128k
Speed
56 tok/s
Input
$0.44 / 1M tokens
Output
$1.75 / 1M tokens
Dedicated
8 × B200 · $54.40/hr
DeepSeek V3Text
deepseek-v3

671B mixture-of-experts with 37B active. Frontier quality at open-weight pricing.

MoEflagship
Context
128k
Speed
62 tok/s
Input
$0.27 / 1M tokens
Output
$1.1 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
Llama 3.3 70B InstructText
llama-3.3-70b

The default general-purpose workhorse. Broad tool-calling support.

generaltools
Context
128k
Speed
94 tok/s
Input
$0.23 / 1M tokens
Output
$0.4 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen2.5 72B InstructText
qwen2.5-72b

Strong multilingual and structured-output behavior.

multilingual
Context
128k
Speed
88 tok/s
Input
$0.35 / 1M tokens
Output
$0.4 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Mixtral 8x22BText
mixtral-8x22b

Sparse MoE with fast decode and permissive licensing.

MoEApache-2.0
Context
64k
Speed
77 tok/s
Input
$0.65 / 1M tokens
Output
$0.65 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
Phi-4Text
phi-4

14B punching well above weight on reasoning benchmarks. Cheapest per token.

smallcheap
Context
16k
Speed
210 tok/s
Input
$0.07 / 1M tokens
Output
$0.14 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Gemma 2 27BText
gemma-2-27b

Efficient mid-size instruct model with a permissive license.

small
Context
8k
Speed
148 tok/s
Input
$0.11 / 1M tokens
Output
$0.22 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Kimi K2 ThinkingReasoning
kimi-k2-thinking

1T-parameter MoE with 32B active. Interleaves reasoning with tool calls across long horizons. Bill by output — traces count.

MoECoTtoolslong-horizon
Context
256k
Speed
44 tok/s
Input
$0.52 / 1M tokens
Output
$2.4 / 1M tokens
Dedicated
8 × B200 · $54.40/hr
DeepSeek R1 Distill 32BReasoning
r1-distill-32b

Emits explicit reasoning traces before the answer. Bill by output — traces count.

CoTtraces
Context
64k
Speed
58 tok/s
Input
$0.18 / 1M tokens
Output
$0.72 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
QwQ 32BReasoning
qwq-32b

Long-horizon deliberate reasoning for maths and planning.

maths
Context
32k
Speed
61 tok/s
Input
$0.2 / 1M tokens
Output
$0.6 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen2.5 Coder 32BCode
qwen2.5-coder-32b

Fill-in-middle and repo-scale completion. The default for editor integrations.

FIM
Context
128k
Speed
128 tok/s
Input
$0.16 / 1M tokens
Output
$0.16 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
DeepSeek Coder V2Code
deepseek-coder-v2

236B MoE trained for agentic software tasks.

MoE
Context
128k
Speed
70 tok/s
Input
$0.28 / 1M tokens
Output
$1.12 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
StarCoder2 15BCode
starcoder2-15b

Permissively licensed base model for completion at volume.

cheapbase
Context
16k
Speed
236 tok/s
Input
$0.06 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenHands LM 32BAgent
openhands-lm-32b

Tuned for multi-step software agents — planning, editing, running tests.

SWE-benchtools
Context
128k
Speed
66 tok/s
Input
$0.24 / 1M tokens
Output
$0.96 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Hermes 3 70BAgent
hermes-3-70b

Reliable structured tool-calling and function schemas for orchestration.

function-calling
Context
128k
Speed
84 tok/s
Input
$0.3 / 1M tokens
Output
$0.45 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Watt Tool 70BAgent
watt-tool-70b

Specialized for API selection and argument construction in long tool chains.

BFCL
Context
128k
Speed
80 tok/s
Input
$0.32 / 1M tokens
Output
$0.48 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen2-VL 72BVision
qwen2-vl-72b

Document VQA, charts and screen understanding at high resolution.

VQA
Context
128k
Speed
44 tok/s
Input
$0.55 / 1M tokens
Output
$0.55 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Pixtral 12BVision
pixtral-12b

Fast, cheap captioning and multimodal classification.

fast
Context
128k
Speed
132 tok/s
Input
$0.12 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
InternVL2 76BVision
internvl2-76b

Strongest open VLM on fine-grained perception benchmarks.

flagship
Context
32k
Speed
38 tok/s
Input
$0.62 / 1M tokens
Output
$0.62 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
GOT-OCR 2.0Document / OCR
got-ocr2

End-to-end OCR emitting Markdown, LaTeX and tables from scans.

LaTeXtables
Context
—
Speed
1.4 s/page
Input
$0.004 / page
Output
$0.004 / page
Dedicated
1 × H100 · $2.14/hr
SuryaDocument / OCR
surya-ocr

Line-level detection and reading order across 90+ languages.

layout
Context
—
Speed
0.6 s/page
Input
$0.002 / page
Output
$0.002 / page
Dedicated
1 × H100 · $2.14/hr
FLUX.1 devImage
flux-1-dev

Best-in-class prompt adherence and typography for open image models.

flagship
Context
—
Speed
2.1 s @ 1024²
Input
$0.025 / image
Output
$0.025 / image
Dedicated
8 × H200 · $23.84/hr
Stable Diffusion 3.5 LargeImage
sd-3.5-large

Strong photographic realism, permissive community license.

photoreal
Context
—
Speed
1.7 s @ 1024²
Input
$0.018 / image
Output
$0.018 / image
Dedicated
1 × H100 · $2.14/hr
KolorsImage
kolors

Fast bilingual generation with ControlNet and IP-Adapter support.

cheapcontrolnet
Context
—
Speed
0.9 s @ 1024²
Input
$0.009 / image
Output
$0.009 / image
Dedicated
1 × H100 · $2.14/hr
HunyuanVideoVideo
hunyuan-video

13B text-to-video with strong motion coherence.

T2V720p
Context
—
Speed
48 s @ 5 s clip
Input
$0.42 / second of video
Output
$0.42 / second of video
Dedicated
1 × H100 · $2.14/hr
Mochi 1Video
mochi-1

Apache-licensed 10B diffusion transformer with high motion fidelity.

Apache-2.0
Context
—
Speed
34 s @ 5 s clip
Input
$0.31 / second of video
Output
$0.31 / second of video
Dedicated
1 × H100 · $2.14/hr
LTX-VideoVideo
ltx-video

Realtime-class generation for previews and iteration.

fast
Context
—
Speed
4 s @ 5 s clip
Input
$0.09 / second of video
Output
$0.09 / second of video
Dedicated
1 × H100 · $2.14/hr
Cosmos Predict 14BWorld model
cosmos-predict-14b

Physics-aware world simulation. Rolls a future forward from a frame plus an action.

physical AIvideo2world
Context
—
Speed
62 s @ 5 s rollout
Input
$0.88 / second of rollout
Output
$0.88 / second of rollout
Dedicated
1 × H100 · $2.14/hr
Cosmos Text2World 7BWorld model
cosmos-text2world-7b

Generates navigable, physically plausible environments from a description.

simulation
Context
—
Speed
38 s @ 5 s rollout
Input
$0.54 / second of rollout
Output
$0.54 / second of rollout
Dedicated
1 × H100 · $2.14/hr
Cosmos TokenizerWorld model
cosmos-tokenizer

Continuous video tokeniser used to condition world models and policies.

tokeniser
Context
—
Speed
310× realtime
Input
$0.02 / second of video
Output
$0.02 / second of video
Dedicated
1 × H100 · $2.14/hr
Hunyuan3D 2.03D
hunyuan3d-2

Image or text to textured mesh, production-ready topology.

meshPBR
Context
—
Speed
26 s/asset
Input
$0.19 / asset
Output
$0.19 / asset
Dedicated
1 × H100 · $2.14/hr
TRELLIS3D
trellis

Structured latents producing radiance fields, meshes and Gaussians from one pass.

gaussians
Context
—
Speed
18 s/asset
Input
$0.14 / asset
Output
$0.14 / asset
Dedicated
1 × H100 · $2.14/hr
Whisper large-v3Speech
whisper-large-v3

Multilingual transcription with word-level timestamps.

ASR99 langs
Context
—
Speed
142× realtime
Input
$0.0035 / minute of audio
Output
$0.0035 / minute of audio
Dedicated
1 × H100 · $2.14/hr
Parakeet TDT 0.6BSpeech
parakeet-tdt

Streaming English ASR with the lowest cost per minute in the catalog.

streaming
Context
—
Speed
310× realtime
Input
$0.0016 / minute of audio
Output
$0.0016 / minute of audio
Dedicated
1 × H100 · $2.14/hr
XTTS v2Speech
xtts-v2

Voice cloning and multilingual synthesis from a 6-second reference.

TTScloning
Context
—
Speed
88× realtime
Input
$0.012 / 1k characters
Output
$0.012 / 1k characters
Dedicated
1 × H100 · $2.14/hr
SeamlessM4T v2Speech
seamless-m4t-v2

Speech-to-speech and speech-to-text translation across 100 languages.

S2ST
Context
—
Speed
46× realtime
Input
$0.008 / minute of audio
Output
$0.008 / minute of audio
Dedicated
8 × B200 · $54.40/hr
MusicGen LargeAudio / Music
musicgen-large

Text-conditioned music generation with optional melody conditioning.

melody
Context
—
Speed
11× realtime
Input
$0.021 / second of audio
Output
$0.021 / second of audio
Dedicated
1 × H100 · $2.14/hr
Stable Audio OpenAudio / Music
stable-audio-open

Sound effects and loops up to 47 seconds at 44.1 kHz stereo.

SFX44.1kHz
Context
—
Speed
19× realtime
Input
$0.014 / second of audio
Output
$0.014 / second of audio
Dedicated
1 × H100 · $2.14/hr
BGE-M3Embedding
bge-m3

Dense, sparse and multi-vector retrieval from one pass. 100+ languages.

hybrid
Context
8k
Speed
14.8k emb/s
Input
$0.012 / 1M tokens
Output
—
Dedicated
1 × H100 · $2.14/hr
E5 Mistral 7BEmbedding
e5-mistral-7b

Instruction-tuned embeddings for task-specific retrieval.

instruct
Context
32k
Speed
3.9k emb/s
Input
$0.028 / 1M tokens
Output
—
Dedicated
1 × H100 · $2.14/hr
Jina CLIP v2Embedding
jina-clip-v2

Shared text and image embedding space for multimodal search.

multimodal
Context
8k
Speed
9.1k emb/s
Input
$0.018 / 1M tokens
Output
—
Dedicated
1 × H100 · $2.14/hr
BGE Reranker v2 m3Rerank
bge-reranker-v2-m3

Cross-encoder reranking. The cheapest large accuracy win over pure vector search.

cross-encoder
Context
8k
Speed
6.2k pair/s
Input
$0.014 / 1M tokens
Output
—
Dedicated
1 × H100 · $2.14/hr
Jina Reranker v2Rerank
jina-reranker-v2

Multilingual reranking with strong function-calling and code retrieval.

multilingual
Context
8k
Speed
5.4k pair/s
Input
$0.016 / 1M tokens
Output
—
Dedicated
1 × H100 · $2.14/hr
TabPFN v2Tabular
tabpfn-v2

In-context tabular prediction — no training step, fit in a single forward pass.

no-train
Context
—
Speed
148k row/s
Input
$0.0009 / 1k rows
Output
—
Dedicated
1 × H100 · $2.14/hr
AutoGluon TabularTabular
autogluon-tabular

Stacked ensemble AutoML over GBDT and neural baselines.

AutoML
Context
—
Speed
92k row/s
Input
$0.0014 / 1k rows
Output
—
Dedicated
1 × H100 · $2.14/hr
Chronos-Bolt BaseTime series
chronos-bolt

Zero-shot probabilistic forecasting. No per-series fitting.

zero-shot
Context
2k steps
Speed
44k series/s
Input
$0.0011 / 1k forecasts
Output
—
Dedicated
1 × H100 · $2.14/hr
TimesFM 2.0Time series
timesfm-2

Decoder-only forecasting foundation model with long horizons.

long-horizon
Context
2k steps
Speed
31k series/s
Input
$0.0015 / 1k forecasts
Output
—
Dedicated
1 × H100 · $2.14/hr
Llama Guard 3 8BSafety
llama-guard-3-8b

Input and output classification against a configurable hazard taxonomy.

moderation
Context
8k
Speed
184 tok/s
Input
$0.04 / 1M tokens
Output
$0.04 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ShieldGemma 9BSafety
shieldgemma-9b

Policy-conditioned safety scoring with calibrated probabilities.

calibrated
Context
8k
Speed
152 tok/s
Input
$0.05 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ESM-2 15BScience
esm2-15b

Protein language model embeddings for structure and function prediction.

protein
Context
—
Speed
260 seq/s
Input
$0.09 / 1k sequences
Output
—
Dedicated
1 × H100 · $2.14/hr
Boltz-2Science
boltz-2

Open biomolecular structure and binding-affinity prediction.

structure
Context
—
Speed
41 seq/s
Input
$0.42 / 1k sequences
Output
—
Dedicated
1 × H100 · $2.14/hr
Llama 3.1 8BText
llama-3-1-8b

Meta's general-purpose instruct family. Broad tool-calling support.

generaltools
Context
128k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.1-8B-Instruct
Llama 3.1 8B FP8Text
llama-3-1-8b-fp8

FP8 serving variant. Meta's general-purpose instruct family. Broad tool-calling support.

generaltoolsFP8
Context
128k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.1-8B-Instruct-FP8
Llama 3.1 8B INT4 AWQText
llama-3-1-8b-awq

INT4 AWQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.

generaltoolsINT4
Context
128k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.1-8B-Instruct-AWQ
Llama 3.1 8B INT8 GPTQText
llama-3-1-8b-gptq

INT8 GPTQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.

generaltoolsINT8
Context
128k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.1-8B-Instruct-GPTQ
Llama 3.1 70BText
llama-3-1-70b

Meta's general-purpose instruct family. Broad tool-calling support.

generaltools
Context
128k
Speed
29 tok/s
Input
$0.321 / 1M tokens
Output
$0.77 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
meta-llama/Llama-3.1-70B-Instruct
Llama 3.1 70B FP8Text
llama-3-1-70b-fp8

FP8 serving variant. Meta's general-purpose instruct family. Broad tool-calling support.

generaltoolsFP8
Context
128k
Speed
44 tok/s
Input
$0.199 / 1M tokens
Output
$0.478 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
meta-llama/Llama-3.1-70B-Instruct-FP8
Llama 3.1 70B INT4 AWQText
llama-3-1-70b-awq

INT4 AWQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.

generaltoolsINT4
Context
128k
Speed
66 tok/s
Input
$0.122 / 1M tokens
Output
$0.293 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
meta-llama/Llama-3.1-70B-Instruct-AWQ
Llama 3.1 70B INT8 GPTQText
llama-3-1-70b-gptq

INT8 GPTQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.

generaltoolsINT8
Context
128k
Speed
53 tok/s
Input
$0.161 / 1M tokens
Output
$0.385 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
meta-llama/Llama-3.1-70B-Instruct-GPTQ
Llama 3.1 405BText
llama-3-1-405b

Meta's general-purpose instruct family. Broad tool-calling support.

generaltools
Context
128k
Speed
8 tok/s
Input
$1.762 / 1M tokens
Output
$4.229 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
meta-llama/Llama-3.1-405B-Instruct
Llama 3.1 405B FP8Text
llama-3-1-405b-fp8

FP8 serving variant. Meta's general-purpose instruct family. Broad tool-calling support.

generaltoolsFP8
Context
128k
Speed
8 tok/s
Input
$1.092 / 1M tokens
Output
$2.622 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
meta-llama/Llama-3.1-405B-Instruct-FP8
Llama 3.1 405B INT4 AWQText
llama-3-1-405b-awq

INT4 AWQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.

generaltoolsINT4
Context
128k
Speed
13 tok/s
Input
$0.67 / 1M tokens
Output
$1.607 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
meta-llama/Llama-3.1-405B-Instruct-AWQ
Llama 3.1 405B INT8 GPTQText
llama-3-1-405b-gptq

INT8 GPTQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.

generaltoolsINT8
Context
128k
Speed
10 tok/s
Input
$0.881 / 1M tokens
Output
$2.114 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
meta-llama/Llama-3.1-405B-Instruct-GPTQ
Llama 3.2 1BText
llama-3-2-1b

Small-footprint Llama for edge and high-throughput serving.

smalledge
Context
128k
Speed
280 tok/s
Input
$0.024 / 1M tokens
Output
$0.058 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.2-1B-Instruct
Llama 3.2 3BText
llama-3-2-3b

Small-footprint Llama for edge and high-throughput serving.

smalledge
Context
128k
Speed
224 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.2-3B-Instruct
Llama 3.2 Vision 11BVision
llama-3-2-vision-11b

Llama with an image encoder for document and chart understanding.

vision
Context
128k
Speed
124 tok/s
Input
$0.067 / 1M tokens
Output
$0.161 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.2-11B-Vision-Instruct
Llama 3.2 Vision 11B FP8Vision
llama-3-2-vision-11b-fp8

FP8 serving variant. Llama with an image encoder for document and chart understanding.

visionFP8
Context
128k
Speed
162 tok/s
Input
$0.042 / 1M tokens
Output
$0.1 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.2-11B-Vision-Instruct-FP8
Llama 3.2 Vision 11B INT4 AWQVision
llama-3-2-vision-11b-awq

INT4 AWQ serving variant. Llama with an image encoder for document and chart understanding.

visionINT4
Context
128k
Speed
200 tok/s
Input
$0.025 / 1M tokens
Output
$0.061 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.2-11B-Vision-Instruct-AWQ
Llama 3.2 Vision 11B INT8 GPTQVision
llama-3-2-vision-11b-gptq

INT8 GPTQ serving variant. Llama with an image encoder for document and chart understanding.

visionINT8
Context
128k
Speed
179 tok/s
Input
$0.034 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-3.2-11B-Vision-Instruct-GPTQ
Llama 3.2 Vision 90BVision
llama-3-2-vision-90b

Llama with an image encoder for document and chart understanding.

vision
Context
128k
Speed
23 tok/s
Input
$0.407 / 1M tokens
Output
$0.977 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
meta-llama/Llama-3.2-90B-Vision-Instruct
Llama 3.2 Vision 90B FP8Vision
llama-3-2-vision-90b-fp8

FP8 serving variant. Llama with an image encoder for document and chart understanding.

visionFP8
Context
128k
Speed
35 tok/s
Input
$0.252 / 1M tokens
Output
$0.606 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
meta-llama/Llama-3.2-90B-Vision-Instruct-FP8
Llama 3.2 Vision 90B INT4 AWQVision
llama-3-2-vision-90b-awq

INT4 AWQ serving variant. Llama with an image encoder for document and chart understanding.

visionINT4
Context
128k
Speed
54 tok/s
Input
$0.155 / 1M tokens
Output
$0.371 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
meta-llama/Llama-3.2-90B-Vision-Instruct-AWQ
Llama 3.2 Vision 90B INT8 GPTQVision
llama-3-2-vision-90b-gptq

INT8 GPTQ serving variant. Llama with an image encoder for document and chart understanding.

visionINT8
Context
128k
Speed
43 tok/s
Input
$0.203 / 1M tokens
Output
$0.488 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
meta-llama/Llama-3.2-90B-Vision-Instruct-GPTQ
Qwen2.5 0.5BText
qwen2-5-0-5b

Strong multilingual and structured-output behavior across the size range.

multilingual
Context
128k
Speed
298 tok/s
Input
$0.022 / 1M tokens
Output
$0.053 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-0.5B-Instruct
Qwen2.5 1.5BText
qwen2-5-1-5b

Strong multilingual and structured-output behavior across the size range.

multilingual
Context
128k
Speed
263 tok/s
Input
$0.026 / 1M tokens
Output
$0.062 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-1.5B-Instruct
Qwen2.5 3BText
qwen2-5-3b

Strong multilingual and structured-output behavior across the size range.

multilingual
Context
128k
Speed
224 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-3B-Instruct
Qwen2.5 7BText
qwen2-5-7b

Strong multilingual and structured-output behavior across the size range.

multilingual
Context
128k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-7B-Instruct
Qwen2.5 7B FP8Text
qwen2-5-7b-fp8

FP8 serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualFP8
Context
128k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-7B-Instruct-FP8
Qwen2.5 7B INT4 AWQText
qwen2-5-7b-awq

INT4 AWQ serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualINT4
Context
128k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-7B-Instruct-AWQ
Qwen2.5 7B INT8 GPTQText
qwen2-5-7b-gptq

INT8 GPTQ serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualINT8
Context
128k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-7B-Instruct-GPTQ
Qwen2.5 14BText
qwen2-5-14b

Strong multilingual and structured-output behavior across the size range.

multilingual
Context
128k
Speed
106 tok/s
Input
$0.08 / 1M tokens
Output
$0.192 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-14B-Instruct
Qwen2.5 14B FP8Text
qwen2-5-14b-fp8

FP8 serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualFP8
Context
128k
Speed
142 tok/s
Input
$0.05 / 1M tokens
Output
$0.119 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-14B-Instruct-FP8
Qwen2.5 14B INT4 AWQText
qwen2-5-14b-awq

INT4 AWQ serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualINT4
Context
128k
Speed
181 tok/s
Input
$0.03 / 1M tokens
Output
$0.073 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-14B-Instruct-AWQ
Qwen2.5 14B INT8 GPTQText
qwen2-5-14b-gptq

INT8 GPTQ serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualINT8
Context
128k
Speed
160 tok/s
Input
$0.04 / 1M tokens
Output
$0.096 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-14B-Instruct-GPTQ
Qwen2.5 32BText
qwen2-5-32b

Strong multilingual and structured-output behavior across the size range.

multilingual
Context
128k
Speed
57 tok/s
Input
$0.158 / 1M tokens
Output
$0.379 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-32B-Instruct
Qwen2.5 32B FP8Text
qwen2-5-32b-fp8

FP8 serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualFP8
Context
128k
Speed
83 tok/s
Input
$0.098 / 1M tokens
Output
$0.235 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-32B-Instruct-FP8
Qwen2.5 32B INT4 AWQText
qwen2-5-32b-awq

INT4 AWQ serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualINT4
Context
128k
Speed
116 tok/s
Input
$0.06 / 1M tokens
Output
$0.144 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-32B-Instruct-AWQ
Qwen2.5 32B INT8 GPTQText
qwen2-5-32b-gptq

INT8 GPTQ serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualINT8
Context
128k
Speed
97 tok/s
Input
$0.079 / 1M tokens
Output
$0.19 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-32B-Instruct-GPTQ
Qwen2.5 72BText
qwen2-5-72b

Strong multilingual and structured-output behavior across the size range.

multilingual
Context
128k
Speed
28 tok/s
Input
$0.33 / 1M tokens
Output
$0.792 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-72B-Instruct
Qwen2.5 72B FP8Text
qwen2-5-72b-fp8

FP8 serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualFP8
Context
128k
Speed
43 tok/s
Input
$0.205 / 1M tokens
Output
$0.491 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-72B-Instruct-FP8
Qwen2.5 72B INT4 AWQText
qwen2-5-72b-awq

INT4 AWQ serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualINT4
Context
128k
Speed
65 tok/s
Input
$0.125 / 1M tokens
Output
$0.301 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-72B-Instruct-AWQ
Qwen2.5 72B INT8 GPTQText
qwen2-5-72b-gptq

INT8 GPTQ serving variant. Strong multilingual and structured-output behavior across the size range.

multilingualINT8
Context
128k
Speed
52 tok/s
Input
$0.165 / 1M tokens
Output
$0.396 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-72B-Instruct-GPTQ
Qwen3 0.6BText
qwen3-0-6b

Hybrid reasoning family with a switchable thinking mode.

reasoninghybrid
Context
128k
Speed
294 tok/s
Input
$0.023 / 1M tokens
Output
$0.055 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-0.6B
Qwen3 1.7BText
qwen3-1-7b

Hybrid reasoning family with a switchable thinking mode.

reasoninghybrid
Context
128k
Speed
257 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-1.7B
Qwen3 4BText
qwen3-4b

Hybrid reasoning family with a switchable thinking mode.

reasoninghybrid
Context
128k
Speed
203 tok/s
Input
$0.037 / 1M tokens
Output
$0.089 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-4B
Qwen3 8BText
qwen3-8b

Hybrid reasoning family with a switchable thinking mode.

reasoninghybrid
Context
128k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-8B
Qwen3 8B FP8Text
qwen3-8b-fp8

FP8 serving variant. Hybrid reasoning family with a switchable thinking mode.

reasoninghybridFP8
Context
128k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-8B-FP8
Qwen3 8B INT4 AWQText
qwen3-8b-awq

INT4 AWQ serving variant. Hybrid reasoning family with a switchable thinking mode.

reasoninghybridINT4
Context
128k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-8B-AWQ
Qwen3 8B INT8 GPTQText
qwen3-8b-gptq

INT8 GPTQ serving variant. Hybrid reasoning family with a switchable thinking mode.

reasoninghybridINT8
Context
128k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-8B-GPTQ
Qwen3 14BText
qwen3-14b

Hybrid reasoning family with a switchable thinking mode.

reasoninghybrid
Context
128k
Speed
106 tok/s
Input
$0.08 / 1M tokens
Output
$0.192 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-14B
Qwen3 14B FP8Text
qwen3-14b-fp8

FP8 serving variant. Hybrid reasoning family with a switchable thinking mode.

reasoninghybridFP8
Context
128k
Speed
142 tok/s
Input
$0.05 / 1M tokens
Output
$0.119 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-14B-FP8
Qwen3 14B INT4 AWQText
qwen3-14b-awq

INT4 AWQ serving variant. Hybrid reasoning family with a switchable thinking mode.

reasoninghybridINT4
Context
128k
Speed
181 tok/s
Input
$0.03 / 1M tokens
Output
$0.073 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-14B-AWQ
Qwen3 14B INT8 GPTQText
qwen3-14b-gptq

INT8 GPTQ serving variant. Hybrid reasoning family with a switchable thinking mode.

reasoninghybridINT8
Context
128k
Speed
160 tok/s
Input
$0.04 / 1M tokens
Output
$0.096 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-14B-GPTQ
Qwen3 32BText
qwen3-32b

Hybrid reasoning family with a switchable thinking mode.

reasoninghybrid
Context
128k
Speed
57 tok/s
Input
$0.158 / 1M tokens
Output
$0.379 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-32B
Qwen3 32B FP8Text
qwen3-32b-fp8

FP8 serving variant. Hybrid reasoning family with a switchable thinking mode.

reasoninghybridFP8
Context
128k
Speed
83 tok/s
Input
$0.098 / 1M tokens
Output
$0.235 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-32B-FP8
Qwen3 32B INT4 AWQText
qwen3-32b-awq

INT4 AWQ serving variant. Hybrid reasoning family with a switchable thinking mode.

reasoninghybridINT4
Context
128k
Speed
116 tok/s
Input
$0.06 / 1M tokens
Output
$0.144 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-32B-AWQ
Qwen3 32B INT8 GPTQText
qwen3-32b-gptq

INT8 GPTQ serving variant. Hybrid reasoning family with a switchable thinking mode.

reasoninghybridINT8
Context
128k
Speed
97 tok/s
Input
$0.079 / 1M tokens
Output
$0.19 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-32B-GPTQ
Qwen3 MoE 30B-A3BText
qwen3-moe-30b-a3b

Sparse Qwen3 — large total parameters, small active set per token.

MoE
Context
128k
Speed
224 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen3-30B-A3B
Qwen3 MoE 235B-A22BText
qwen3-moe-235b-a22b

Sparse Qwen3 — large total parameters, small active set per token.

MoE
Context
128k
Speed
77 tok/s
Input
$0.115 / 1M tokens
Output
$0.276 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
Qwen/Qwen3-235B-A22B
Qwen3 MoE 235B-A22B FP8Text
qwen3-moe-235b-a22b-fp8

FP8 serving variant. Sparse Qwen3 — large total parameters, small active set per token.

MoEFP8
Context
128k
Speed
108 tok/s
Input
$0.071 / 1M tokens
Output
$0.171 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
Qwen/Qwen3-235B-A22B-FP8
Qwen3 MoE 235B-A22B INT4 AWQText
qwen3-moe-235b-a22b-awq

INT4 AWQ serving variant. Sparse Qwen3 — large total parameters, small active set per token.

MoEINT4
Context
128k
Speed
145 tok/s
Input
$0.044 / 1M tokens
Output
$0.105 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
Qwen/Qwen3-235B-A22B-AWQ
Qwen3 MoE 235B-A22B INT8 GPTQText
qwen3-moe-235b-a22b-gptq

INT8 GPTQ serving variant. Sparse Qwen3 — large total parameters, small active set per token.

MoEINT8
Context
128k
Speed
124 tok/s
Input
$0.058 / 1M tokens
Output
$0.138 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
Qwen/Qwen3-235B-A22B-GPTQ
Qwen2.5 Coder 0.5BCode
qwen2-5-coder-0-5b

Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIM
Context
128k
Speed
298 tok/s
Input
$0.022 / 1M tokens
Output
$0.053 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-0.5B-Instruct
Qwen2.5 Coder 1.5BCode
qwen2-5-coder-1-5b

Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIM
Context
128k
Speed
263 tok/s
Input
$0.026 / 1M tokens
Output
$0.062 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-1.5B-Instruct
Qwen2.5 Coder 3BCode
qwen2-5-coder-3b

Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIM
Context
128k
Speed
224 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-3B-Instruct
Qwen2.5 Coder 7BCode
qwen2-5-coder-7b

Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIM
Context
128k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-7B-Instruct
Qwen2.5 Coder 7B FP8Code
qwen2-5-coder-7b-fp8

FP8 serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIMFP8
Context
128k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-7B-Instruct-FP8
Qwen2.5 Coder 7B INT4 AWQCode
qwen2-5-coder-7b-awq

INT4 AWQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIMINT4
Context
128k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-7B-Instruct-AWQ
Qwen2.5 Coder 7B INT8 GPTQCode
qwen2-5-coder-7b-gptq

INT8 GPTQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIMINT8
Context
128k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-7B-Instruct-GPTQ
Qwen2.5 Coder 14BCode
qwen2-5-coder-14b

Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIM
Context
128k
Speed
106 tok/s
Input
$0.08 / 1M tokens
Output
$0.192 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-14B-Instruct
Qwen2.5 Coder 14B FP8Code
qwen2-5-coder-14b-fp8

FP8 serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIMFP8
Context
128k
Speed
142 tok/s
Input
$0.05 / 1M tokens
Output
$0.119 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-14B-Instruct-FP8
Qwen2.5 Coder 14B INT4 AWQCode
qwen2-5-coder-14b-awq

INT4 AWQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIMINT4
Context
128k
Speed
181 tok/s
Input
$0.03 / 1M tokens
Output
$0.073 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-14B-Instruct-AWQ
Qwen2.5 Coder 14B INT8 GPTQCode
qwen2-5-coder-14b-gptq

INT8 GPTQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIMINT8
Context
128k
Speed
160 tok/s
Input
$0.04 / 1M tokens
Output
$0.096 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-14B-Instruct-GPTQ
Qwen2.5 Coder 32BCode
qwen2-5-coder-32b

Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIM
Context
128k
Speed
57 tok/s
Input
$0.158 / 1M tokens
Output
$0.379 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-32B-Instruct
Qwen2.5 Coder 32B FP8Code
qwen2-5-coder-32b-fp8

FP8 serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIMFP8
Context
128k
Speed
83 tok/s
Input
$0.098 / 1M tokens
Output
$0.235 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-32B-Instruct-FP8
Qwen2.5 Coder 32B INT4 AWQCode
qwen2-5-coder-32b-awq

INT4 AWQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIMINT4
Context
128k
Speed
116 tok/s
Input
$0.06 / 1M tokens
Output
$0.144 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-32B-Instruct-AWQ
Qwen2.5 Coder 32B INT8 GPTQCode
qwen2-5-coder-32b-gptq

INT8 GPTQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.

codeFIMINT8
Context
128k
Speed
97 tok/s
Input
$0.079 / 1M tokens
Output
$0.19 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Coder-32B-Instruct-GPTQ
Qwen2.5 Math 1.5BReasoning
qwen2-5-math-1-5b

Mathematical reasoning with chain-of-thought and tool-integrated solving.

math
Context
4k
Speed
263 tok/s
Input
$0.026 / 1M tokens
Output
$0.062 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Math-1.5B-Instruct
Qwen2.5 Math 7BReasoning
qwen2-5-math-7b

Mathematical reasoning with chain-of-thought and tool-integrated solving.

math
Context
4k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Math-7B-Instruct
Qwen2.5 Math 7B FP8Reasoning
qwen2-5-math-7b-fp8

FP8 serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.

mathFP8
Context
4k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Math-7B-Instruct-FP8
Qwen2.5 Math 7B INT4 AWQReasoning
qwen2-5-math-7b-awq

INT4 AWQ serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.

mathINT4
Context
4k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Math-7B-Instruct-AWQ
Qwen2.5 Math 7B INT8 GPTQReasoning
qwen2-5-math-7b-gptq

INT8 GPTQ serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.

mathINT8
Context
4k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-Math-7B-Instruct-GPTQ
Qwen2.5 Math 72BReasoning
qwen2-5-math-72b

Mathematical reasoning with chain-of-thought and tool-integrated solving.

math
Context
4k
Speed
28 tok/s
Input
$0.33 / 1M tokens
Output
$0.792 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-Math-72B-Instruct
Qwen2.5 Math 72B FP8Reasoning
qwen2-5-math-72b-fp8

FP8 serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.

mathFP8
Context
4k
Speed
43 tok/s
Input
$0.205 / 1M tokens
Output
$0.491 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-Math-72B-Instruct-FP8
Qwen2.5 Math 72B INT4 AWQReasoning
qwen2-5-math-72b-awq

INT4 AWQ serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.

mathINT4
Context
4k
Speed
65 tok/s
Input
$0.125 / 1M tokens
Output
$0.301 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-Math-72B-Instruct-AWQ
Qwen2.5 Math 72B INT8 GPTQReasoning
qwen2-5-math-72b-gptq

INT8 GPTQ serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.

mathINT8
Context
4k
Speed
52 tok/s
Input
$0.165 / 1M tokens
Output
$0.396 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-Math-72B-Instruct-GPTQ
Qwen2.5 VL 3BVision
qwen2-5-vl-3b

Vision-language Qwen with document parsing and grounding.

visionOCR
Context
128k
Speed
224 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-VL-3B-Instruct
Qwen2.5 VL 7BVision
qwen2-5-vl-7b

Vision-language Qwen with document parsing and grounding.

visionOCR
Context
128k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-VL-7B-Instruct
Qwen2.5 VL 7B FP8Vision
qwen2-5-vl-7b-fp8

FP8 serving variant. Vision-language Qwen with document parsing and grounding.

visionOCRFP8
Context
128k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-VL-7B-Instruct-FP8
Qwen2.5 VL 7B INT4 AWQVision
qwen2-5-vl-7b-awq

INT4 AWQ serving variant. Vision-language Qwen with document parsing and grounding.

visionOCRINT4
Context
128k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-VL-7B-Instruct-AWQ
Qwen2.5 VL 7B INT8 GPTQVision
qwen2-5-vl-7b-gptq

INT8 GPTQ serving variant. Vision-language Qwen with document parsing and grounding.

visionOCRINT8
Context
128k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-VL-7B-Instruct-GPTQ
Qwen2.5 VL 32BVision
qwen2-5-vl-32b

Vision-language Qwen with document parsing and grounding.

visionOCR
Context
128k
Speed
57 tok/s
Input
$0.158 / 1M tokens
Output
$0.379 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-VL-32B-Instruct
Qwen2.5 VL 32B FP8Vision
qwen2-5-vl-32b-fp8

FP8 serving variant. Vision-language Qwen with document parsing and grounding.

visionOCRFP8
Context
128k
Speed
83 tok/s
Input
$0.098 / 1M tokens
Output
$0.235 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-VL-32B-Instruct-FP8
Qwen2.5 VL 32B INT4 AWQVision
qwen2-5-vl-32b-awq

INT4 AWQ serving variant. Vision-language Qwen with document parsing and grounding.

visionOCRINT4
Context
128k
Speed
116 tok/s
Input
$0.06 / 1M tokens
Output
$0.144 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-VL-32B-Instruct-AWQ
Qwen2.5 VL 32B INT8 GPTQVision
qwen2-5-vl-32b-gptq

INT8 GPTQ serving variant. Vision-language Qwen with document parsing and grounding.

visionOCRINT8
Context
128k
Speed
97 tok/s
Input
$0.079 / 1M tokens
Output
$0.19 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/Qwen2.5-VL-32B-Instruct-GPTQ
Qwen2.5 VL 72BVision
qwen2-5-vl-72b

Vision-language Qwen with document parsing and grounding.

visionOCR
Context
128k
Speed
28 tok/s
Input
$0.33 / 1M tokens
Output
$0.792 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-VL-72B-Instruct
Qwen2.5 VL 72B FP8Vision
qwen2-5-vl-72b-fp8

FP8 serving variant. Vision-language Qwen with document parsing and grounding.

visionOCRFP8
Context
128k
Speed
43 tok/s
Input
$0.205 / 1M tokens
Output
$0.491 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-VL-72B-Instruct-FP8
Qwen2.5 VL 72B INT4 AWQVision
qwen2-5-vl-72b-awq

INT4 AWQ serving variant. Vision-language Qwen with document parsing and grounding.

visionOCRINT4
Context
128k
Speed
65 tok/s
Input
$0.125 / 1M tokens
Output
$0.301 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-VL-72B-Instruct-AWQ
Qwen2.5 VL 72B INT8 GPTQVision
qwen2-5-vl-72b-gptq

INT8 GPTQ serving variant. Vision-language Qwen with document parsing and grounding.

visionOCRINT8
Context
128k
Speed
52 tok/s
Input
$0.165 / 1M tokens
Output
$0.396 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
Qwen/Qwen2.5-VL-72B-Instruct-GPTQ
DeepSeek R1 Distill Qwen-1.5BReasoning
deepseek-r1-distill-qwen-1-5b

Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistill
Context
128k
Speed
263 tok/s
Input
$0.026 / 1M tokens
Output
$0.062 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
DeepSeek R1 Distill Qwen-7BReasoning
deepseek-r1-distill-qwen-7b

Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistill
Context
128k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
DeepSeek R1 Distill Qwen-7B FP8Reasoning
deepseek-r1-distill-qwen-7b-fp8

FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillFP8
Context
128k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B-FP8
DeepSeek R1 Distill Qwen-7B INT4 AWQReasoning
deepseek-r1-distill-qwen-7b-awq

INT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT4
Context
128k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B-AWQ
DeepSeek R1 Distill Qwen-7B INT8 GPTQReasoning
deepseek-r1-distill-qwen-7b-gptq

INT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT8
Context
128k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B-GPTQ
DeepSeek R1 Distill Qwen-14BReasoning
deepseek-r1-distill-qwen-14b

Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistill
Context
128k
Speed
106 tok/s
Input
$0.08 / 1M tokens
Output
$0.192 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
DeepSeek R1 Distill Qwen-14B FP8Reasoning
deepseek-r1-distill-qwen-14b-fp8

FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillFP8
Context
128k
Speed
142 tok/s
Input
$0.05 / 1M tokens
Output
$0.119 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B-FP8
DeepSeek R1 Distill Qwen-14B INT4 AWQReasoning
deepseek-r1-distill-qwen-14b-awq

INT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT4
Context
128k
Speed
181 tok/s
Input
$0.03 / 1M tokens
Output
$0.073 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B-AWQ
DeepSeek R1 Distill Qwen-14B INT8 GPTQReasoning
deepseek-r1-distill-qwen-14b-gptq

INT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT8
Context
128k
Speed
160 tok/s
Input
$0.04 / 1M tokens
Output
$0.096 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B-GPTQ
DeepSeek R1 Distill Qwen-32BReasoning
deepseek-r1-distill-qwen-32b

Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistill
Context
128k
Speed
57 tok/s
Input
$0.158 / 1M tokens
Output
$0.379 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
DeepSeek R1 Distill Qwen-32B FP8Reasoning
deepseek-r1-distill-qwen-32b-fp8

FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillFP8
Context
128k
Speed
83 tok/s
Input
$0.098 / 1M tokens
Output
$0.235 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B-FP8
DeepSeek R1 Distill Qwen-32B INT4 AWQReasoning
deepseek-r1-distill-qwen-32b-awq

INT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT4
Context
128k
Speed
116 tok/s
Input
$0.06 / 1M tokens
Output
$0.144 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B-AWQ
DeepSeek R1 Distill Qwen-32B INT8 GPTQReasoning
deepseek-r1-distill-qwen-32b-gptq

INT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT8
Context
128k
Speed
97 tok/s
Input
$0.079 / 1M tokens
Output
$0.19 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B-GPTQ
DeepSeek R1 Distill Llama-8BReasoning
deepseek-r1-distill-llama-8b

Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistill
Context
128k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Llama-8B
DeepSeek R1 Distill Llama-8B FP8Reasoning
deepseek-r1-distill-llama-8b-fp8

FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillFP8
Context
128k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Llama-8B-FP8
DeepSeek R1 Distill Llama-8B INT4 AWQReasoning
deepseek-r1-distill-llama-8b-awq

INT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT4
Context
128k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Llama-8B-AWQ
DeepSeek R1 Distill Llama-8B INT8 GPTQReasoning
deepseek-r1-distill-llama-8b-gptq

INT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT8
Context
128k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
deepseek-ai/DeepSeek-R1-Distill-Llama-8B-GPTQ
DeepSeek R1 Distill Llama-70BReasoning
deepseek-r1-distill-llama-70b

Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistill
Context
128k
Speed
29 tok/s
Input
$0.321 / 1M tokens
Output
$0.77 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
deepseek-ai/DeepSeek-R1-Distill-Llama-70B
DeepSeek R1 Distill Llama-70B FP8Reasoning
deepseek-r1-distill-llama-70b-fp8

FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillFP8
Context
128k
Speed
44 tok/s
Input
$0.199 / 1M tokens
Output
$0.478 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
deepseek-ai/DeepSeek-R1-Distill-Llama-70B-FP8
DeepSeek R1 Distill Llama-70B INT4 AWQReasoning
deepseek-r1-distill-llama-70b-awq

INT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT4
Context
128k
Speed
66 tok/s
Input
$0.122 / 1M tokens
Output
$0.293 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
deepseek-ai/DeepSeek-R1-Distill-Llama-70B-AWQ
DeepSeek R1 Distill Llama-70B INT8 GPTQReasoning
deepseek-r1-distill-llama-70b-gptq

INT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.

reasoningdistillINT8
Context
128k
Speed
53 tok/s
Input
$0.161 / 1M tokens
Output
$0.385 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
deepseek-ai/DeepSeek-R1-Distill-Llama-70B-GPTQ
DeepSeek Coder V2 LiteCode
deepseek-coder-v2-lite

Repository-scale code MoE with strong completion and repair.

codeMoE
Context
128k
Speed
238 tok/s
Input
$0.03 / 1M tokens
Output
$0.072 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct
DeepSeek Coder V2 236BCode
deepseek-coder-v2-236b

Repository-scale code MoE with strong completion and repair.

codeMoE
Context
128k
Speed
80 tok/s
Input
$0.11 / 1M tokens
Output
$0.264 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
deepseek-ai/DeepSeek-Coder-V2-Instruct
DeepSeek Coder V2 236B FP8Code
deepseek-coder-v2-236b-fp8

FP8 serving variant. Repository-scale code MoE with strong completion and repair.

codeMoEFP8
Context
128k
Speed
111 tok/s
Input
$0.068 / 1M tokens
Output
$0.164 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
deepseek-ai/DeepSeek-Coder-V2-Instruct-FP8
DeepSeek Coder V2 236B INT4 AWQCode
deepseek-coder-v2-236b-awq

INT4 AWQ serving variant. Repository-scale code MoE with strong completion and repair.

codeMoEINT4
Context
128k
Speed
149 tok/s
Input
$0.042 / 1M tokens
Output
$0.1 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
deepseek-ai/DeepSeek-Coder-V2-Instruct-AWQ
DeepSeek Coder V2 236B INT8 GPTQCode
deepseek-coder-v2-236b-gptq

INT8 GPTQ serving variant. Repository-scale code MoE with strong completion and repair.

codeMoEINT8
Context
128k
Speed
128 tok/s
Input
$0.055 / 1M tokens
Output
$0.132 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
deepseek-ai/DeepSeek-Coder-V2-Instruct-GPTQ
Gemma 2 2bText
gemma-2-2b

Google's open dense family. Efficient at small and mid scale.

general
Context
8k
Speed
248 tok/s
Input
$0.029 / 1M tokens
Output
$0.07 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-2-2b-it
Gemma 2 9bText
gemma-2-9b

Google's open dense family. Efficient at small and mid scale.

general
Context
8k
Speed
140 tok/s
Input
$0.059 / 1M tokens
Output
$0.142 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-2-9b-it
Gemma 2 9b FP8Text
gemma-2-9b-fp8

FP8 serving variant. Google's open dense family. Efficient at small and mid scale.

generalFP8
Context
8k
Speed
178 tok/s
Input
$0.037 / 1M tokens
Output
$0.088 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-2-9b-it-FP8
Gemma 2 9b INT4 AWQText
gemma-2-9b-awq

INT4 AWQ serving variant. Google's open dense family. Efficient at small and mid scale.

generalINT4
Context
8k
Speed
214 tok/s
Input
$0.022 / 1M tokens
Output
$0.054 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-2-9b-it-AWQ
Gemma 2 9b INT8 GPTQText
gemma-2-9b-gptq

INT8 GPTQ serving variant. Google's open dense family. Efficient at small and mid scale.

generalINT8
Context
8k
Speed
194 tok/s
Input
$0.029 / 1M tokens
Output
$0.071 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-2-9b-it-GPTQ
Gemma 2 27b FP8Text
gemma-2-27b-fp8

FP8 serving variant. Google's open dense family. Efficient at small and mid scale.

generalFP8
Context
8k
Speed
94 tok/s
Input
$0.084 / 1M tokens
Output
$0.202 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-2-27b-it-FP8
Gemma 2 27b INT4 AWQText
gemma-2-27b-awq

INT4 AWQ serving variant. Google's open dense family. Efficient at small and mid scale.

generalINT4
Context
8k
Speed
129 tok/s
Input
$0.052 / 1M tokens
Output
$0.124 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-2-27b-it-AWQ
Gemma 2 27b INT8 GPTQText
gemma-2-27b-gptq

INT8 GPTQ serving variant. Google's open dense family. Efficient at small and mid scale.

generalINT8
Context
8k
Speed
109 tok/s
Input
$0.068 / 1M tokens
Output
$0.163 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-2-27b-it-GPTQ
Gemma 3 1bText
gemma-3-1b

Multimodal Gemma with a long context window and wide language coverage.

multimodal
Context
128k
Speed
280 tok/s
Input
$0.024 / 1M tokens
Output
$0.058 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-1b-it
Gemma 3 4bText
gemma-3-4b

Multimodal Gemma with a long context window and wide language coverage.

multimodal
Context
128k
Speed
203 tok/s
Input
$0.037 / 1M tokens
Output
$0.089 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-4b-it
Gemma 3 12bText
gemma-3-12b

Multimodal Gemma with a long context window and wide language coverage.

multimodal
Context
128k
Speed
117 tok/s
Input
$0.072 / 1M tokens
Output
$0.173 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-12b-it
Gemma 3 12b FP8Text
gemma-3-12b-fp8

FP8 serving variant. Multimodal Gemma with a long context window and wide language coverage.

multimodalFP8
Context
128k
Speed
155 tok/s
Input
$0.045 / 1M tokens
Output
$0.107 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-12b-it-FP8
Gemma 3 12b INT4 AWQText
gemma-3-12b-awq

INT4 AWQ serving variant. Multimodal Gemma with a long context window and wide language coverage.

multimodalINT4
Context
128k
Speed
193 tok/s
Input
$0.027 / 1M tokens
Output
$0.066 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-12b-it-AWQ
Gemma 3 12b INT8 GPTQText
gemma-3-12b-gptq

INT8 GPTQ serving variant. Multimodal Gemma with a long context window and wide language coverage.

multimodalINT8
Context
128k
Speed
172 tok/s
Input
$0.036 / 1M tokens
Output
$0.086 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-12b-it-GPTQ
Gemma 3 27bText
gemma-3-27b

Multimodal Gemma with a long context window and wide language coverage.

multimodal
Context
128k
Speed
65 tok/s
Input
$0.136 / 1M tokens
Output
$0.326 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-27b-it
Gemma 3 27b FP8Text
gemma-3-27b-fp8

FP8 serving variant. Multimodal Gemma with a long context window and wide language coverage.

multimodalFP8
Context
128k
Speed
94 tok/s
Input
$0.084 / 1M tokens
Output
$0.202 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-27b-it-FP8
Gemma 3 27b INT4 AWQText
gemma-3-27b-awq

INT4 AWQ serving variant. Multimodal Gemma with a long context window and wide language coverage.

multimodalINT4
Context
128k
Speed
129 tok/s
Input
$0.052 / 1M tokens
Output
$0.124 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-27b-it-AWQ
Gemma 3 27b INT8 GPTQText
gemma-3-27b-gptq

INT8 GPTQ serving variant. Multimodal Gemma with a long context window and wide language coverage.

multimodalINT8
Context
128k
Speed
109 tok/s
Input
$0.068 / 1M tokens
Output
$0.163 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/gemma-3-27b-it-GPTQ
Mistral 7B v0.3Text
mistral-7b-v0-3

Compact, permissively licensed dense models.

Apache-2.0
Context
32k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-7B-Instruct-v0.3
Mistral 7B v0.3 FP8Text
mistral-7b-v0-3-fp8

FP8 serving variant. Compact, permissively licensed dense models.

Apache-2.0FP8
Context
32k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-7B-Instruct-v0.3-FP8
Mistral 7B v0.3 INT4 AWQText
mistral-7b-v0-3-awq

INT4 AWQ serving variant. Compact, permissively licensed dense models.

Apache-2.0INT4
Context
32k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-7B-Instruct-v0.3-AWQ
Mistral 7B v0.3 INT8 GPTQText
mistral-7b-v0-3-gptq

INT8 GPTQ serving variant. Compact, permissively licensed dense models.

Apache-2.0INT8
Context
32k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-7B-Instruct-v0.3-GPTQ
Mistral Nemo 12BText
mistral-nemo-12b

Compact, permissively licensed dense models.

Apache-2.0
Context
32k
Speed
117 tok/s
Input
$0.072 / 1M tokens
Output
$0.173 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-Nemo-Instruct-2407
Mistral Nemo 12B FP8Text
mistral-nemo-12b-fp8

FP8 serving variant. Compact, permissively licensed dense models.

Apache-2.0FP8
Context
32k
Speed
155 tok/s
Input
$0.045 / 1M tokens
Output
$0.107 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-Nemo-Instruct-2407-FP8
Mistral Nemo 12B INT4 AWQText
mistral-nemo-12b-awq

INT4 AWQ serving variant. Compact, permissively licensed dense models.

Apache-2.0INT4
Context
32k
Speed
193 tok/s
Input
$0.027 / 1M tokens
Output
$0.066 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-Nemo-Instruct-2407-AWQ
Mistral Nemo 12B INT8 GPTQText
mistral-nemo-12b-gptq

INT8 GPTQ serving variant. Compact, permissively licensed dense models.

Apache-2.0INT8
Context
32k
Speed
172 tok/s
Input
$0.036 / 1M tokens
Output
$0.086 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-Nemo-Instruct-2407-GPTQ
Mistral Small 24BText
mistral-small-24b

Compact, permissively licensed dense models.

Apache-2.0
Context
32k
Speed
72 tok/s
Input
$0.123 / 1M tokens
Output
$0.295 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-Small-24B-Instruct-2501
Mistral Small 24B FP8Text
mistral-small-24b-fp8

FP8 serving variant. Compact, permissively licensed dense models.

Apache-2.0FP8
Context
32k
Speed
102 tok/s
Input
$0.076 / 1M tokens
Output
$0.183 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-Small-24B-Instruct-2501-FP8
Mistral Small 24B INT4 AWQText
mistral-small-24b-awq

INT4 AWQ serving variant. Compact, permissively licensed dense models.

Apache-2.0INT4
Context
32k
Speed
138 tok/s
Input
$0.047 / 1M tokens
Output
$0.112 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-Small-24B-Instruct-2501-AWQ
Mistral Small 24B INT8 GPTQText
mistral-small-24b-gptq

INT8 GPTQ serving variant. Compact, permissively licensed dense models.

Apache-2.0INT8
Context
32k
Speed
117 tok/s
Input
$0.061 / 1M tokens
Output
$0.148 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Mistral-Small-24B-Instruct-2501-GPTQ
Mistral Large 123BText
mistral-large-123b

Compact, permissively licensed dense models.

Apache-2.0
Context
32k
Speed
17 tok/s
Input
$0.549 / 1M tokens
Output
$1.318 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
mistralai/Mistral-Large-Instruct-2407
Mistral Large 123B FP8Text
mistral-large-123b-fp8

FP8 serving variant. Compact, permissively licensed dense models.

Apache-2.0FP8
Context
32k
Speed
26 tok/s
Input
$0.34 / 1M tokens
Output
$0.817 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
mistralai/Mistral-Large-Instruct-2407-FP8
Mistral Large 123B INT4 AWQText
mistral-large-123b-awq

INT4 AWQ serving variant. Compact, permissively licensed dense models.

Apache-2.0INT4
Context
32k
Speed
41 tok/s
Input
$0.209 / 1M tokens
Output
$0.501 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
mistralai/Mistral-Large-Instruct-2407-AWQ
Mistral Large 123B INT8 GPTQText
mistral-large-123b-gptq

INT8 GPTQ serving variant. Compact, permissively licensed dense models.

Apache-2.0INT8
Context
32k
Speed
32 tok/s
Input
$0.275 / 1M tokens
Output
$0.659 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
mistralai/Mistral-Large-Instruct-2407-GPTQ
Ministral 3BText
ministral-3b

Edge-class Mistral for on-device and high-throughput serving.

smalledge
Context
128k
Speed
224 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Ministral-3B-Instruct
Ministral 8BText
ministral-8b

Edge-class Mistral for on-device and high-throughput serving.

smalledge
Context
128k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Ministral-8B-Instruct-2410
Ministral 8B FP8Text
ministral-8b-fp8

FP8 serving variant. Edge-class Mistral for on-device and high-throughput serving.

smalledgeFP8
Context
128k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Ministral-8B-Instruct-2410-FP8
Ministral 8B INT4 AWQText
ministral-8b-awq

INT4 AWQ serving variant. Edge-class Mistral for on-device and high-throughput serving.

smalledgeINT4
Context
128k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Ministral-8B-Instruct-2410-AWQ
Ministral 8B INT8 GPTQText
ministral-8b-gptq

INT8 GPTQ serving variant. Edge-class Mistral for on-device and high-throughput serving.

smalledgeINT8
Context
128k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Ministral-8B-Instruct-2410-GPTQ
Codestral 22BCode
codestral-22b

Mistral's code model with fill-in-the-middle across 80+ languages.

codeFIM
Context
32k
Speed
77 tok/s
Input
$0.115 / 1M tokens
Output
$0.276 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Codestral-22B-v0.1
Codestral 22B FP8Code
codestral-22b-fp8

FP8 serving variant. Mistral's code model with fill-in-the-middle across 80+ languages.

codeFIMFP8
Context
32k
Speed
108 tok/s
Input
$0.071 / 1M tokens
Output
$0.171 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Codestral-22B-v0.1-FP8
Codestral 22B INT4 AWQCode
codestral-22b-awq

INT4 AWQ serving variant. Mistral's code model with fill-in-the-middle across 80+ languages.

codeFIMINT4
Context
32k
Speed
145 tok/s
Input
$0.044 / 1M tokens
Output
$0.105 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Codestral-22B-v0.1-AWQ
Codestral 22B INT8 GPTQCode
codestral-22b-gptq

INT8 GPTQ serving variant. Mistral's code model with fill-in-the-middle across 80+ languages.

codeFIMINT8
Context
32k
Speed
124 tok/s
Input
$0.058 / 1M tokens
Output
$0.138 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Codestral-22B-v0.1-GPTQ
Phi-3 miniText
phi-3-mini

Small models trained on heavily filtered data. Strong reasoning per parameter.

smallcheap
Context
128k
Speed
207 tok/s
Input
$0.036 / 1M tokens
Output
$0.086 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3-mini-128k-instruct
Phi-3 smallText
phi-3-small

Small models trained on heavily filtered data. Strong reasoning per parameter.

smallcheap
Context
128k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3-small-128k-instruct
Phi-3 small FP8Text
phi-3-small-fp8

FP8 serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.

smallcheapFP8
Context
128k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3-small-128k-instruct-FP8
Phi-3 small INT4 AWQText
phi-3-small-awq

INT4 AWQ serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.

smallcheapINT4
Context
128k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3-small-128k-instruct-AWQ
Phi-3 small INT8 GPTQText
phi-3-small-gptq

INT8 GPTQ serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.

smallcheapINT8
Context
128k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3-small-128k-instruct-GPTQ
Phi-3 mediumText
phi-3-medium

Small models trained on heavily filtered data. Strong reasoning per parameter.

smallcheap
Context
128k
Speed
106 tok/s
Input
$0.08 / 1M tokens
Output
$0.192 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3-medium-128k-instruct
Phi-3 medium FP8Text
phi-3-medium-fp8

FP8 serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.

smallcheapFP8
Context
128k
Speed
142 tok/s
Input
$0.05 / 1M tokens
Output
$0.119 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3-medium-128k-instruct-FP8
Phi-3 medium INT4 AWQText
phi-3-medium-awq

INT4 AWQ serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.

smallcheapINT4
Context
128k
Speed
181 tok/s
Input
$0.03 / 1M tokens
Output
$0.073 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3-medium-128k-instruct-AWQ
Phi-3 medium INT8 GPTQText
phi-3-medium-gptq

INT8 GPTQ serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.

smallcheapINT8
Context
128k
Speed
160 tok/s
Input
$0.04 / 1M tokens
Output
$0.096 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3-medium-128k-instruct-GPTQ
Phi-3.5 miniText
phi-3-5-mini

Refreshed Phi with a sparse MoE variant and vision sibling.

small
Context
128k
Speed
207 tok/s
Input
$0.036 / 1M tokens
Output
$0.086 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3.5-mini-instruct
Phi-3.5 MoEText
phi-3-5-moe

Refreshed Phi with a sparse MoE variant and vision sibling.

small
Context
128k
Speed
164 tok/s
Input
$0.048 / 1M tokens
Output
$0.115 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3.5-MoE-instruct
Phi-3.5 visionText
phi-3-5-vision

Refreshed Phi with a sparse MoE variant and vision sibling.

small
Context
128k
Speed
200 tok/s
Input
$0.038 / 1M tokens
Output
$0.091 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-3.5-vision-instruct
Yi 1.5 6BText
yi-1-5-6b

Bilingual Chinese/English dense family.

multilingual
Context
32k
Speed
172 tok/s
Input
$0.046 / 1M tokens
Output
$0.11 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
01-ai/Yi-1.5-6B-Chat
Yi 1.5 9BText
yi-1-5-9b

Bilingual Chinese/English dense family.

multilingual
Context
32k
Speed
140 tok/s
Input
$0.059 / 1M tokens
Output
$0.142 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
01-ai/Yi-1.5-9B-Chat
Yi 1.5 9B FP8Text
yi-1-5-9b-fp8

FP8 serving variant. Bilingual Chinese/English dense family.

multilingualFP8
Context
32k
Speed
178 tok/s
Input
$0.037 / 1M tokens
Output
$0.088 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
01-ai/Yi-1.5-9B-Chat-FP8
Yi 1.5 9B INT4 AWQText
yi-1-5-9b-awq

INT4 AWQ serving variant. Bilingual Chinese/English dense family.

multilingualINT4
Context
32k
Speed
214 tok/s
Input
$0.022 / 1M tokens
Output
$0.054 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
01-ai/Yi-1.5-9B-Chat-AWQ
Yi 1.5 9B INT8 GPTQText
yi-1-5-9b-gptq

INT8 GPTQ serving variant. Bilingual Chinese/English dense family.

multilingualINT8
Context
32k
Speed
194 tok/s
Input
$0.029 / 1M tokens
Output
$0.071 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
01-ai/Yi-1.5-9B-Chat-GPTQ
Yi 1.5 34BText
yi-1-5-34b

Bilingual Chinese/English dense family.

multilingual
Context
32k
Speed
54 tok/s
Input
$0.166 / 1M tokens
Output
$0.398 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
01-ai/Yi-1.5-34B-Chat
Yi 1.5 34B FP8Text
yi-1-5-34b-fp8

FP8 serving variant. Bilingual Chinese/English dense family.

multilingualFP8
Context
32k
Speed
79 tok/s
Input
$0.103 / 1M tokens
Output
$0.247 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
01-ai/Yi-1.5-34B-Chat-FP8
Yi 1.5 34B INT4 AWQText
yi-1-5-34b-awq

INT4 AWQ serving variant. Bilingual Chinese/English dense family.

multilingualINT4
Context
32k
Speed
112 tok/s
Input
$0.063 / 1M tokens
Output
$0.151 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
01-ai/Yi-1.5-34B-Chat-AWQ
Yi 1.5 34B INT8 GPTQText
yi-1-5-34b-gptq

INT8 GPTQ serving variant. Bilingual Chinese/English dense family.

multilingualINT8
Context
32k
Speed
93 tok/s
Input
$0.083 / 1M tokens
Output
$0.199 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
01-ai/Yi-1.5-34B-Chat-GPTQ
GLM-4 9BText
glm-4-9b

Zhipu's bilingual family with agent tuning.

agentic
Context
128k
Speed
140 tok/s
Input
$0.059 / 1M tokens
Output
$0.142 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
THUDM/glm-4-9b-chat
GLM-4 9B FP8Text
glm-4-9b-fp8

FP8 serving variant. Zhipu's bilingual family with agent tuning.

agenticFP8
Context
128k
Speed
178 tok/s
Input
$0.037 / 1M tokens
Output
$0.088 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
THUDM/glm-4-9b-chat-FP8
GLM-4 9B INT4 AWQText
glm-4-9b-awq

INT4 AWQ serving variant. Zhipu's bilingual family with agent tuning.

agenticINT4
Context
128k
Speed
214 tok/s
Input
$0.022 / 1M tokens
Output
$0.054 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
THUDM/glm-4-9b-chat-AWQ
GLM-4 9B INT8 GPTQText
glm-4-9b-gptq

INT8 GPTQ serving variant. Zhipu's bilingual family with agent tuning.

agenticINT8
Context
128k
Speed
194 tok/s
Input
$0.029 / 1M tokens
Output
$0.071 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
THUDM/glm-4-9b-chat-GPTQ
GLM-4 9B 1MText
glm-4-9b-1m

Zhipu's bilingual family with agent tuning.

agentic
Context
128k
Speed
140 tok/s
Input
$0.059 / 1M tokens
Output
$0.142 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
THUDM/glm-4-9b-chat-1m
GLM-4 9B 1M FP8Text
glm-4-9b-1m-fp8

FP8 serving variant. Zhipu's bilingual family with agent tuning.

agenticFP8
Context
128k
Speed
178 tok/s
Input
$0.037 / 1M tokens
Output
$0.088 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
THUDM/glm-4-9b-chat-1m-FP8
GLM-4 9B 1M INT4 AWQText
glm-4-9b-1m-awq

INT4 AWQ serving variant. Zhipu's bilingual family with agent tuning.

agenticINT4
Context
128k
Speed
214 tok/s
Input
$0.022 / 1M tokens
Output
$0.054 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
THUDM/glm-4-9b-chat-1m-AWQ
GLM-4 9B 1M INT8 GPTQText
glm-4-9b-1m-gptq

INT8 GPTQ serving variant. Zhipu's bilingual family with agent tuning.

agenticINT8
Context
128k
Speed
194 tok/s
Input
$0.029 / 1M tokens
Output
$0.071 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
THUDM/glm-4-9b-chat-1m-GPTQ
InternLM2.5 1_8bText
internlm2-5-1-8b

Long-context Chinese/English models with tool-use training.

long-context
Context
1M
Speed
254 tok/s
Input
$0.028 / 1M tokens
Output
$0.067 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
internlm/internlm2_5-1_8b-chat
InternLM2.5 7bText
internlm2-5-7b

Long-context Chinese/English models with tool-use training.

long-context
Context
1M
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
internlm/internlm2_5-7b-chat
InternLM2.5 7b FP8Text
internlm2-5-7b-fp8

FP8 serving variant. Long-context Chinese/English models with tool-use training.

long-contextFP8
Context
1M
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
internlm/internlm2_5-7b-chat-FP8
InternLM2.5 7b INT4 AWQText
internlm2-5-7b-awq

INT4 AWQ serving variant. Long-context Chinese/English models with tool-use training.

long-contextINT4
Context
1M
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
internlm/internlm2_5-7b-chat-AWQ
InternLM2.5 7b INT8 GPTQText
internlm2-5-7b-gptq

INT8 GPTQ serving variant. Long-context Chinese/English models with tool-use training.

long-contextINT8
Context
1M
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
internlm/internlm2_5-7b-chat-GPTQ
InternLM2.5 20bText
internlm2-5-20b

Long-context Chinese/English models with tool-use training.

long-context
Context
1M
Speed
82 tok/s
Input
$0.106 / 1M tokens
Output
$0.254 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
internlm/internlm2_5-20b-chat
InternLM2.5 20b FP8Text
internlm2-5-20b-fp8

FP8 serving variant. Long-context Chinese/English models with tool-use training.

long-contextFP8
Context
1M
Speed
115 tok/s
Input
$0.066 / 1M tokens
Output
$0.158 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
internlm/internlm2_5-20b-chat-FP8
InternLM2.5 20b INT4 AWQText
internlm2-5-20b-awq

INT4 AWQ serving variant. Long-context Chinese/English models with tool-use training.

long-contextINT4
Context
1M
Speed
153 tok/s
Input
$0.04 / 1M tokens
Output
$0.097 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
internlm/internlm2_5-20b-chat-AWQ
InternLM2.5 20b INT8 GPTQText
internlm2-5-20b-gptq

INT8 GPTQ serving variant. Long-context Chinese/English models with tool-use training.

long-contextINT8
Context
1M
Speed
131 tok/s
Input
$0.053 / 1M tokens
Output
$0.127 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
internlm/internlm2_5-20b-chat-GPTQ
MiniCPM3 4BText
minicpm3-4b

Small models tuned for on-device deployment.

edge
Context
32k
Speed
203 tok/s
Input
$0.037 / 1M tokens
Output
$0.089 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
openbmb/MiniCPM3-4B
OLMo 2 7BText
olmo-2-7b

Fully open training data, code and checkpoints. The reproducibility baseline.

fully-open
Context
4k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/OLMo-2-1124-7B-Instruct
OLMo 2 7B FP8Text
olmo-2-7b-fp8

FP8 serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.

fully-openFP8
Context
4k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/OLMo-2-1124-7B-Instruct-FP8
OLMo 2 7B INT4 AWQText
olmo-2-7b-awq

INT4 AWQ serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.

fully-openINT4
Context
4k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/OLMo-2-1124-7B-Instruct-AWQ
OLMo 2 7B INT8 GPTQText
olmo-2-7b-gptq

INT8 GPTQ serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.

fully-openINT8
Context
4k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/OLMo-2-1124-7B-Instruct-GPTQ
OLMo 2 13BText
olmo-2-13b

Fully open training data, code and checkpoints. The reproducibility baseline.

fully-open
Context
4k
Speed
112 tok/s
Input
$0.076 / 1M tokens
Output
$0.182 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/OLMo-2-1124-13B-Instruct
OLMo 2 13B FP8Text
olmo-2-13b-fp8

FP8 serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.

fully-openFP8
Context
4k
Speed
148 tok/s
Input
$0.047 / 1M tokens
Output
$0.113 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/OLMo-2-1124-13B-Instruct-FP8
OLMo 2 13B INT4 AWQText
olmo-2-13b-awq

INT4 AWQ serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.

fully-openINT4
Context
4k
Speed
187 tok/s
Input
$0.029 / 1M tokens
Output
$0.069 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/OLMo-2-1124-13B-Instruct-AWQ
OLMo 2 13B INT8 GPTQText
olmo-2-13b-gptq

INT8 GPTQ serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.

fully-openINT8
Context
4k
Speed
165 tok/s
Input
$0.038 / 1M tokens
Output
$0.091 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/OLMo-2-1124-13B-Instruct-GPTQ
Falcon 3 1BText
falcon-3-1b

TII's efficient dense family.

general
Context
32k
Speed
280 tok/s
Input
$0.024 / 1M tokens
Output
$0.058 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-1B-Instruct
Falcon 3 3BText
falcon-3-3b

TII's efficient dense family.

general
Context
32k
Speed
224 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-3B-Instruct
Falcon 3 7BText
falcon-3-7b

TII's efficient dense family.

general
Context
32k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-7B-Instruct
Falcon 3 7B FP8Text
falcon-3-7b-fp8

FP8 serving variant. TII's efficient dense family.

generalFP8
Context
32k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-7B-Instruct-FP8
Falcon 3 7B INT4 AWQText
falcon-3-7b-awq

INT4 AWQ serving variant. TII's efficient dense family.

generalINT4
Context
32k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-7B-Instruct-AWQ
Falcon 3 7B INT8 GPTQText
falcon-3-7b-gptq

INT8 GPTQ serving variant. TII's efficient dense family.

generalINT8
Context
32k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-7B-Instruct-GPTQ
Falcon 3 10BText
falcon-3-10b

TII's efficient dense family.

general
Context
32k
Speed
131 tok/s
Input
$0.063 / 1M tokens
Output
$0.151 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-10B-Instruct
Falcon 3 10B FP8Text
falcon-3-10b-fp8

FP8 serving variant. TII's efficient dense family.

generalFP8
Context
32k
Speed
169 tok/s
Input
$0.039 / 1M tokens
Output
$0.094 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-10B-Instruct-FP8
Falcon 3 10B INT4 AWQText
falcon-3-10b-awq

INT4 AWQ serving variant. TII's efficient dense family.

generalINT4
Context
32k
Speed
207 tok/s
Input
$0.024 / 1M tokens
Output
$0.057 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-10B-Instruct-AWQ
Falcon 3 10B INT8 GPTQText
falcon-3-10b-gptq

INT8 GPTQ serving variant. TII's efficient dense family.

generalINT8
Context
32k
Speed
186 tok/s
Input
$0.032 / 1M tokens
Output
$0.076 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
tiiuae/Falcon3-10B-Instruct-GPTQ
Command R R 35BText
command-r-r-35b

Retrieval-augmented generation and citation-grounded answering.

RAGtools
Context
128k
Speed
53 tok/s
Input
$0.17 / 1M tokens
Output
$0.408 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
CohereForAI/c4ai-command-r-v01
Command R R 35B FP8Text
command-r-r-35b-fp8

FP8 serving variant. Retrieval-augmented generation and citation-grounded answering.

RAGtoolsFP8
Context
128k
Speed
78 tok/s
Input
$0.105 / 1M tokens
Output
$0.253 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
CohereForAI/c4ai-command-r-v01-FP8
Command R R 35B INT4 AWQText
command-r-r-35b-awq

INT4 AWQ serving variant. Retrieval-augmented generation and citation-grounded answering.

RAGtoolsINT4
Context
128k
Speed
110 tok/s
Input
$0.065 / 1M tokens
Output
$0.155 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
CohereForAI/c4ai-command-r-v01-AWQ
Command R R 35B INT8 GPTQText
command-r-r-35b-gptq

INT8 GPTQ serving variant. Retrieval-augmented generation and citation-grounded answering.

RAGtoolsINT8
Context
128k
Speed
91 tok/s
Input
$0.085 / 1M tokens
Output
$0.204 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
CohereForAI/c4ai-command-r-v01-GPTQ
Command R R+ 104BText
command-r-r-104b

Retrieval-augmented generation and citation-grounded answering.

RAGtools
Context
128k
Speed
20 tok/s
Input
$0.467 / 1M tokens
Output
$1.121 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
CohereForAI/c4ai-command-r-plus
Command R R+ 104B FP8Text
command-r-r-104b-fp8

FP8 serving variant. Retrieval-augmented generation and citation-grounded answering.

RAGtoolsFP8
Context
128k
Speed
31 tok/s
Input
$0.29 / 1M tokens
Output
$0.695 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
CohereForAI/c4ai-command-r-plus-FP8
Command R R+ 104B INT4 AWQText
command-r-r-104b-awq

INT4 AWQ serving variant. Retrieval-augmented generation and citation-grounded answering.

RAGtoolsINT4
Context
128k
Speed
48 tok/s
Input
$0.177 / 1M tokens
Output
$0.426 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
CohereForAI/c4ai-command-r-plus-AWQ
Command R R+ 104B INT8 GPTQText
command-r-r-104b-gptq

INT8 GPTQ serving variant. Retrieval-augmented generation and citation-grounded answering.

RAGtoolsINT8
Context
128k
Speed
37 tok/s
Input
$0.234 / 1M tokens
Output
$0.56 / 1M tokens
Dedicated
4 × H100 · $8.56/hr
CohereForAI/c4ai-command-r-plus-GPTQ
Command R R7BText
command-r-r7b

Retrieval-augmented generation and citation-grounded answering.

RAGtools
Context
128k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/c4ai-command-r7b-12-2024
Command R R7B FP8Text
command-r-r7b-fp8

FP8 serving variant. Retrieval-augmented generation and citation-grounded answering.

RAGtoolsFP8
Context
128k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/c4ai-command-r7b-12-2024-FP8
Command R R7B INT4 AWQText
command-r-r7b-awq

INT4 AWQ serving variant. Retrieval-augmented generation and citation-grounded answering.

RAGtoolsINT4
Context
128k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/c4ai-command-r7b-12-2024-AWQ
Command R R7B INT8 GPTQText
command-r-r7b-gptq

INT8 GPTQ serving variant. Retrieval-augmented generation and citation-grounded answering.

RAGtoolsINT8
Context
128k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/c4ai-command-r7b-12-2024-GPTQ
Aya Expanse 8BText
aya-expanse-8b

Massively multilingual coverage across 23 languages.

multilingual
Context
128k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/aya-expanse-8b
Aya Expanse 8B FP8Text
aya-expanse-8b-fp8

FP8 serving variant. Massively multilingual coverage across 23 languages.

multilingualFP8
Context
128k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/aya-expanse-8b-FP8
Aya Expanse 8B INT4 AWQText
aya-expanse-8b-awq

INT4 AWQ serving variant. Massively multilingual coverage across 23 languages.

multilingualINT4
Context
128k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/aya-expanse-8b-AWQ
Aya Expanse 8B INT8 GPTQText
aya-expanse-8b-gptq

INT8 GPTQ serving variant. Massively multilingual coverage across 23 languages.

multilingualINT8
Context
128k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/aya-expanse-8b-GPTQ
Aya Expanse 32BText
aya-expanse-32b

Massively multilingual coverage across 23 languages.

multilingual
Context
128k
Speed
57 tok/s
Input
$0.158 / 1M tokens
Output
$0.379 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/aya-expanse-32b
Aya Expanse 32B FP8Text
aya-expanse-32b-fp8

FP8 serving variant. Massively multilingual coverage across 23 languages.

multilingualFP8
Context
128k
Speed
83 tok/s
Input
$0.098 / 1M tokens
Output
$0.235 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/aya-expanse-32b-FP8
Aya Expanse 32B INT4 AWQText
aya-expanse-32b-awq

INT4 AWQ serving variant. Massively multilingual coverage across 23 languages.

multilingualINT4
Context
128k
Speed
116 tok/s
Input
$0.06 / 1M tokens
Output
$0.144 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/aya-expanse-32b-AWQ
Aya Expanse 32B INT8 GPTQText
aya-expanse-32b-gptq

INT8 GPTQ serving variant. Massively multilingual coverage across 23 languages.

multilingualINT8
Context
128k
Speed
97 tok/s
Input
$0.079 / 1M tokens
Output
$0.19 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
CohereForAI/aya-expanse-32b-GPTQ
Granite 3 2bText
granite-3-2b

IBM's enterprise family with an explicit indemnified license.

enterprise
Context
128k
Speed
248 tok/s
Input
$0.029 / 1M tokens
Output
$0.07 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ibm-granite/granite-3.1-2b-instruct
Granite 3 8bText
granite-3-8b

IBM's enterprise family with an explicit indemnified license.

enterprise
Context
128k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ibm-granite/granite-3.1-8b-instruct
Granite 3 8b FP8Text
granite-3-8b-fp8

FP8 serving variant. IBM's enterprise family with an explicit indemnified license.

enterpriseFP8
Context
128k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ibm-granite/granite-3.1-8b-instruct-FP8
Granite 3 8b INT4 AWQText
granite-3-8b-awq

INT4 AWQ serving variant. IBM's enterprise family with an explicit indemnified license.

enterpriseINT4
Context
128k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ibm-granite/granite-3.1-8b-instruct-AWQ
Granite 3 8b INT8 GPTQText
granite-3-8b-gptq

INT8 GPTQ serving variant. IBM's enterprise family with an explicit indemnified license.

enterpriseINT8
Context
128k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ibm-granite/granite-3.1-8b-instruct-GPTQ
Granite 3 1b-a400mText
granite-3-1b-a400m

IBM's enterprise family with an explicit indemnified license.

enterprise
Context
128k
Speed
302 tok/s
Input
$0.022 / 1M tokens
Output
$0.053 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ibm-granite/granite-3.1-1b-a400m-instruct
Granite 3 3b-a800mText
granite-3-3b-a800m

IBM's enterprise family with an explicit indemnified license.

enterprise
Context
128k
Speed
287 tok/s
Input
$0.023 / 1M tokens
Output
$0.055 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ibm-granite/granite-3.1-3b-a800m-instruct
Nemotron 70BText
nemotron-70b

NVIDIA's reward-tuned Llama derivatives.

RLHF
Context
128k
Speed
29 tok/s
Input
$0.321 / 1M tokens
Output
$0.77 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
nvidia/Llama-3.1-Nemotron-70B-Instruct-HF
Nemotron 70B FP8Text
nemotron-70b-fp8

FP8 serving variant. NVIDIA's reward-tuned Llama derivatives.

RLHFFP8
Context
128k
Speed
44 tok/s
Input
$0.199 / 1M tokens
Output
$0.478 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
nvidia/Llama-3.1-Nemotron-70B-Instruct-HF-FP8
Nemotron 70B INT4 AWQText
nemotron-70b-awq

INT4 AWQ serving variant. NVIDIA's reward-tuned Llama derivatives.

RLHFINT4
Context
128k
Speed
66 tok/s
Input
$0.122 / 1M tokens
Output
$0.293 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
nvidia/Llama-3.1-Nemotron-70B-Instruct-HF-AWQ
Nemotron 70B INT8 GPTQText
nemotron-70b-gptq

INT8 GPTQ serving variant. NVIDIA's reward-tuned Llama derivatives.

RLHFINT8
Context
128k
Speed
53 tok/s
Input
$0.161 / 1M tokens
Output
$0.385 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
nvidia/Llama-3.1-Nemotron-70B-Instruct-HF-GPTQ
Nemotron 51BText
nemotron-51b

NVIDIA's reward-tuned Llama derivatives.

RLHF
Context
128k
Speed
38 tok/s
Input
$0.239 / 1M tokens
Output
$0.574 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
nvidia/Llama-3_1-Nemotron-51B-Instruct
Nemotron 51B FP8Text
nemotron-51b-fp8

FP8 serving variant. NVIDIA's reward-tuned Llama derivatives.

RLHFFP8
Context
128k
Speed
58 tok/s
Input
$0.148 / 1M tokens
Output
$0.356 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
nvidia/Llama-3_1-Nemotron-51B-Instruct-FP8
Nemotron 51B INT4 AWQText
nemotron-51b-awq

INT4 AWQ serving variant. NVIDIA's reward-tuned Llama derivatives.

RLHFINT4
Context
128k
Speed
84 tok/s
Input
$0.091 / 1M tokens
Output
$0.218 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
nvidia/Llama-3_1-Nemotron-51B-Instruct-AWQ
Nemotron 51B INT8 GPTQText
nemotron-51b-gptq

INT8 GPTQ serving variant. NVIDIA's reward-tuned Llama derivatives.

RLHFINT8
Context
128k
Speed
68 tok/s
Input
$0.119 / 1M tokens
Output
$0.287 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
nvidia/Llama-3_1-Nemotron-51B-Instruct-GPTQ
Nemotron Mini 4BText
nemotron-mini-4b

NVIDIA's reward-tuned Llama derivatives.

RLHF
Context
128k
Speed
203 tok/s
Input
$0.037 / 1M tokens
Output
$0.089 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
nvidia/Nemotron-Mini-4B-Instruct
SmolLM2 135MText
smollm2-135m

Tiny models that still follow instructions. The bottom of the size curve.

tinyedge
Context
8k
Speed
313 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceTB/SmolLM2-135M-Instruct
SmolLM2 360MText
smollm2-360m

Tiny models that still follow instructions. The bottom of the size curve.

tinyedge
Context
8k
Speed
304 tok/s
Input
$0.022 / 1M tokens
Output
$0.053 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceTB/SmolLM2-360M-Instruct
SmolLM2 1.7BText
smollm2-1-7b

Tiny models that still follow instructions. The bottom of the size curve.

tinyedge
Context
8k
Speed
257 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceTB/SmolLM2-1.7B-Instruct
StarCoder2 3bCode
starcoder2-3b

Permissively trained code models over 600+ languages.

codeopen-data
Context
16k
Speed
224 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
bigcode/starcoder2-3b
StarCoder2 7bCode
starcoder2-7b

Permissively trained code models over 600+ languages.

codeopen-data
Context
16k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
bigcode/starcoder2-7b
StarCoder2 7b FP8Code
starcoder2-7b-fp8

FP8 serving variant. Permissively trained code models over 600+ languages.

codeopen-dataFP8
Context
16k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
bigcode/starcoder2-7b-FP8
StarCoder2 7b INT4 AWQCode
starcoder2-7b-awq

INT4 AWQ serving variant. Permissively trained code models over 600+ languages.

codeopen-dataINT4
Context
16k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
bigcode/starcoder2-7b-AWQ
StarCoder2 7b INT8 GPTQCode
starcoder2-7b-gptq

INT8 GPTQ serving variant. Permissively trained code models over 600+ languages.

codeopen-dataINT8
Context
16k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
bigcode/starcoder2-7b-GPTQ
StarCoder2 15b FP8Code
starcoder2-15b-fp8

FP8 serving variant. Permissively trained code models over 600+ languages.

codeopen-dataFP8
Context
16k
Speed
137 tok/s
Input
$0.053 / 1M tokens
Output
$0.126 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
bigcode/starcoder2-15b-FP8
StarCoder2 15b INT4 AWQCode
starcoder2-15b-awq

INT4 AWQ serving variant. Permissively trained code models over 600+ languages.

codeopen-dataINT4
Context
16k
Speed
176 tok/s
Input
$0.032 / 1M tokens
Output
$0.078 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
bigcode/starcoder2-15b-AWQ
StarCoder2 15b INT8 GPTQCode
starcoder2-15b-gptq

INT8 GPTQ serving variant. Permissively trained code models over 600+ languages.

codeopen-dataINT8
Context
16k
Speed
154 tok/s
Input
$0.043 / 1M tokens
Output
$0.102 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
bigcode/starcoder2-15b-GPTQ
CodeLlama 7bCode
codellama-7b

Long-standing code baseline with infilling and Python specialisation.

code
Context
16k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-7b-Instruct-hf
CodeLlama 7b FP8Code
codellama-7b-fp8

FP8 serving variant. Long-standing code baseline with infilling and Python specialisation.

codeFP8
Context
16k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-7b-Instruct-hf-FP8
CodeLlama 7b INT4 AWQCode
codellama-7b-awq

INT4 AWQ serving variant. Long-standing code baseline with infilling and Python specialisation.

codeINT4
Context
16k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-7b-Instruct-hf-AWQ
CodeLlama 7b INT8 GPTQCode
codellama-7b-gptq

INT8 GPTQ serving variant. Long-standing code baseline with infilling and Python specialisation.

codeINT8
Context
16k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-7b-Instruct-hf-GPTQ
CodeLlama 13bCode
codellama-13b

Long-standing code baseline with infilling and Python specialisation.

code
Context
16k
Speed
112 tok/s
Input
$0.076 / 1M tokens
Output
$0.182 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-13b-Instruct-hf
CodeLlama 13b FP8Code
codellama-13b-fp8

FP8 serving variant. Long-standing code baseline with infilling and Python specialisation.

codeFP8
Context
16k
Speed
148 tok/s
Input
$0.047 / 1M tokens
Output
$0.113 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-13b-Instruct-hf-FP8
CodeLlama 13b INT4 AWQCode
codellama-13b-awq

INT4 AWQ serving variant. Long-standing code baseline with infilling and Python specialisation.

codeINT4
Context
16k
Speed
187 tok/s
Input
$0.029 / 1M tokens
Output
$0.069 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-13b-Instruct-hf-AWQ
CodeLlama 13b INT8 GPTQCode
codellama-13b-gptq

INT8 GPTQ serving variant. Long-standing code baseline with infilling and Python specialisation.

codeINT8
Context
16k
Speed
165 tok/s
Input
$0.038 / 1M tokens
Output
$0.091 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-13b-Instruct-hf-GPTQ
CodeLlama 34bCode
codellama-34b

Long-standing code baseline with infilling and Python specialisation.

code
Context
16k
Speed
54 tok/s
Input
$0.166 / 1M tokens
Output
$0.398 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-34b-Instruct-hf
CodeLlama 34b FP8Code
codellama-34b-fp8

FP8 serving variant. Long-standing code baseline with infilling and Python specialisation.

codeFP8
Context
16k
Speed
79 tok/s
Input
$0.103 / 1M tokens
Output
$0.247 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-34b-Instruct-hf-FP8
CodeLlama 34b INT4 AWQCode
codellama-34b-awq

INT4 AWQ serving variant. Long-standing code baseline with infilling and Python specialisation.

codeINT4
Context
16k
Speed
112 tok/s
Input
$0.063 / 1M tokens
Output
$0.151 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-34b-Instruct-hf-AWQ
CodeLlama 34b INT8 GPTQCode
codellama-34b-gptq

INT8 GPTQ serving variant. Long-standing code baseline with infilling and Python specialisation.

codeINT8
Context
16k
Speed
93 tok/s
Input
$0.083 / 1M tokens
Output
$0.199 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-34b-Instruct-hf-GPTQ
CodeLlama 70bCode
codellama-70b

Long-standing code baseline with infilling and Python specialisation.

code
Context
16k
Speed
29 tok/s
Input
$0.321 / 1M tokens
Output
$0.77 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-70b-Instruct-hf
CodeLlama 70b FP8Code
codellama-70b-fp8

FP8 serving variant. Long-standing code baseline with infilling and Python specialisation.

codeFP8
Context
16k
Speed
44 tok/s
Input
$0.199 / 1M tokens
Output
$0.478 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-70b-Instruct-hf-FP8
CodeLlama 70b INT4 AWQCode
codellama-70b-awq

INT4 AWQ serving variant. Long-standing code baseline with infilling and Python specialisation.

codeINT4
Context
16k
Speed
66 tok/s
Input
$0.122 / 1M tokens
Output
$0.293 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-70b-Instruct-hf-AWQ
CodeLlama 70b INT8 GPTQCode
codellama-70b-gptq

INT8 GPTQ serving variant. Long-standing code baseline with infilling and Python specialisation.

codeINT8
Context
16k
Speed
53 tok/s
Input
$0.161 / 1M tokens
Output
$0.385 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/CodeLlama-70b-Instruct-hf-GPTQ
InternVL2.5 1BVision
internvl2-5-1b

Vision-language models with strong chart, table and document reading.

visionOCR
Context
32k
Speed
280 tok/s
Input
$0.024 / 1M tokens
Output
$0.058 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-1B
InternVL2.5 2BVision
internvl2-5-2b

Vision-language models with strong chart, table and document reading.

visionOCR
Context
32k
Speed
248 tok/s
Input
$0.029 / 1M tokens
Output
$0.07 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-2B
InternVL2.5 4BVision
internvl2-5-4b

Vision-language models with strong chart, table and document reading.

visionOCR
Context
32k
Speed
203 tok/s
Input
$0.037 / 1M tokens
Output
$0.089 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-4B
InternVL2.5 8BVision
internvl2-5-8b

Vision-language models with strong chart, table and document reading.

visionOCR
Context
32k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-8B
InternVL2.5 8B FP8Vision
internvl2-5-8b-fp8

FP8 serving variant. Vision-language models with strong chart, table and document reading.

visionOCRFP8
Context
32k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-8B-FP8
InternVL2.5 8B INT4 AWQVision
internvl2-5-8b-awq

INT4 AWQ serving variant. Vision-language models with strong chart, table and document reading.

visionOCRINT4
Context
32k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-8B-AWQ
InternVL2.5 8B INT8 GPTQVision
internvl2-5-8b-gptq

INT8 GPTQ serving variant. Vision-language models with strong chart, table and document reading.

visionOCRINT8
Context
32k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-8B-GPTQ
InternVL2.5 26BVision
internvl2-5-26b

Vision-language models with strong chart, table and document reading.

visionOCR
Context
32k
Speed
67 tok/s
Input
$0.132 / 1M tokens
Output
$0.317 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-26B
InternVL2.5 26B FP8Vision
internvl2-5-26b-fp8

FP8 serving variant. Vision-language models with strong chart, table and document reading.

visionOCRFP8
Context
32k
Speed
96 tok/s
Input
$0.082 / 1M tokens
Output
$0.196 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-26B-FP8
InternVL2.5 26B INT4 AWQVision
internvl2-5-26b-awq

INT4 AWQ serving variant. Vision-language models with strong chart, table and document reading.

visionOCRINT4
Context
32k
Speed
132 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-26B-AWQ
InternVL2.5 26B INT8 GPTQVision
internvl2-5-26b-gptq

INT8 GPTQ serving variant. Vision-language models with strong chart, table and document reading.

visionOCRINT8
Context
32k
Speed
112 tok/s
Input
$0.066 / 1M tokens
Output
$0.158 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
OpenGVLab/InternVL2_5-26B-GPTQ
InternVL2.5 38BVision
internvl2-5-38b

Vision-language models with strong chart, table and document reading.

visionOCR
Context
32k
Speed
49 tok/s
Input
$0.183 / 1M tokens
Output
$0.439 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
OpenGVLab/InternVL2_5-38B
InternVL2.5 38B FP8Vision
internvl2-5-38b-fp8

FP8 serving variant. Vision-language models with strong chart, table and document reading.

visionOCRFP8
Context
32k
Speed
73 tok/s
Input
$0.113 / 1M tokens
Output
$0.272 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
OpenGVLab/InternVL2_5-38B-FP8
InternVL2.5 38B INT4 AWQVision
internvl2-5-38b-awq

INT4 AWQ serving variant. Vision-language models with strong chart, table and document reading.

visionOCRINT4
Context
32k
Speed
104 tok/s
Input
$0.07 / 1M tokens
Output
$0.167 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
OpenGVLab/InternVL2_5-38B-AWQ
InternVL2.5 38B INT8 GPTQVision
internvl2-5-38b-gptq

INT8 GPTQ serving variant. Vision-language models with strong chart, table and document reading.

visionOCRINT8
Context
32k
Speed
86 tok/s
Input
$0.091 / 1M tokens
Output
$0.22 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
OpenGVLab/InternVL2_5-38B-GPTQ
InternVL2.5 78BVision
internvl2-5-78b

Vision-language models with strong chart, table and document reading.

visionOCR
Context
32k
Speed
26 tok/s
Input
$0.355 / 1M tokens
Output
$0.852 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
OpenGVLab/InternVL2_5-78B
InternVL2.5 78B FP8Vision
internvl2-5-78b-fp8

FP8 serving variant. Vision-language models with strong chart, table and document reading.

visionOCRFP8
Context
32k
Speed
40 tok/s
Input
$0.22 / 1M tokens
Output
$0.528 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
OpenGVLab/InternVL2_5-78B-FP8
InternVL2.5 78B INT4 AWQVision
internvl2-5-78b-awq

INT4 AWQ serving variant. Vision-language models with strong chart, table and document reading.

visionOCRINT4
Context
32k
Speed
61 tok/s
Input
$0.135 / 1M tokens
Output
$0.324 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
OpenGVLab/InternVL2_5-78B-AWQ
InternVL2.5 78B INT8 GPTQVision
internvl2-5-78b-gptq

INT8 GPTQ serving variant. Vision-language models with strong chart, table and document reading.

visionOCRINT8
Context
32k
Speed
48 tok/s
Input
$0.177 / 1M tokens
Output
$0.426 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
OpenGVLab/InternVL2_5-78B-GPTQ
LLaVA 1.6 Mistral 7BVision
llava-1-6-mistral-7b

The open vision-language baseline. Broad ecosystem support.

vision
Context
32k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-mistral-7b-hf
LLaVA 1.6 Mistral 7B FP8Vision
llava-1-6-mistral-7b-fp8

FP8 serving variant. The open vision-language baseline. Broad ecosystem support.

visionFP8
Context
32k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-mistral-7b-hf-FP8
LLaVA 1.6 Mistral 7B INT4 AWQVision
llava-1-6-mistral-7b-awq

INT4 AWQ serving variant. The open vision-language baseline. Broad ecosystem support.

visionINT4
Context
32k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-mistral-7b-hf-AWQ
LLaVA 1.6 Mistral 7B INT8 GPTQVision
llava-1-6-mistral-7b-gptq

INT8 GPTQ serving variant. The open vision-language baseline. Broad ecosystem support.

visionINT8
Context
32k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-mistral-7b-hf-GPTQ
LLaVA 1.6 Vicuna 13BVision
llava-1-6-vicuna-13b

The open vision-language baseline. Broad ecosystem support.

vision
Context
32k
Speed
112 tok/s
Input
$0.076 / 1M tokens
Output
$0.182 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-vicuna-13b-hf
LLaVA 1.6 Vicuna 13B FP8Vision
llava-1-6-vicuna-13b-fp8

FP8 serving variant. The open vision-language baseline. Broad ecosystem support.

visionFP8
Context
32k
Speed
148 tok/s
Input
$0.047 / 1M tokens
Output
$0.113 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-vicuna-13b-hf-FP8
LLaVA 1.6 Vicuna 13B INT4 AWQVision
llava-1-6-vicuna-13b-awq

INT4 AWQ serving variant. The open vision-language baseline. Broad ecosystem support.

visionINT4
Context
32k
Speed
187 tok/s
Input
$0.029 / 1M tokens
Output
$0.069 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-vicuna-13b-hf-AWQ
LLaVA 1.6 Vicuna 13B INT8 GPTQVision
llava-1-6-vicuna-13b-gptq

INT8 GPTQ serving variant. The open vision-language baseline. Broad ecosystem support.

visionINT8
Context
32k
Speed
165 tok/s
Input
$0.038 / 1M tokens
Output
$0.091 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-vicuna-13b-hf-GPTQ
LLaVA 1.6 34BVision
llava-1-6-34b

The open vision-language baseline. Broad ecosystem support.

vision
Context
32k
Speed
54 tok/s
Input
$0.166 / 1M tokens
Output
$0.398 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-34b-hf
LLaVA 1.6 34B FP8Vision
llava-1-6-34b-fp8

FP8 serving variant. The open vision-language baseline. Broad ecosystem support.

visionFP8
Context
32k
Speed
79 tok/s
Input
$0.103 / 1M tokens
Output
$0.247 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-34b-hf-FP8
LLaVA 1.6 34B INT4 AWQVision
llava-1-6-34b-awq

INT4 AWQ serving variant. The open vision-language baseline. Broad ecosystem support.

visionINT4
Context
32k
Speed
112 tok/s
Input
$0.063 / 1M tokens
Output
$0.151 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-34b-hf-AWQ
LLaVA 1.6 34B INT8 GPTQVision
llava-1-6-34b-gptq

INT8 GPTQ serving variant. The open vision-language baseline. Broad ecosystem support.

visionINT8
Context
32k
Speed
93 tok/s
Input
$0.083 / 1M tokens
Output
$0.199 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
llava-hf/llava-v1.6-34b-hf-GPTQ
MiniCPM-V 2.6Vision
minicpm-v-2-6

Small vision-language models that run on a single consumer GPU.

visionedge
Context
32k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
openbmb/MiniCPM-V-2_6
MiniCPM-V 2.6 FP8Vision
minicpm-v-2-6-fp8

FP8 serving variant. Small vision-language models that run on a single consumer GPU.

visionedgeFP8
Context
32k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
openbmb/MiniCPM-V-2_6-FP8
MiniCPM-V 2.6 INT4 AWQVision
minicpm-v-2-6-awq

INT4 AWQ serving variant. Small vision-language models that run on a single consumer GPU.

visionedgeINT4
Context
32k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
openbmb/MiniCPM-V-2_6-AWQ
MiniCPM-V 2.6 INT8 GPTQVision
minicpm-v-2-6-gptq

INT8 GPTQ serving variant. Small vision-language models that run on a single consumer GPU.

visionedgeINT8
Context
32k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
openbmb/MiniCPM-V-2_6-GPTQ
Pixtral 12B FP8Vision
pixtral-12b-fp8

FP8 serving variant. Mistral's multimodal model, native variable-resolution image input.

visionFP8
Context
128k
Speed
155 tok/s
Input
$0.045 / 1M tokens
Output
$0.107 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Pixtral-12B-2409-FP8
Pixtral 12B INT4 AWQVision
pixtral-12b-awq

INT4 AWQ serving variant. Mistral's multimodal model, native variable-resolution image input.

visionINT4
Context
128k
Speed
193 tok/s
Input
$0.027 / 1M tokens
Output
$0.066 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Pixtral-12B-2409-AWQ
Pixtral 12B INT8 GPTQVision
pixtral-12b-gptq

INT8 GPTQ serving variant. Mistral's multimodal model, native variable-resolution image input.

visionINT8
Context
128k
Speed
172 tok/s
Input
$0.036 / 1M tokens
Output
$0.086 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mistralai/Pixtral-12B-2409-GPTQ
Molmo 7B-DVision
molmo-7b-d

Open vision-language models with pointing and grounding supervision.

visiongrounding
Context
32k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/Molmo-7B-D-0924
Molmo 7B-D FP8Vision
molmo-7b-d-fp8

FP8 serving variant. Open vision-language models with pointing and grounding supervision.

visiongroundingFP8
Context
32k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/Molmo-7B-D-0924-FP8
Molmo 7B-D INT4 AWQVision
molmo-7b-d-awq

INT4 AWQ serving variant. Open vision-language models with pointing and grounding supervision.

visiongroundingINT4
Context
32k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/Molmo-7B-D-0924-AWQ
Molmo 7B-D INT8 GPTQVision
molmo-7b-d-gptq

INT8 GPTQ serving variant. Open vision-language models with pointing and grounding supervision.

visiongroundingINT8
Context
32k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
allenai/Molmo-7B-D-0924-GPTQ
Molmo 72BVision
molmo-72b

Open vision-language models with pointing and grounding supervision.

visiongrounding
Context
32k
Speed
28 tok/s
Input
$0.33 / 1M tokens
Output
$0.792 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
allenai/Molmo-72B-0924
Molmo 72B FP8Vision
molmo-72b-fp8

FP8 serving variant. Open vision-language models with pointing and grounding supervision.

visiongroundingFP8
Context
32k
Speed
43 tok/s
Input
$0.205 / 1M tokens
Output
$0.491 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
allenai/Molmo-72B-0924-FP8
Molmo 72B INT4 AWQVision
molmo-72b-awq

INT4 AWQ serving variant. Open vision-language models with pointing and grounding supervision.

visiongroundingINT4
Context
32k
Speed
65 tok/s
Input
$0.125 / 1M tokens
Output
$0.301 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
allenai/Molmo-72B-0924-AWQ
Molmo 72B INT8 GPTQVision
molmo-72b-gptq

INT8 GPTQ serving variant. Open vision-language models with pointing and grounding supervision.

visiongroundingINT8
Context
32k
Speed
52 tok/s
Input
$0.165 / 1M tokens
Output
$0.396 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
allenai/Molmo-72B-0924-GPTQ
Idefics3 8BVision
idefics3-8b

Document-heavy multimodal model built on Llama.

visiondocuments
Context
16k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceM4/Idefics3-8B-Llama3
Idefics3 8B FP8Vision
idefics3-8b-fp8

FP8 serving variant. Document-heavy multimodal model built on Llama.

visiondocumentsFP8
Context
16k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceM4/Idefics3-8B-Llama3-FP8
Idefics3 8B INT4 AWQVision
idefics3-8b-awq

INT4 AWQ serving variant. Document-heavy multimodal model built on Llama.

visiondocumentsINT4
Context
16k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceM4/Idefics3-8B-Llama3-AWQ
Idefics3 8B INT8 GPTQVision
idefics3-8b-gptq

INT8 GPTQ serving variant. Document-heavy multimodal model built on Llama.

visiondocumentsINT8
Context
16k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceM4/Idefics3-8B-Llama3-GPTQ
GOT-OCR2 0.6BDocument / OCR
got-ocr2-0-6b

End-to-end OCR covering formulas, tables, sheet music and charts.

OCR
Context
8k
Speed
294 tok/s
Input
$0.023 / page
Output
$0.055 / page
Dedicated
1 × H100 · $2.14/hr
stepfun-ai/GOT-OCR2_0
Florence-2 baseVision
florence-2-base

Compact unified vision model — captioning, detection and segmentation.

vision
Context
4k
Speed
309 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Florence-2-base
Florence-2 largeVision
florence-2-large

Compact unified vision model — captioning, detection and segmentation.

vision
Context
4k
Speed
288 tok/s
Input
$0.023 / 1M tokens
Output
$0.055 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Florence-2-large
BGE M3 m3Embedding
bge-m3-m3

Multilingual, multi-granularity retrieval embeddings.

retrievalmultilingual
Context
8k
Speed
295 tok/s
Input
$0.022 / 1M tokens
Output
$0.053 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
BAAI/bge-m3
BGE smallEmbedding
bge-small

The retrieval workhorse. Dense embeddings across three sizes.

retrieval
Context
512
Speed
318 tok/s
Input
$0.02 / 1M tokens
Output
$0.048 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
BAAI/bge-small-en-v1.5
BGE baseEmbedding
bge-base

The retrieval workhorse. Dense embeddings across three sizes.

retrieval
Context
512
Speed
315 tok/s
Input
$0.02 / 1M tokens
Output
$0.048 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
BAAI/bge-base-en-v1.5
BGE largeEmbedding
bge-large

The retrieval workhorse. Dense embeddings across three sizes.

retrieval
Context
512
Speed
305 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
BAAI/bge-large-en-v1.5
BGE Reranker baseRerank
bge-reranker-base

Cross-encoder reranking — the expensive half of a good retrieval stack.

rerank
Context
8k
Speed
307 tok/s
Input
$0.021 / 1K queries
Output
$0.05 / 1K queries
Dedicated
1 × H100 · $2.14/hr
BAAI/bge-reranker-base
BGE Reranker largeRerank
bge-reranker-large

Cross-encoder reranking — the expensive half of a good retrieval stack.

rerank
Context
8k
Speed
296 tok/s
Input
$0.022 / 1K queries
Output
$0.053 / 1K queries
Dedicated
1 × H100 · $2.14/hr
BAAI/bge-reranker-large
E5 smallEmbedding
e5-small

Instruction-tuned embeddings with strong zero-shot retrieval.

retrieval
Context
512
Speed
314 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
intfloat/multilingual-e5-small
E5 baseEmbedding
e5-base

Instruction-tuned embeddings with strong zero-shot retrieval.

retrieval
Context
512
Speed
307 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
intfloat/multilingual-e5-base
E5 largeEmbedding
e5-large

Instruction-tuned embeddings with strong zero-shot retrieval.

retrieval
Context
512
Speed
296 tok/s
Input
$0.022 / 1M tokens
Output
$0.053 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
intfloat/multilingual-e5-large
GTE baseEmbedding
gte-base

Long-context general text embeddings.

retrievallong-context
Context
8k
Speed
313 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Alibaba-NLP/gte-base-en-v1.5
GTE largeEmbedding
gte-large

Long-context general text embeddings.

retrievallong-context
Context
8k
Speed
301 tok/s
Input
$0.022 / 1M tokens
Output
$0.053 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Alibaba-NLP/gte-large-en-v1.5
GTE Qwen2-7BEmbedding
gte-qwen2-7b

Long-context general text embeddings.

retrievallong-context
Context
8k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Alibaba-NLP/gte-Qwen2-7B-instruct
Nomic Embed v1.5Embedding
nomic-embed-v1-5

Fully open embeddings with reproducible training data.

retrievalopen-data
Context
8k
Speed
313 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
nomic-ai/nomic-embed-text-v1.5
Jina Embeddings v3Embedding
jina-embeddings-v3

Bilingual and code-aware embedding models.

retrieval
Context
8k
Speed
295 tok/s
Input
$0.022 / 1M tokens
Output
$0.053 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
jinaai/jina-embeddings-v3
Jina Embeddings v2 base codeEmbedding
jina-embeddings-v2-base-code

Bilingual and code-aware embedding models.

retrieval
Context
8k
Speed
312 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
jinaai/jina-embeddings-v2-base-code
mxbai large v1Embedding
mxbai-large-v1

Compact embeddings tuned for Matryoshka truncation.

retrieval
Context
512
Speed
305 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
mixedbread-ai/mxbai-embed-large-v1
Stella 400M v5Embedding
stella-400m-v5

High-ranking compact retrieval embeddings.

retrieval
Context
8k
Speed
302 tok/s
Input
$0.022 / 1M tokens
Output
$0.053 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
dunzhang/stella_en_400M_v5
Stella 1.5B v5Embedding
stella-1-5b-v5

High-ranking compact retrieval embeddings.

retrieval
Context
8k
Speed
263 tok/s
Input
$0.026 / 1M tokens
Output
$0.062 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
dunzhang/stella_en_1.5B_v5
Whisper tinySpeech
whisper-tiny

The transcription baseline. 99 languages with timestamped output.

ASR
Context
30s window
Speed
318 tok/s
Input
$0.02 / audio min
Output
$0.048 / audio min
Dedicated
1 × H100 · $2.14/hr
openai/whisper-tiny
Whisper baseSpeech
whisper-base

The transcription baseline. 99 languages with timestamped output.

ASR
Context
30s window
Speed
316 tok/s
Input
$0.02 / audio min
Output
$0.048 / audio min
Dedicated
1 × H100 · $2.14/hr
openai/whisper-base
Whisper smallSpeech
whisper-small

The transcription baseline. 99 languages with timestamped output.

ASR
Context
30s window
Speed
309 tok/s
Input
$0.021 / audio min
Output
$0.05 / audio min
Dedicated
1 × H100 · $2.14/hr
openai/whisper-small
Whisper mediumSpeech
whisper-medium

The transcription baseline. 99 languages with timestamped output.

ASR
Context
30s window
Speed
288 tok/s
Input
$0.023 / audio min
Output
$0.055 / audio min
Dedicated
1 × H100 · $2.14/hr
openai/whisper-medium
Whisper large-v3-turboSpeech
whisper-large-v3-turbo

The transcription baseline. 99 languages with timestamped output.

ASR
Context
30s window
Speed
286 tok/s
Input
$0.023 / audio min
Output
$0.055 / audio min
Dedicated
1 × H100 · $2.14/hr
openai/whisper-large-v3-turbo
Distil-Whisper small.enSpeech
distil-whisper-small-en

Distilled Whisper — most of the accuracy at a fraction of the decode cost.

ASRfast
Context
30s window
Speed
312 tok/s
Input
$0.021 / audio min
Output
$0.05 / audio min
Dedicated
1 × H100 · $2.14/hr
distil-whisper/distil-small.en
Distil-Whisper medium.enSpeech
distil-whisper-medium-en

Distilled Whisper — most of the accuracy at a fraction of the decode cost.

ASRfast
Context
30s window
Speed
302 tok/s
Input
$0.022 / audio min
Output
$0.053 / audio min
Dedicated
1 × H100 · $2.14/hr
distil-whisper/distil-medium.en
Distil-Whisper large-v3Speech
distil-whisper-large-v3

Distilled Whisper — most of the accuracy at a fraction of the decode cost.

ASRfast
Context
30s window
Speed
288 tok/s
Input
$0.023 / audio min
Output
$0.055 / audio min
Dedicated
1 × H100 · $2.14/hr
distil-whisper/distil-large-v3
Parakeet TDT 0.6BSpeech
parakeet-tdt-0-6b

Streaming ASR built for realtime transcription.

ASRrealtime
Context
streaming
Speed
294 tok/s
Input
$0.023 / audio min
Output
$0.055 / audio min
Dedicated
1 × H100 · $2.14/hr
nvidia/parakeet-tdt-0.6b-v2
Parakeet CTC 1.1BSpeech
parakeet-ctc-1-1b

Streaming ASR built for realtime transcription.

ASRrealtime
Context
streaming
Speed
276 tok/s
Input
$0.025 / audio min
Output
$0.06 / audio min
Dedicated
1 × H100 · $2.14/hr
nvidia/parakeet-ctc-1.1b
Canary 1BSpeech
canary-1b

Multilingual ASR with speech translation in the same pass.

ASRtranslation
Context
streaming
Speed
280 tok/s
Input
$0.024 / audio min
Output
$0.058 / audio min
Dedicated
1 × H100 · $2.14/hr
nvidia/canary-1b
Canary 180M flashSpeech
canary-180m-flash

Multilingual ASR with speech translation in the same pass.

ASRtranslation
Context
streaming
Speed
311 tok/s
Input
$0.021 / audio min
Output
$0.05 / audio min
Dedicated
1 × H100 · $2.14/hr
nvidia/canary-180m-flash
SeamlessM4T v2 largeSpeech
seamlessm4t-v2-large

Speech-to-speech and speech-to-text translation across 100 languages.

translation
Context
streaming
Speed
240 tok/s
Input
$0.03 / audio min
Output
$0.072 / audio min
Dedicated
8 × B200 · $54.40/hr
facebook/seamless-m4t-v2-large
MMS 1B ASRSpeech
mms-1b-asr

Massively multilingual speech — ASR and TTS across 1,000+ languages.

multilingual
Context
30s window
Speed
280 tok/s
Input
$0.024 / audio min
Output
$0.058 / audio min
Dedicated
1 × H100 · $2.14/hr
facebook/mms-1b-all
MMS TTSSpeech
mms-tts

Massively multilingual speech — ASR and TTS across 1,000+ languages.

multilingual
Context
30s window
Speed
304 tok/s
Input
$0.022 / audio min
Output
$0.053 / audio min
Dedicated
1 × H100 · $2.14/hr
facebook/mms-tts
Kokoro 82MSpeech
kokoro-82m

Very small TTS with natural prosody. Cheap enough for realtime at scale.

TTSsmall
Context
streaming
Speed
316 tok/s
Input
$0.02 / audio min
Output
$0.048 / audio min
Dedicated
1 × H100 · $2.14/hr
hexgrad/Kokoro-82M
Bark smallSpeech
bark-small

Generative audio — speech, sound effects and non-verbal sounds.

TTS
Context
14s window
Speed
302 tok/s
Input
$0.022 / audio min
Output
$0.053 / audio min
Dedicated
1 × H100 · $2.14/hr
suno/bark-small
Bark baseSpeech
bark-base

Generative audio — speech, sound effects and non-verbal sounds.

TTS
Context
14s window
Speed
283 tok/s
Input
$0.024 / audio min
Output
$0.058 / audio min
Dedicated
1 × H100 · $2.14/hr
suno/bark
MusicGen smallAudio / Music
musicgen-small

Text-conditioned music generation with optional melody conditioning.

music
Context
30s window
Speed
306 tok/s
Input
$0.021 / audio min
Output
$0.05 / audio min
Dedicated
1 × H100 · $2.14/hr
facebook/musicgen-small
MusicGen mediumAudio / Music
musicgen-medium

Text-conditioned music generation with optional melody conditioning.

music
Context
30s window
Speed
263 tok/s
Input
$0.026 / audio min
Output
$0.062 / audio min
Dedicated
1 × H100 · $2.14/hr
facebook/musicgen-medium
Stable Audio open 1.0Audio / Music
stable-audio-open-1-0

Text-to-audio for sound design and loops.

musicsfx
Context
47s window
Speed
276 tok/s
Input
$0.025 / audio min
Output
$0.06 / audio min
Dedicated
1 × H100 · $2.14/hr
stabilityai/stable-audio-open-1.0
Stable Diffusion 1.5Image
stable-diffusion-1-5

The open image baseline with the deepest fine-tune ecosystem.

image
Context
—
Speed
284 tok/s
Input
$0.024 / image
Output
$0.058 / image
Dedicated
1 × H100 · $2.14/hr
stabilityai/stable-diffusion-v1-5
Stable Diffusion 2.1Image
stable-diffusion-2-1

The open image baseline with the deepest fine-tune ecosystem.

image
Context
—
Speed
284 tok/s
Input
$0.024 / image
Output
$0.058 / image
Dedicated
1 × H100 · $2.14/hr
stabilityai/stable-diffusion-2-1
Stable Diffusion XLImage
stable-diffusion-xl

The open image baseline with the deepest fine-tune ecosystem.

image
Context
—
Speed
213 tok/s
Input
$0.035 / image
Output
$0.084 / image
Dedicated
1 × H100 · $2.14/hr
stabilityai/stable-diffusion-xl-base-1.0
Stable Diffusion 3 MediumImage
stable-diffusion-3-medium

The open image baseline with the deepest fine-tune ecosystem.

image
Context
—
Speed
248 tok/s
Input
$0.029 / image
Output
$0.07 / image
Dedicated
1 × H100 · $2.14/hr
stabilityai/stable-diffusion-3-medium
Stable Diffusion 3.5 LargeImage
stable-diffusion-3-5-large

The open image baseline with the deepest fine-tune ecosystem.

image
Context
—
Speed
149 tok/s
Input
$0.054 / image
Output
$0.13 / image
Dedicated
1 × H100 · $2.14/hr
stabilityai/stable-diffusion-3.5-large
Stable Diffusion 3.5 Large TurboImage
stable-diffusion-3-5-large-turbo

The open image baseline with the deepest fine-tune ecosystem.

image
Context
—
Speed
149 tok/s
Input
$0.054 / image
Output
$0.13 / image
Dedicated
1 × H100 · $2.14/hr
stabilityai/stable-diffusion-3.5-large-turbo
FLUX.1 schnellImage
flux-1-schnell

Rectified-flow transformer. Current open quality leader for text rendering.

image
Context
—
Speed
117 tok/s
Input
$0.072 / image
Output
$0.173 / image
Dedicated
1 × H100 · $2.14/hr
black-forest-labs/FLUX.1-schnell
SANA 600MImage
sana-600m

Linear-attention diffusion — 4K images on modest hardware.

imagefast
Context
—
Speed
294 tok/s
Input
$0.023 / image
Output
$0.055 / image
Dedicated
1 × H100 · $2.14/hr
Efficient-Large-Model/Sana_600M_1024px_diffusers
SANA 1.6BImage
sana-1-6b

Linear-attention diffusion — 4K images on modest hardware.

imagefast
Context
—
Speed
260 tok/s
Input
$0.027 / image
Output
$0.065 / image
Dedicated
1 × H100 · $2.14/hr
Efficient-Large-Model/Sana_1600M_1024px_diffusers
Kandinsky 2.2Image
kandinsky-2-2

Latent diffusion with a strong prior model.

image
Context
—
Speed
193 tok/s
Input
$0.04 / image
Output
$0.096 / image
Dedicated
1 × H100 · $2.14/hr
kandinsky-community/kandinsky-2-2-decoder
Kandinsky 3Image
kandinsky-3

Latent diffusion with a strong prior model.

image
Context
—
Speed
118 tok/s
Input
$0.071 / image
Output
$0.17 / image
Dedicated
1 × H100 · $2.14/hr
kandinsky-community/kandinsky-3
Playground v2.5Image
playground-v2-5

Aesthetic-tuned SDXL derivative.

imageaesthetic
Context
—
Speed
213 tok/s
Input
$0.035 / image
Output
$0.084 / image
Dedicated
1 × H100 · $2.14/hr
playgroundai/playground-v2.5-1024px-aesthetic
Kolors 1.0Image
kolors-1-0

Bilingual text-to-image with strong Chinese prompt handling.

imagebilingual
Context
—
Speed
233 tok/s
Input
$0.031 / image
Output
$0.074 / image
Dedicated
1 × H100 · $2.14/hr
Kwai-Kolors/Kolors-diffusers
AuraFlow 0.3Image
auraflow-0-3

Fully open flow-based text-to-image.

imageopen
Context
—
Speed
162 tok/s
Input
$0.049 / image
Output
$0.118 / image
Dedicated
1 × H100 · $2.14/hr
fal/AuraFlow-v0.3
Stable Video Diffusion XTVideo
stable-video-diffusion-xt

Image-to-video with camera motion control.

video
Context
25 frames
Speed
263 tok/s
Input
$0.026 / second
Output
$0.062 / second
Dedicated
1 × H100 · $2.14/hr
stabilityai/stable-video-diffusion-img2vid-xt
CogVideoX 2BVideo
cogvideox-2b

Text-to-video diffusion transformer with usable motion coherence.

video
Context
6s
Speed
248 tok/s
Input
$0.029 / second
Output
$0.07 / second
Dedicated
1 × H100 · $2.14/hr
THUDM/CogVideoX-2b
CogVideoX 5BVideo
cogvideox-5b

Text-to-video diffusion transformer with usable motion coherence.

video
Context
6s
Speed
186 tok/s
Input
$0.041 / second
Output
$0.098 / second
Dedicated
1 × H100 · $2.14/hr
THUDM/CogVideoX-5b
Mochi 1 previewVideo
mochi-1-preview

Open video generation with strong prompt adherence.

video
Context
5s
Speed
131 tok/s
Input
$0.063 / second
Output
$0.151 / second
Dedicated
1 × H100 · $2.14/hr
genmo/mochi-1-preview
LTX Video 2BVideo
ltx-video-2b

Realtime-class video generation — faster than playback on a single GPU.

videofast
Context
5s
Speed
248 tok/s
Input
$0.029 / second
Output
$0.07 / second
Dedicated
1 × H100 · $2.14/hr
Lightricks/LTX-Video
HunyuanVideo 13BVideo
hunyuanvideo-13b

Large open video model with cinematic motion.

video
Context
5s
Speed
112 tok/s
Input
$0.076 / second
Output
$0.182 / second
Dedicated
1 × H100 · $2.14/hr
tencent/HunyuanVideo
Wan 2.1 1.3BVideo
wan-2-1-1-3b

Text and image conditioned video with a small variant for consumer GPUs.

video
Context
5s
Speed
269 tok/s
Input
$0.026 / second
Output
$0.062 / second
Dedicated
1 × H100 · $2.14/hr
Wan-AI/Wan2.1-T2V-1.3B
Wan 2.1 14BVideo
wan-2-1-14b

Text and image conditioned video with a small variant for consumer GPUs.

video
Context
5s
Speed
106 tok/s
Input
$0.08 / second
Output
$0.192 / second
Dedicated
1 × H100 · $2.14/hr
Wan-AI/Wan2.1-T2V-14B
TripoSR base3D
triposr-base

Single-image to 3D mesh in under a second.

3Dfast
Context
single image
Speed
298 tok/s
Input
$0.022 / asset
Output
$0.053 / asset
Dedicated
1 × H100 · $2.14/hr
stabilityai/TripoSR
Hunyuan3D 2.03D
hunyuan3d-2-0

Image-to-3D with texture synthesis.

3D
Context
single image
Speed
276 tok/s
Input
$0.025 / asset
Output
$0.06 / asset
Dedicated
1 × H100 · $2.14/hr
tencent/Hunyuan3D-2
InstantMesh base3D
instantmesh-base

Feed-forward sparse-view reconstruction to mesh.

3D
Context
single image
Speed
294 tok/s
Input
$0.023 / asset
Output
$0.055 / asset
Dedicated
1 × H100 · $2.14/hr
TencentARC/InstantMesh
Shap-E base3D
shap-e-base

Text-to-3D implicit function generation.

3D
Context
text prompt
Speed
306 tok/s
Input
$0.021 / asset
Output
$0.05 / asset
Dedicated
1 × H100 · $2.14/hr
openai/shap-e
Genie-style world model baseWorld model
genie-style-world-model-base

Learned environment model for policy training and evaluation.

worldrobotics
Context
rollout
Speed
280 tok/s
Input
$0.024 / rollout-hr
Output
$0.058 / rollout-hr
Dedicated
1 × H100 · $2.14/hr
1x-technologies/worldmodel
Llama Guard 3 1BSafety
llama-guard-3-1b

Input and output safety classification against a configurable taxonomy.

safety
Context
8k
Speed
280 tok/s
Input
$0.024 / 1M tokens
Output
$0.058 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-Guard-3-1B
Llama Guard 4 12BSafety
llama-guard-4-12b

Input and output safety classification against a configurable taxonomy.

safety
Context
8k
Speed
117 tok/s
Input
$0.072 / 1M tokens
Output
$0.173 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
meta-llama/Llama-Guard-4-12B
ShieldGemma 2bSafety
shieldgemma-2b

Gemma-based content safety classifiers.

safety
Context
8k
Speed
248 tok/s
Input
$0.029 / 1M tokens
Output
$0.07 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/shieldgemma-2b
ShieldGemma 27bSafety
shieldgemma-27b

Gemma-based content safety classifiers.

safety
Context
8k
Speed
65 tok/s
Input
$0.136 / 1M tokens
Output
$0.326 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
google/shieldgemma-27b
Granite Guardian 3.1 2BSafety
granite-guardian-3-1-2b

Enterprise risk detection including groundedness and jailbreak checks.

safety
Context
8k
Speed
248 tok/s
Input
$0.029 / 1M tokens
Output
$0.07 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ibm-granite/granite-guardian-3.1-2b
Granite Guardian 3.1 8BSafety
granite-guardian-3-1-8b

Enterprise risk detection including groundedness and jailbreak checks.

safety
Context
8k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
ibm-granite/granite-guardian-3.1-8b
ESM-2 8MScience
esm-2-8m

Protein language model — structure and function prediction from sequence.

biology
Context
1k residues
Speed
319 tok/s
Input
$0.02 / 1M tokens
Output
$0.048 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
facebook/esm2_t6_8M_UR50D
ESM-2 35MScience
esm-2-35m

Protein language model — structure and function prediction from sequence.

biology
Context
1k residues
Speed
318 tok/s
Input
$0.02 / 1M tokens
Output
$0.048 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
facebook/esm2_t12_35M_UR50D
ESM-2 150MScience
esm-2-150m

Protein language model — structure and function prediction from sequence.

biology
Context
1k residues
Speed
313 tok/s
Input
$0.021 / 1M tokens
Output
$0.05 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
facebook/esm2_t30_150M_UR50D
ESM-2 650MScience
esm-2-650m

Protein language model — structure and function prediction from sequence.

biology
Context
1k residues
Speed
292 tok/s
Input
$0.023 / 1M tokens
Output
$0.055 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
facebook/esm2_t33_650M_UR50D
ESM-2 3BScience
esm-2-3b

Protein language model — structure and function prediction from sequence.

biology
Context
1k residues
Speed
224 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
facebook/esm2_t36_3B_UR50D
ChemBERTa 77M MTRScience
chemberta-77m-mtr

Molecular property prediction from SMILES.

chemistry
Context
512
Speed
316 tok/s
Input
$0.02 / 1M tokens
Output
$0.048 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
DeepChem/ChemBERTa-77M-MTR
Chronos tinyTime series
chronos-tiny

Pretrained forecasting on tokenised time series. Zero-shot on new series.

forecasting
Context
512 ctx
Speed
319 tok/s
Input
$0.02 / 1M rows
Output
$0.048 / 1M rows
Dedicated
1 × H100 · $2.14/hr
amazon/chronos-t5-tiny
Chronos miniTime series
chronos-mini

Pretrained forecasting on tokenised time series. Zero-shot on new series.

forecasting
Context
512 ctx
Speed
319 tok/s
Input
$0.02 / 1M rows
Output
$0.048 / 1M rows
Dedicated
1 × H100 · $2.14/hr
amazon/chronos-t5-mini
Chronos smallTime series
chronos-small

Pretrained forecasting on tokenised time series. Zero-shot on new series.

forecasting
Context
512 ctx
Speed
317 tok/s
Input
$0.02 / 1M rows
Output
$0.048 / 1M rows
Dedicated
1 × H100 · $2.14/hr
amazon/chronos-t5-small
Chronos baseTime series
chronos-base

Pretrained forecasting on tokenised time series. Zero-shot on new series.

forecasting
Context
512 ctx
Speed
311 tok/s
Input
$0.021 / 1M rows
Output
$0.05 / 1M rows
Dedicated
1 × H100 · $2.14/hr
amazon/chronos-t5-base
Chronos largeTime series
chronos-large

Pretrained forecasting on tokenised time series. Zero-shot on new series.

forecasting
Context
512 ctx
Speed
290 tok/s
Input
$0.023 / 1M rows
Output
$0.055 / 1M rows
Dedicated
1 × H100 · $2.14/hr
amazon/chronos-t5-large
Chronos-Bolt tinyTime series
chronos-bolt-tiny

Faster Chronos variant with direct multi-step decoding.

forecastingfast
Context
2k ctx
Speed
319 tok/s
Input
$0.02 / 1M rows
Output
$0.048 / 1M rows
Dedicated
1 × H100 · $2.14/hr
amazon/chronos-bolt-tiny
Chronos-Bolt miniTime series
chronos-bolt-mini

Faster Chronos variant with direct multi-step decoding.

forecastingfast
Context
2k ctx
Speed
319 tok/s
Input
$0.02 / 1M rows
Output
$0.048 / 1M rows
Dedicated
1 × H100 · $2.14/hr
amazon/chronos-bolt-mini
Chronos-Bolt smallTime series
chronos-bolt-small

Faster Chronos variant with direct multi-step decoding.

forecastingfast
Context
2k ctx
Speed
317 tok/s
Input
$0.02 / 1M rows
Output
$0.048 / 1M rows
Dedicated
1 × H100 · $2.14/hr
amazon/chronos-bolt-small
Chronos-Bolt baseTime series
chronos-bolt-base

Faster Chronos variant with direct multi-step decoding.

forecastingfast
Context
2k ctx
Speed
310 tok/s
Input
$0.021 / 1M rows
Output
$0.05 / 1M rows
Dedicated
1 × H100 · $2.14/hr
amazon/chronos-bolt-base
Moirai smallTime series
moirai-small

Universal forecasting across frequencies and domains.

forecasting
Context
5k ctx
Speed
319 tok/s
Input
$0.02 / 1M rows
Output
$0.048 / 1M rows
Dedicated
1 × H100 · $2.14/hr
Salesforce/moirai-1.1-R-small
Moirai baseTime series
moirai-base

Universal forecasting across frequencies and domains.

forecasting
Context
5k ctx
Speed
315 tok/s
Input
$0.02 / 1M rows
Output
$0.048 / 1M rows
Dedicated
1 × H100 · $2.14/hr
Salesforce/moirai-1.1-R-base
Moirai largeTime series
moirai-large

Universal forecasting across frequencies and domains.

forecasting
Context
5k ctx
Speed
306 tok/s
Input
$0.021 / 1M rows
Output
$0.05 / 1M rows
Dedicated
1 × H100 · $2.14/hr
Salesforce/moirai-1.1-R-large
QwQ 32B FP8Reasoning
qwq-32b-fp8

FP8 serving variant. Reasoning-first Qwen that thinks before answering.

reasoningFP8
Context
128k
Speed
83 tok/s
Input
$0.098 / 1M tokens
Output
$0.235 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/QwQ-32B-FP8
QwQ 32B INT4 AWQReasoning
qwq-32b-awq

INT4 AWQ serving variant. Reasoning-first Qwen that thinks before answering.

reasoningINT4
Context
128k
Speed
116 tok/s
Input
$0.06 / 1M tokens
Output
$0.144 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/QwQ-32B-AWQ
QwQ 32B INT8 GPTQReasoning
qwq-32b-gptq

INT8 GPTQ serving variant. Reasoning-first Qwen that thinks before answering.

reasoningINT8
Context
128k
Speed
97 tok/s
Input
$0.079 / 1M tokens
Output
$0.19 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
Qwen/QwQ-32B-GPTQ
Hermes 3 8BText
hermes-3-8b

Community-tuned Llama with steerable system prompts.

communitytools
Context
128k
Speed
149 tok/s
Input
$0.054 / 1M tokens
Output
$0.13 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
NousResearch/Hermes-3-Llama-3.1-8B
Hermes 3 8B FP8Text
hermes-3-8b-fp8

FP8 serving variant. Community-tuned Llama with steerable system prompts.

communitytoolsFP8
Context
128k
Speed
187 tok/s
Input
$0.033 / 1M tokens
Output
$0.08 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
NousResearch/Hermes-3-Llama-3.1-8B-FP8
Hermes 3 8B INT4 AWQText
hermes-3-8b-awq

INT4 AWQ serving variant. Community-tuned Llama with steerable system prompts.

communitytoolsINT4
Context
128k
Speed
223 tok/s
Input
$0.021 / 1M tokens
Output
$0.049 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
NousResearch/Hermes-3-Llama-3.1-8B-AWQ
Hermes 3 8B INT8 GPTQText
hermes-3-8b-gptq

INT8 GPTQ serving variant. Community-tuned Llama with steerable system prompts.

communitytoolsINT8
Context
128k
Speed
203 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
NousResearch/Hermes-3-Llama-3.1-8B-GPTQ
Hermes 3 70B FP8Text
hermes-3-70b-fp8

FP8 serving variant. Community-tuned Llama with steerable system prompts.

communitytoolsFP8
Context
128k
Speed
44 tok/s
Input
$0.199 / 1M tokens
Output
$0.478 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
NousResearch/Hermes-3-Llama-3.1-70B-FP8
Hermes 3 70B INT4 AWQText
hermes-3-70b-awq

INT4 AWQ serving variant. Community-tuned Llama with steerable system prompts.

communitytoolsINT4
Context
128k
Speed
66 tok/s
Input
$0.122 / 1M tokens
Output
$0.293 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
NousResearch/Hermes-3-Llama-3.1-70B-AWQ
Hermes 3 70B INT8 GPTQText
hermes-3-70b-gptq

INT8 GPTQ serving variant. Community-tuned Llama with steerable system prompts.

communitytoolsINT8
Context
128k
Speed
53 tok/s
Input
$0.161 / 1M tokens
Output
$0.385 / 1M tokens
Dedicated
2 × H100 · $4.28/hr
NousResearch/Hermes-3-Llama-3.1-70B-GPTQ
Zephyr 7B betaText
zephyr-7b-beta

DPO-tuned Mistral. The open alignment reference point.

DPO
Context
32k
Speed
160 tok/s
Input
$0.05 / 1M tokens
Output
$0.12 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceH4/zephyr-7b-beta
Zephyr 7B beta FP8Text
zephyr-7b-beta-fp8

FP8 serving variant. DPO-tuned Mistral. The open alignment reference point.

DPOFP8
Context
32k
Speed
197 tok/s
Input
$0.031 / 1M tokens
Output
$0.074 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceH4/zephyr-7b-beta-FP8
Zephyr 7B beta INT4 AWQText
zephyr-7b-beta-awq

INT4 AWQ serving variant. DPO-tuned Mistral. The open alignment reference point.

DPOINT4
Context
32k
Speed
231 tok/s
Input
$0.019 / 1M tokens
Output
$0.046 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceH4/zephyr-7b-beta-AWQ
Zephyr 7B beta INT8 GPTQText
zephyr-7b-beta-gptq

INT8 GPTQ serving variant. DPO-tuned Mistral. The open alignment reference point.

DPOINT8
Context
32k
Speed
213 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
HuggingFaceH4/zephyr-7b-beta-GPTQ
SOLAR 10.7BText
solar-10-7b

Depth-upscaled 10.7B that outperforms its parameter count.

upscaled
Context
4k
Speed
126 tok/s
Input
$0.066 / 1M tokens
Output
$0.158 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
upstage/SOLAR-10.7B-Instruct-v1.0
SOLAR 10.7B FP8Text
solar-10-7b-fp8

FP8 serving variant. Depth-upscaled 10.7B that outperforms its parameter count.

upscaledFP8
Context
4k
Speed
164 tok/s
Input
$0.041 / 1M tokens
Output
$0.098 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
upstage/SOLAR-10.7B-Instruct-v1.0-FP8
SOLAR 10.7B INT4 AWQText
solar-10-7b-awq

INT4 AWQ serving variant. Depth-upscaled 10.7B that outperforms its parameter count.

upscaledINT4
Context
4k
Speed
202 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
upstage/SOLAR-10.7B-Instruct-v1.0-AWQ
SOLAR 10.7B INT8 GPTQText
solar-10-7b-gptq

INT8 GPTQ serving variant. Depth-upscaled 10.7B that outperforms its parameter count.

upscaledINT8
Context
4k
Speed
181 tok/s
Input
$0.033 / 1M tokens
Output
$0.079 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
upstage/SOLAR-10.7B-Instruct-v1.0-GPTQ
StableLM 2 1.6BText
stablelm-2-1-6b

Compact multilingual models trained on open data.

small
Context
16k
Speed
260 tok/s
Input
$0.027 / 1M tokens
Output
$0.065 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
stabilityai/stablelm-2-1_6b-chat
StableLM 2 12BText
stablelm-2-12b

Compact multilingual models trained on open data.

small
Context
16k
Speed
117 tok/s
Input
$0.072 / 1M tokens
Output
$0.173 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
stabilityai/stablelm-2-12b-chat
StableLM 2 12B FP8Text
stablelm-2-12b-fp8

FP8 serving variant. Compact multilingual models trained on open data.

smallFP8
Context
16k
Speed
155 tok/s
Input
$0.045 / 1M tokens
Output
$0.107 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
stabilityai/stablelm-2-12b-chat-FP8
StableLM 2 12B INT4 AWQText
stablelm-2-12b-awq

INT4 AWQ serving variant. Compact multilingual models trained on open data.

smallINT4
Context
16k
Speed
193 tok/s
Input
$0.027 / 1M tokens
Output
$0.066 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
stabilityai/stablelm-2-12b-chat-AWQ
StableLM 2 12B INT8 GPTQText
stablelm-2-12b-gptq

INT8 GPTQ serving variant. Compact multilingual models trained on open data.

smallINT8
Context
16k
Speed
172 tok/s
Input
$0.036 / 1M tokens
Output
$0.086 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
stabilityai/stablelm-2-12b-chat-GPTQ
TinyLlama 1.1BText
tinyllama-1-1b

1.1B chat model — the floor of the usable size range.

tinyedge
Context
2k
Speed
276 tok/s
Input
$0.025 / 1M tokens
Output
$0.06 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
TinyLlama/TinyLlama-1.1B-Chat-v1.0
Phi-4 mini 3.8BText
phi-4-mini-3-8b

The smallest Phi-4. Reasoning quality at edge cost.

small
Context
128k
Speed
207 tok/s
Input
$0.036 / 1M tokens
Output
$0.086 / 1M tokens
Dedicated
1 × H100 · $2.14/hr
microsoft/Phi-4-mini-instruct
DeepSeek VL2 tinyVision
deepseek-vl2-tiny

Mixture-of-experts vision-language for dense documents.

visionMoE
Context
4k
Speed
280 tok/s
Input
$0.024 / 1M tokens
Output
$0.058 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
deepseek-ai/deepseek-vl2-tiny
DeepSeek VL2 smallVision
deepseek-vl2-small

Mixture-of-experts vision-language for dense documents.

visionMoE
Context
4k
Speed
228 tok/s
Input
$0.032 / 1M tokens
Output
$0.077 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
deepseek-ai/deepseek-vl2-small
DeepSeek VL2 baseVision
deepseek-vl2-base

Mixture-of-experts vision-language for dense documents.

visionMoE
Context
4k
Speed
194 tok/s
Input
$0.039 / 1M tokens
Output
$0.094 / 1M tokens
Dedicated
8 × H200 · $23.84/hr
deepseek-ai/deepseek-vl2