kimi-k3Moonshot's flagship MoE. Long-context agentic work with native tool use.
- Context
- 256k
- Speed
- 48 tok/s
- Input
- $0.58 / 1M tokens
- Output
- $2.2 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
kimi-k21T-parameter MoE with 32B active. Strong tool-calling and code agents.
- Context
- 128k
- Speed
- 56 tok/s
- Input
- $0.44 / 1M tokens
- Output
- $1.75 / 1M tokens
- Dedicated
- 8 × B200 · $54.40/hr
deepseek-v3671B mixture-of-experts with 37B active. Frontier quality at open-weight pricing.
- Context
- 128k
- Speed
- 62 tok/s
- Input
- $0.27 / 1M tokens
- Output
- $1.1 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
llama-3.3-70bThe default general-purpose workhorse. Broad tool-calling support.
- Context
- 128k
- Speed
- 94 tok/s
- Input
- $0.23 / 1M tokens
- Output
- $0.4 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2.5-72bStrong multilingual and structured-output behavior.
- Context
- 128k
- Speed
- 88 tok/s
- Input
- $0.35 / 1M tokens
- Output
- $0.4 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
mixtral-8x22bSparse MoE with fast decode and permissive licensing.
- Context
- 64k
- Speed
- 77 tok/s
- Input
- $0.65 / 1M tokens
- Output
- $0.65 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
phi-414B punching well above weight on reasoning benchmarks. Cheapest per token.
- Context
- 16k
- Speed
- 210 tok/s
- Input
- $0.07 / 1M tokens
- Output
- $0.14 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-2-27bEfficient mid-size instruct model with a permissive license.
- Context
- 8k
- Speed
- 148 tok/s
- Input
- $0.11 / 1M tokens
- Output
- $0.22 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
kimi-k2-thinking1T-parameter MoE with 32B active. Interleaves reasoning with tool calls across long horizons. Bill by output — traces count.
- Context
- 256k
- Speed
- 44 tok/s
- Input
- $0.52 / 1M tokens
- Output
- $2.4 / 1M tokens
- Dedicated
- 8 × B200 · $54.40/hr
r1-distill-32bEmits explicit reasoning traces before the answer. Bill by output — traces count.
- Context
- 64k
- Speed
- 58 tok/s
- Input
- $0.18 / 1M tokens
- Output
- $0.72 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwq-32bLong-horizon deliberate reasoning for maths and planning.
- Context
- 32k
- Speed
- 61 tok/s
- Input
- $0.2 / 1M tokens
- Output
- $0.6 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2.5-coder-32bFill-in-middle and repo-scale completion. The default for editor integrations.
- Context
- 128k
- Speed
- 128 tok/s
- Input
- $0.16 / 1M tokens
- Output
- $0.16 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-coder-v2236B MoE trained for agentic software tasks.
- Context
- 128k
- Speed
- 70 tok/s
- Input
- $0.28 / 1M tokens
- Output
- $1.12 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
starcoder2-15bPermissively licensed base model for completion at volume.
- Context
- 16k
- Speed
- 236 tok/s
- Input
- $0.06 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
openhands-lm-32bTuned for multi-step software agents — planning, editing, running tests.
- Context
- 128k
- Speed
- 66 tok/s
- Input
- $0.24 / 1M tokens
- Output
- $0.96 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
hermes-3-70bReliable structured tool-calling and function schemas for orchestration.
- Context
- 128k
- Speed
- 84 tok/s
- Input
- $0.3 / 1M tokens
- Output
- $0.45 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
watt-tool-70bSpecialized for API selection and argument construction in long tool chains.
- Context
- 128k
- Speed
- 80 tok/s
- Input
- $0.32 / 1M tokens
- Output
- $0.48 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-vl-72bDocument VQA, charts and screen understanding at high resolution.
- Context
- 128k
- Speed
- 44 tok/s
- Input
- $0.55 / 1M tokens
- Output
- $0.55 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
pixtral-12bFast, cheap captioning and multimodal classification.
- Context
- 128k
- Speed
- 132 tok/s
- Input
- $0.12 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-76bStrongest open VLM on fine-grained perception benchmarks.
- Context
- 32k
- Speed
- 38 tok/s
- Input
- $0.62 / 1M tokens
- Output
- $0.62 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
got-ocr2End-to-end OCR emitting Markdown, LaTeX and tables from scans.
- Context
- —
- Speed
- 1.4 s/page
- Input
- $0.004 / page
- Output
- $0.004 / page
- Dedicated
- 1 × H100 · $2.14/hr
surya-ocrLine-level detection and reading order across 90+ languages.
- Context
- —
- Speed
- 0.6 s/page
- Input
- $0.002 / page
- Output
- $0.002 / page
- Dedicated
- 1 × H100 · $2.14/hr
flux-1-devBest-in-class prompt adherence and typography for open image models.
- Context
- —
- Speed
- 2.1 s @ 1024²
- Input
- $0.025 / image
- Output
- $0.025 / image
- Dedicated
- 8 × H200 · $23.84/hr
sd-3.5-largeStrong photographic realism, permissive community license.
- Context
- —
- Speed
- 1.7 s @ 1024²
- Input
- $0.018 / image
- Output
- $0.018 / image
- Dedicated
- 1 × H100 · $2.14/hr
kolorsFast bilingual generation with ControlNet and IP-Adapter support.
- Context
- —
- Speed
- 0.9 s @ 1024²
- Input
- $0.009 / image
- Output
- $0.009 / image
- Dedicated
- 1 × H100 · $2.14/hr
hunyuan-video13B text-to-video with strong motion coherence.
- Context
- —
- Speed
- 48 s @ 5 s clip
- Input
- $0.42 / second of video
- Output
- $0.42 / second of video
- Dedicated
- 1 × H100 · $2.14/hr
mochi-1Apache-licensed 10B diffusion transformer with high motion fidelity.
- Context
- —
- Speed
- 34 s @ 5 s clip
- Input
- $0.31 / second of video
- Output
- $0.31 / second of video
- Dedicated
- 1 × H100 · $2.14/hr
ltx-videoRealtime-class generation for previews and iteration.
- Context
- —
- Speed
- 4 s @ 5 s clip
- Input
- $0.09 / second of video
- Output
- $0.09 / second of video
- Dedicated
- 1 × H100 · $2.14/hr
cosmos-predict-14bPhysics-aware world simulation. Rolls a future forward from a frame plus an action.
- Context
- —
- Speed
- 62 s @ 5 s rollout
- Input
- $0.88 / second of rollout
- Output
- $0.88 / second of rollout
- Dedicated
- 1 × H100 · $2.14/hr
cosmos-text2world-7bGenerates navigable, physically plausible environments from a description.
- Context
- —
- Speed
- 38 s @ 5 s rollout
- Input
- $0.54 / second of rollout
- Output
- $0.54 / second of rollout
- Dedicated
- 1 × H100 · $2.14/hr
cosmos-tokenizerContinuous video tokeniser used to condition world models and policies.
- Context
- —
- Speed
- 310× realtime
- Input
- $0.02 / second of video
- Output
- $0.02 / second of video
- Dedicated
- 1 × H100 · $2.14/hr
hunyuan3d-2Image or text to textured mesh, production-ready topology.
- Context
- —
- Speed
- 26 s/asset
- Input
- $0.19 / asset
- Output
- $0.19 / asset
- Dedicated
- 1 × H100 · $2.14/hr
trellisStructured latents producing radiance fields, meshes and Gaussians from one pass.
- Context
- —
- Speed
- 18 s/asset
- Input
- $0.14 / asset
- Output
- $0.14 / asset
- Dedicated
- 1 × H100 · $2.14/hr
whisper-large-v3Multilingual transcription with word-level timestamps.
- Context
- —
- Speed
- 142× realtime
- Input
- $0.0035 / minute of audio
- Output
- $0.0035 / minute of audio
- Dedicated
- 1 × H100 · $2.14/hr
parakeet-tdtStreaming English ASR with the lowest cost per minute in the catalog.
- Context
- —
- Speed
- 310× realtime
- Input
- $0.0016 / minute of audio
- Output
- $0.0016 / minute of audio
- Dedicated
- 1 × H100 · $2.14/hr
xtts-v2Voice cloning and multilingual synthesis from a 6-second reference.
- Context
- —
- Speed
- 88× realtime
- Input
- $0.012 / 1k characters
- Output
- $0.012 / 1k characters
- Dedicated
- 1 × H100 · $2.14/hr
seamless-m4t-v2Speech-to-speech and speech-to-text translation across 100 languages.
- Context
- —
- Speed
- 46× realtime
- Input
- $0.008 / minute of audio
- Output
- $0.008 / minute of audio
- Dedicated
- 8 × B200 · $54.40/hr
musicgen-largeText-conditioned music generation with optional melody conditioning.
- Context
- —
- Speed
- 11× realtime
- Input
- $0.021 / second of audio
- Output
- $0.021 / second of audio
- Dedicated
- 1 × H100 · $2.14/hr
stable-audio-openSound effects and loops up to 47 seconds at 44.1 kHz stereo.
- Context
- —
- Speed
- 19× realtime
- Input
- $0.014 / second of audio
- Output
- $0.014 / second of audio
- Dedicated
- 1 × H100 · $2.14/hr
bge-m3Dense, sparse and multi-vector retrieval from one pass. 100+ languages.
- Context
- 8k
- Speed
- 14.8k emb/s
- Input
- $0.012 / 1M tokens
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
e5-mistral-7bInstruction-tuned embeddings for task-specific retrieval.
- Context
- 32k
- Speed
- 3.9k emb/s
- Input
- $0.028 / 1M tokens
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
jina-clip-v2Shared text and image embedding space for multimodal search.
- Context
- 8k
- Speed
- 9.1k emb/s
- Input
- $0.018 / 1M tokens
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
bge-reranker-v2-m3Cross-encoder reranking. The cheapest large accuracy win over pure vector search.
- Context
- 8k
- Speed
- 6.2k pair/s
- Input
- $0.014 / 1M tokens
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
jina-reranker-v2Multilingual reranking with strong function-calling and code retrieval.
- Context
- 8k
- Speed
- 5.4k pair/s
- Input
- $0.016 / 1M tokens
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
tabpfn-v2In-context tabular prediction — no training step, fit in a single forward pass.
- Context
- —
- Speed
- 148k row/s
- Input
- $0.0009 / 1k rows
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
autogluon-tabularStacked ensemble AutoML over GBDT and neural baselines.
- Context
- —
- Speed
- 92k row/s
- Input
- $0.0014 / 1k rows
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
chronos-boltZero-shot probabilistic forecasting. No per-series fitting.
- Context
- 2k steps
- Speed
- 44k series/s
- Input
- $0.0011 / 1k forecasts
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
timesfm-2Decoder-only forecasting foundation model with long horizons.
- Context
- 2k steps
- Speed
- 31k series/s
- Input
- $0.0015 / 1k forecasts
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
llama-guard-3-8bInput and output classification against a configurable hazard taxonomy.
- Context
- 8k
- Speed
- 184 tok/s
- Input
- $0.04 / 1M tokens
- Output
- $0.04 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
shieldgemma-9bPolicy-conditioned safety scoring with calibrated probabilities.
- Context
- 8k
- Speed
- 152 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
esm2-15bProtein language model embeddings for structure and function prediction.
- Context
- —
- Speed
- 260 seq/s
- Input
- $0.09 / 1k sequences
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
boltz-2Open biomolecular structure and binding-affinity prediction.
- Context
- —
- Speed
- 41 seq/s
- Input
- $0.42 / 1k sequences
- Output
- —
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-1-8bMeta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-1-8b-fp8FP8 serving variant. Meta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-1-8b-awqINT4 AWQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-1-8b-gptqINT8 GPTQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-1-70bMeta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 29 tok/s
- Input
- $0.321 / 1M tokens
- Output
- $0.77 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
llama-3-1-70b-fp8FP8 serving variant. Meta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 44 tok/s
- Input
- $0.199 / 1M tokens
- Output
- $0.478 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
llama-3-1-70b-awqINT4 AWQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 66 tok/s
- Input
- $0.122 / 1M tokens
- Output
- $0.293 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
llama-3-1-70b-gptqINT8 GPTQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 53 tok/s
- Input
- $0.161 / 1M tokens
- Output
- $0.385 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
llama-3-1-405bMeta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 8 tok/s
- Input
- $1.762 / 1M tokens
- Output
- $4.229 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
llama-3-1-405b-fp8FP8 serving variant. Meta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 8 tok/s
- Input
- $1.092 / 1M tokens
- Output
- $2.622 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
llama-3-1-405b-awqINT4 AWQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 13 tok/s
- Input
- $0.67 / 1M tokens
- Output
- $1.607 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
llama-3-1-405b-gptqINT8 GPTQ serving variant. Meta's general-purpose instruct family. Broad tool-calling support.
- Context
- 128k
- Speed
- 10 tok/s
- Input
- $0.881 / 1M tokens
- Output
- $2.114 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
llama-3-2-1bSmall-footprint Llama for edge and high-throughput serving.
- Context
- 128k
- Speed
- 280 tok/s
- Input
- $0.024 / 1M tokens
- Output
- $0.058 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-2-3bSmall-footprint Llama for edge and high-throughput serving.
- Context
- 128k
- Speed
- 224 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-2-vision-11bLlama with an image encoder for document and chart understanding.
- Context
- 128k
- Speed
- 124 tok/s
- Input
- $0.067 / 1M tokens
- Output
- $0.161 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-2-vision-11b-fp8FP8 serving variant. Llama with an image encoder for document and chart understanding.
- Context
- 128k
- Speed
- 162 tok/s
- Input
- $0.042 / 1M tokens
- Output
- $0.1 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-2-vision-11b-awqINT4 AWQ serving variant. Llama with an image encoder for document and chart understanding.
- Context
- 128k
- Speed
- 200 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.061 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-2-vision-11b-gptqINT8 GPTQ serving variant. Llama with an image encoder for document and chart understanding.
- Context
- 128k
- Speed
- 179 tok/s
- Input
- $0.034 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-3-2-vision-90bLlama with an image encoder for document and chart understanding.
- Context
- 128k
- Speed
- 23 tok/s
- Input
- $0.407 / 1M tokens
- Output
- $0.977 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
llama-3-2-vision-90b-fp8FP8 serving variant. Llama with an image encoder for document and chart understanding.
- Context
- 128k
- Speed
- 35 tok/s
- Input
- $0.252 / 1M tokens
- Output
- $0.606 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
llama-3-2-vision-90b-awqINT4 AWQ serving variant. Llama with an image encoder for document and chart understanding.
- Context
- 128k
- Speed
- 54 tok/s
- Input
- $0.155 / 1M tokens
- Output
- $0.371 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
llama-3-2-vision-90b-gptqINT8 GPTQ serving variant. Llama with an image encoder for document and chart understanding.
- Context
- 128k
- Speed
- 43 tok/s
- Input
- $0.203 / 1M tokens
- Output
- $0.488 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
qwen2-5-0-5bStrong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 298 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.053 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-1-5bStrong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 263 tok/s
- Input
- $0.026 / 1M tokens
- Output
- $0.062 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-3bStrong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 224 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-7bStrong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-7b-fp8FP8 serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-7b-awqINT4 AWQ serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-7b-gptqINT8 GPTQ serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-14bStrong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 106 tok/s
- Input
- $0.08 / 1M tokens
- Output
- $0.192 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-14b-fp8FP8 serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 142 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.119 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-14b-awqINT4 AWQ serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 181 tok/s
- Input
- $0.03 / 1M tokens
- Output
- $0.073 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-14b-gptqINT8 GPTQ serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.04 / 1M tokens
- Output
- $0.096 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-32bStrong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 57 tok/s
- Input
- $0.158 / 1M tokens
- Output
- $0.379 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-32b-fp8FP8 serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 83 tok/s
- Input
- $0.098 / 1M tokens
- Output
- $0.235 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-32b-awqINT4 AWQ serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 116 tok/s
- Input
- $0.06 / 1M tokens
- Output
- $0.144 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-32b-gptqINT8 GPTQ serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 97 tok/s
- Input
- $0.079 / 1M tokens
- Output
- $0.19 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-72bStrong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 28 tok/s
- Input
- $0.33 / 1M tokens
- Output
- $0.792 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-72b-fp8FP8 serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 43 tok/s
- Input
- $0.205 / 1M tokens
- Output
- $0.491 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-72b-awqINT4 AWQ serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 65 tok/s
- Input
- $0.125 / 1M tokens
- Output
- $0.301 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-72b-gptqINT8 GPTQ serving variant. Strong multilingual and structured-output behavior across the size range.
- Context
- 128k
- Speed
- 52 tok/s
- Input
- $0.165 / 1M tokens
- Output
- $0.396 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen3-0-6bHybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 294 tok/s
- Input
- $0.023 / 1M tokens
- Output
- $0.055 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-1-7bHybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 257 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-4bHybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.037 / 1M tokens
- Output
- $0.089 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-8bHybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-8b-fp8FP8 serving variant. Hybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-8b-awqINT4 AWQ serving variant. Hybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-8b-gptqINT8 GPTQ serving variant. Hybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-14bHybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 106 tok/s
- Input
- $0.08 / 1M tokens
- Output
- $0.192 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-14b-fp8FP8 serving variant. Hybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 142 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.119 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-14b-awqINT4 AWQ serving variant. Hybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 181 tok/s
- Input
- $0.03 / 1M tokens
- Output
- $0.073 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-14b-gptqINT8 GPTQ serving variant. Hybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.04 / 1M tokens
- Output
- $0.096 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-32bHybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 57 tok/s
- Input
- $0.158 / 1M tokens
- Output
- $0.379 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-32b-fp8FP8 serving variant. Hybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 83 tok/s
- Input
- $0.098 / 1M tokens
- Output
- $0.235 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-32b-awqINT4 AWQ serving variant. Hybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 116 tok/s
- Input
- $0.06 / 1M tokens
- Output
- $0.144 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-32b-gptqINT8 GPTQ serving variant. Hybrid reasoning family with a switchable thinking mode.
- Context
- 128k
- Speed
- 97 tok/s
- Input
- $0.079 / 1M tokens
- Output
- $0.19 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-moe-30b-a3bSparse Qwen3 — large total parameters, small active set per token.
- Context
- 128k
- Speed
- 224 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen3-moe-235b-a22bSparse Qwen3 — large total parameters, small active set per token.
- Context
- 128k
- Speed
- 77 tok/s
- Input
- $0.115 / 1M tokens
- Output
- $0.276 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
qwen3-moe-235b-a22b-fp8FP8 serving variant. Sparse Qwen3 — large total parameters, small active set per token.
- Context
- 128k
- Speed
- 108 tok/s
- Input
- $0.071 / 1M tokens
- Output
- $0.171 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
qwen3-moe-235b-a22b-awqINT4 AWQ serving variant. Sparse Qwen3 — large total parameters, small active set per token.
- Context
- 128k
- Speed
- 145 tok/s
- Input
- $0.044 / 1M tokens
- Output
- $0.105 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
qwen3-moe-235b-a22b-gptqINT8 GPTQ serving variant. Sparse Qwen3 — large total parameters, small active set per token.
- Context
- 128k
- Speed
- 124 tok/s
- Input
- $0.058 / 1M tokens
- Output
- $0.138 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
qwen2-5-coder-0-5bCode-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 298 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.053 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-1-5bCode-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 263 tok/s
- Input
- $0.026 / 1M tokens
- Output
- $0.062 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-3bCode-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 224 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-7bCode-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-7b-fp8FP8 serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-7b-awqINT4 AWQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-7b-gptqINT8 GPTQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-14bCode-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 106 tok/s
- Input
- $0.08 / 1M tokens
- Output
- $0.192 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-14b-fp8FP8 serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 142 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.119 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-14b-awqINT4 AWQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 181 tok/s
- Input
- $0.03 / 1M tokens
- Output
- $0.073 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-14b-gptqINT8 GPTQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.04 / 1M tokens
- Output
- $0.096 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-32bCode-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 57 tok/s
- Input
- $0.158 / 1M tokens
- Output
- $0.379 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-32b-fp8FP8 serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 83 tok/s
- Input
- $0.098 / 1M tokens
- Output
- $0.235 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-32b-awqINT4 AWQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 116 tok/s
- Input
- $0.06 / 1M tokens
- Output
- $0.144 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-coder-32b-gptqINT8 GPTQ serving variant. Code-specialized Qwen. Fill-in-the-middle and repository-scale context.
- Context
- 128k
- Speed
- 97 tok/s
- Input
- $0.079 / 1M tokens
- Output
- $0.19 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-math-1-5bMathematical reasoning with chain-of-thought and tool-integrated solving.
- Context
- 4k
- Speed
- 263 tok/s
- Input
- $0.026 / 1M tokens
- Output
- $0.062 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-math-7bMathematical reasoning with chain-of-thought and tool-integrated solving.
- Context
- 4k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-math-7b-fp8FP8 serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.
- Context
- 4k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-math-7b-awqINT4 AWQ serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.
- Context
- 4k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-math-7b-gptqINT8 GPTQ serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.
- Context
- 4k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-math-72bMathematical reasoning with chain-of-thought and tool-integrated solving.
- Context
- 4k
- Speed
- 28 tok/s
- Input
- $0.33 / 1M tokens
- Output
- $0.792 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-math-72b-fp8FP8 serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.
- Context
- 4k
- Speed
- 43 tok/s
- Input
- $0.205 / 1M tokens
- Output
- $0.491 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-math-72b-awqINT4 AWQ serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.
- Context
- 4k
- Speed
- 65 tok/s
- Input
- $0.125 / 1M tokens
- Output
- $0.301 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-math-72b-gptqINT8 GPTQ serving variant. Mathematical reasoning with chain-of-thought and tool-integrated solving.
- Context
- 4k
- Speed
- 52 tok/s
- Input
- $0.165 / 1M tokens
- Output
- $0.396 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-vl-3bVision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 224 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-vl-7bVision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-vl-7b-fp8FP8 serving variant. Vision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-vl-7b-awqINT4 AWQ serving variant. Vision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-vl-7b-gptqINT8 GPTQ serving variant. Vision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-vl-32bVision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 57 tok/s
- Input
- $0.158 / 1M tokens
- Output
- $0.379 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-vl-32b-fp8FP8 serving variant. Vision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 83 tok/s
- Input
- $0.098 / 1M tokens
- Output
- $0.235 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-vl-32b-awqINT4 AWQ serving variant. Vision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 116 tok/s
- Input
- $0.06 / 1M tokens
- Output
- $0.144 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-vl-32b-gptqINT8 GPTQ serving variant. Vision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 97 tok/s
- Input
- $0.079 / 1M tokens
- Output
- $0.19 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwen2-5-vl-72bVision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 28 tok/s
- Input
- $0.33 / 1M tokens
- Output
- $0.792 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-vl-72b-fp8FP8 serving variant. Vision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 43 tok/s
- Input
- $0.205 / 1M tokens
- Output
- $0.491 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-vl-72b-awqINT4 AWQ serving variant. Vision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 65 tok/s
- Input
- $0.125 / 1M tokens
- Output
- $0.301 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
qwen2-5-vl-72b-gptqINT8 GPTQ serving variant. Vision-language Qwen with document parsing and grounding.
- Context
- 128k
- Speed
- 52 tok/s
- Input
- $0.165 / 1M tokens
- Output
- $0.396 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
deepseek-r1-distill-qwen-1-5bReasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 263 tok/s
- Input
- $0.026 / 1M tokens
- Output
- $0.062 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-7bReasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-7b-fp8FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-7b-awqINT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-7b-gptqINT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-14bReasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 106 tok/s
- Input
- $0.08 / 1M tokens
- Output
- $0.192 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-14b-fp8FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 142 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.119 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-14b-awqINT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 181 tok/s
- Input
- $0.03 / 1M tokens
- Output
- $0.073 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-14b-gptqINT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.04 / 1M tokens
- Output
- $0.096 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-32bReasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 57 tok/s
- Input
- $0.158 / 1M tokens
- Output
- $0.379 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-32b-fp8FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 83 tok/s
- Input
- $0.098 / 1M tokens
- Output
- $0.235 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-32b-awqINT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 116 tok/s
- Input
- $0.06 / 1M tokens
- Output
- $0.144 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-qwen-32b-gptqINT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 97 tok/s
- Input
- $0.079 / 1M tokens
- Output
- $0.19 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-llama-8bReasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-llama-8b-fp8FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-llama-8b-awqINT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-llama-8b-gptqINT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-r1-distill-llama-70bReasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 29 tok/s
- Input
- $0.321 / 1M tokens
- Output
- $0.77 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
deepseek-r1-distill-llama-70b-fp8FP8 serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 44 tok/s
- Input
- $0.199 / 1M tokens
- Output
- $0.478 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
deepseek-r1-distill-llama-70b-awqINT4 AWQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 66 tok/s
- Input
- $0.122 / 1M tokens
- Output
- $0.293 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
deepseek-r1-distill-llama-70b-gptqINT8 GPTQ serving variant. Reasoning traces from R1 distilled into smaller dense backbones.
- Context
- 128k
- Speed
- 53 tok/s
- Input
- $0.161 / 1M tokens
- Output
- $0.385 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
deepseek-coder-v2-liteRepository-scale code MoE with strong completion and repair.
- Context
- 128k
- Speed
- 238 tok/s
- Input
- $0.03 / 1M tokens
- Output
- $0.072 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
deepseek-coder-v2-236bRepository-scale code MoE with strong completion and repair.
- Context
- 128k
- Speed
- 80 tok/s
- Input
- $0.11 / 1M tokens
- Output
- $0.264 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
deepseek-coder-v2-236b-fp8FP8 serving variant. Repository-scale code MoE with strong completion and repair.
- Context
- 128k
- Speed
- 111 tok/s
- Input
- $0.068 / 1M tokens
- Output
- $0.164 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
deepseek-coder-v2-236b-awqINT4 AWQ serving variant. Repository-scale code MoE with strong completion and repair.
- Context
- 128k
- Speed
- 149 tok/s
- Input
- $0.042 / 1M tokens
- Output
- $0.1 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
deepseek-coder-v2-236b-gptqINT8 GPTQ serving variant. Repository-scale code MoE with strong completion and repair.
- Context
- 128k
- Speed
- 128 tok/s
- Input
- $0.055 / 1M tokens
- Output
- $0.132 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
gemma-2-2bGoogle's open dense family. Efficient at small and mid scale.
- Context
- 8k
- Speed
- 248 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.07 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-2-9bGoogle's open dense family. Efficient at small and mid scale.
- Context
- 8k
- Speed
- 140 tok/s
- Input
- $0.059 / 1M tokens
- Output
- $0.142 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-2-9b-fp8FP8 serving variant. Google's open dense family. Efficient at small and mid scale.
- Context
- 8k
- Speed
- 178 tok/s
- Input
- $0.037 / 1M tokens
- Output
- $0.088 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-2-9b-awqINT4 AWQ serving variant. Google's open dense family. Efficient at small and mid scale.
- Context
- 8k
- Speed
- 214 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.054 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-2-9b-gptqINT8 GPTQ serving variant. Google's open dense family. Efficient at small and mid scale.
- Context
- 8k
- Speed
- 194 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.071 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-2-27b-fp8FP8 serving variant. Google's open dense family. Efficient at small and mid scale.
- Context
- 8k
- Speed
- 94 tok/s
- Input
- $0.084 / 1M tokens
- Output
- $0.202 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-2-27b-awqINT4 AWQ serving variant. Google's open dense family. Efficient at small and mid scale.
- Context
- 8k
- Speed
- 129 tok/s
- Input
- $0.052 / 1M tokens
- Output
- $0.124 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-2-27b-gptqINT8 GPTQ serving variant. Google's open dense family. Efficient at small and mid scale.
- Context
- 8k
- Speed
- 109 tok/s
- Input
- $0.068 / 1M tokens
- Output
- $0.163 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-1bMultimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 280 tok/s
- Input
- $0.024 / 1M tokens
- Output
- $0.058 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-4bMultimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.037 / 1M tokens
- Output
- $0.089 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-12bMultimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 117 tok/s
- Input
- $0.072 / 1M tokens
- Output
- $0.173 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-12b-fp8FP8 serving variant. Multimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 155 tok/s
- Input
- $0.045 / 1M tokens
- Output
- $0.107 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-12b-awqINT4 AWQ serving variant. Multimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 193 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.066 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-12b-gptqINT8 GPTQ serving variant. Multimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 172 tok/s
- Input
- $0.036 / 1M tokens
- Output
- $0.086 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-27bMultimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 65 tok/s
- Input
- $0.136 / 1M tokens
- Output
- $0.326 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-27b-fp8FP8 serving variant. Multimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 94 tok/s
- Input
- $0.084 / 1M tokens
- Output
- $0.202 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-27b-awqINT4 AWQ serving variant. Multimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 129 tok/s
- Input
- $0.052 / 1M tokens
- Output
- $0.124 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gemma-3-27b-gptqINT8 GPTQ serving variant. Multimodal Gemma with a long context window and wide language coverage.
- Context
- 128k
- Speed
- 109 tok/s
- Input
- $0.068 / 1M tokens
- Output
- $0.163 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-7b-v0-3Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-7b-v0-3-fp8FP8 serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-7b-v0-3-awqINT4 AWQ serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-7b-v0-3-gptqINT8 GPTQ serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-nemo-12bCompact, permissively licensed dense models.
- Context
- 32k
- Speed
- 117 tok/s
- Input
- $0.072 / 1M tokens
- Output
- $0.173 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-nemo-12b-fp8FP8 serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 155 tok/s
- Input
- $0.045 / 1M tokens
- Output
- $0.107 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-nemo-12b-awqINT4 AWQ serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 193 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.066 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-nemo-12b-gptqINT8 GPTQ serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 172 tok/s
- Input
- $0.036 / 1M tokens
- Output
- $0.086 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-small-24bCompact, permissively licensed dense models.
- Context
- 32k
- Speed
- 72 tok/s
- Input
- $0.123 / 1M tokens
- Output
- $0.295 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-small-24b-fp8FP8 serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 102 tok/s
- Input
- $0.076 / 1M tokens
- Output
- $0.183 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-small-24b-awqINT4 AWQ serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 138 tok/s
- Input
- $0.047 / 1M tokens
- Output
- $0.112 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-small-24b-gptqINT8 GPTQ serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 117 tok/s
- Input
- $0.061 / 1M tokens
- Output
- $0.148 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mistral-large-123bCompact, permissively licensed dense models.
- Context
- 32k
- Speed
- 17 tok/s
- Input
- $0.549 / 1M tokens
- Output
- $1.318 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
mistral-large-123b-fp8FP8 serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 26 tok/s
- Input
- $0.34 / 1M tokens
- Output
- $0.817 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
mistral-large-123b-awqINT4 AWQ serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 41 tok/s
- Input
- $0.209 / 1M tokens
- Output
- $0.501 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
mistral-large-123b-gptqINT8 GPTQ serving variant. Compact, permissively licensed dense models.
- Context
- 32k
- Speed
- 32 tok/s
- Input
- $0.275 / 1M tokens
- Output
- $0.659 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
ministral-3bEdge-class Mistral for on-device and high-throughput serving.
- Context
- 128k
- Speed
- 224 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
ministral-8bEdge-class Mistral for on-device and high-throughput serving.
- Context
- 128k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
ministral-8b-fp8FP8 serving variant. Edge-class Mistral for on-device and high-throughput serving.
- Context
- 128k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
ministral-8b-awqINT4 AWQ serving variant. Edge-class Mistral for on-device and high-throughput serving.
- Context
- 128k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
ministral-8b-gptqINT8 GPTQ serving variant. Edge-class Mistral for on-device and high-throughput serving.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codestral-22bMistral's code model with fill-in-the-middle across 80+ languages.
- Context
- 32k
- Speed
- 77 tok/s
- Input
- $0.115 / 1M tokens
- Output
- $0.276 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codestral-22b-fp8FP8 serving variant. Mistral's code model with fill-in-the-middle across 80+ languages.
- Context
- 32k
- Speed
- 108 tok/s
- Input
- $0.071 / 1M tokens
- Output
- $0.171 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codestral-22b-awqINT4 AWQ serving variant. Mistral's code model with fill-in-the-middle across 80+ languages.
- Context
- 32k
- Speed
- 145 tok/s
- Input
- $0.044 / 1M tokens
- Output
- $0.105 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codestral-22b-gptqINT8 GPTQ serving variant. Mistral's code model with fill-in-the-middle across 80+ languages.
- Context
- 32k
- Speed
- 124 tok/s
- Input
- $0.058 / 1M tokens
- Output
- $0.138 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-miniSmall models trained on heavily filtered data. Strong reasoning per parameter.
- Context
- 128k
- Speed
- 207 tok/s
- Input
- $0.036 / 1M tokens
- Output
- $0.086 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-smallSmall models trained on heavily filtered data. Strong reasoning per parameter.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-small-fp8FP8 serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.
- Context
- 128k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-small-awqINT4 AWQ serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.
- Context
- 128k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-small-gptqINT8 GPTQ serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.
- Context
- 128k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-mediumSmall models trained on heavily filtered data. Strong reasoning per parameter.
- Context
- 128k
- Speed
- 106 tok/s
- Input
- $0.08 / 1M tokens
- Output
- $0.192 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-medium-fp8FP8 serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.
- Context
- 128k
- Speed
- 142 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.119 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-medium-awqINT4 AWQ serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.
- Context
- 128k
- Speed
- 181 tok/s
- Input
- $0.03 / 1M tokens
- Output
- $0.073 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-medium-gptqINT8 GPTQ serving variant. Small models trained on heavily filtered data. Strong reasoning per parameter.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.04 / 1M tokens
- Output
- $0.096 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-5-miniRefreshed Phi with a sparse MoE variant and vision sibling.
- Context
- 128k
- Speed
- 207 tok/s
- Input
- $0.036 / 1M tokens
- Output
- $0.086 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-5-moeRefreshed Phi with a sparse MoE variant and vision sibling.
- Context
- 128k
- Speed
- 164 tok/s
- Input
- $0.048 / 1M tokens
- Output
- $0.115 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-3-5-visionRefreshed Phi with a sparse MoE variant and vision sibling.
- Context
- 128k
- Speed
- 200 tok/s
- Input
- $0.038 / 1M tokens
- Output
- $0.091 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
yi-1-5-6bBilingual Chinese/English dense family.
- Context
- 32k
- Speed
- 172 tok/s
- Input
- $0.046 / 1M tokens
- Output
- $0.11 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
yi-1-5-9bBilingual Chinese/English dense family.
- Context
- 32k
- Speed
- 140 tok/s
- Input
- $0.059 / 1M tokens
- Output
- $0.142 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
yi-1-5-9b-fp8FP8 serving variant. Bilingual Chinese/English dense family.
- Context
- 32k
- Speed
- 178 tok/s
- Input
- $0.037 / 1M tokens
- Output
- $0.088 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
yi-1-5-9b-awqINT4 AWQ serving variant. Bilingual Chinese/English dense family.
- Context
- 32k
- Speed
- 214 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.054 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
yi-1-5-9b-gptqINT8 GPTQ serving variant. Bilingual Chinese/English dense family.
- Context
- 32k
- Speed
- 194 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.071 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
yi-1-5-34bBilingual Chinese/English dense family.
- Context
- 32k
- Speed
- 54 tok/s
- Input
- $0.166 / 1M tokens
- Output
- $0.398 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
yi-1-5-34b-fp8FP8 serving variant. Bilingual Chinese/English dense family.
- Context
- 32k
- Speed
- 79 tok/s
- Input
- $0.103 / 1M tokens
- Output
- $0.247 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
yi-1-5-34b-awqINT4 AWQ serving variant. Bilingual Chinese/English dense family.
- Context
- 32k
- Speed
- 112 tok/s
- Input
- $0.063 / 1M tokens
- Output
- $0.151 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
yi-1-5-34b-gptqINT8 GPTQ serving variant. Bilingual Chinese/English dense family.
- Context
- 32k
- Speed
- 93 tok/s
- Input
- $0.083 / 1M tokens
- Output
- $0.199 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
glm-4-9bZhipu's bilingual family with agent tuning.
- Context
- 128k
- Speed
- 140 tok/s
- Input
- $0.059 / 1M tokens
- Output
- $0.142 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
glm-4-9b-fp8FP8 serving variant. Zhipu's bilingual family with agent tuning.
- Context
- 128k
- Speed
- 178 tok/s
- Input
- $0.037 / 1M tokens
- Output
- $0.088 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
glm-4-9b-awqINT4 AWQ serving variant. Zhipu's bilingual family with agent tuning.
- Context
- 128k
- Speed
- 214 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.054 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
glm-4-9b-gptqINT8 GPTQ serving variant. Zhipu's bilingual family with agent tuning.
- Context
- 128k
- Speed
- 194 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.071 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
glm-4-9b-1mZhipu's bilingual family with agent tuning.
- Context
- 128k
- Speed
- 140 tok/s
- Input
- $0.059 / 1M tokens
- Output
- $0.142 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
glm-4-9b-1m-fp8FP8 serving variant. Zhipu's bilingual family with agent tuning.
- Context
- 128k
- Speed
- 178 tok/s
- Input
- $0.037 / 1M tokens
- Output
- $0.088 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
glm-4-9b-1m-awqINT4 AWQ serving variant. Zhipu's bilingual family with agent tuning.
- Context
- 128k
- Speed
- 214 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.054 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
glm-4-9b-1m-gptqINT8 GPTQ serving variant. Zhipu's bilingual family with agent tuning.
- Context
- 128k
- Speed
- 194 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.071 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internlm2-5-1-8bLong-context Chinese/English models with tool-use training.
- Context
- 1M
- Speed
- 254 tok/s
- Input
- $0.028 / 1M tokens
- Output
- $0.067 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internlm2-5-7bLong-context Chinese/English models with tool-use training.
- Context
- 1M
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internlm2-5-7b-fp8FP8 serving variant. Long-context Chinese/English models with tool-use training.
- Context
- 1M
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internlm2-5-7b-awqINT4 AWQ serving variant. Long-context Chinese/English models with tool-use training.
- Context
- 1M
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internlm2-5-7b-gptqINT8 GPTQ serving variant. Long-context Chinese/English models with tool-use training.
- Context
- 1M
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internlm2-5-20bLong-context Chinese/English models with tool-use training.
- Context
- 1M
- Speed
- 82 tok/s
- Input
- $0.106 / 1M tokens
- Output
- $0.254 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internlm2-5-20b-fp8FP8 serving variant. Long-context Chinese/English models with tool-use training.
- Context
- 1M
- Speed
- 115 tok/s
- Input
- $0.066 / 1M tokens
- Output
- $0.158 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internlm2-5-20b-awqINT4 AWQ serving variant. Long-context Chinese/English models with tool-use training.
- Context
- 1M
- Speed
- 153 tok/s
- Input
- $0.04 / 1M tokens
- Output
- $0.097 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internlm2-5-20b-gptqINT8 GPTQ serving variant. Long-context Chinese/English models with tool-use training.
- Context
- 1M
- Speed
- 131 tok/s
- Input
- $0.053 / 1M tokens
- Output
- $0.127 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
minicpm3-4bSmall models tuned for on-device deployment.
- Context
- 32k
- Speed
- 203 tok/s
- Input
- $0.037 / 1M tokens
- Output
- $0.089 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
olmo-2-7bFully open training data, code and checkpoints. The reproducibility baseline.
- Context
- 4k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
olmo-2-7b-fp8FP8 serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.
- Context
- 4k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
olmo-2-7b-awqINT4 AWQ serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.
- Context
- 4k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
olmo-2-7b-gptqINT8 GPTQ serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.
- Context
- 4k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
olmo-2-13bFully open training data, code and checkpoints. The reproducibility baseline.
- Context
- 4k
- Speed
- 112 tok/s
- Input
- $0.076 / 1M tokens
- Output
- $0.182 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
olmo-2-13b-fp8FP8 serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.
- Context
- 4k
- Speed
- 148 tok/s
- Input
- $0.047 / 1M tokens
- Output
- $0.113 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
olmo-2-13b-awqINT4 AWQ serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.
- Context
- 4k
- Speed
- 187 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.069 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
olmo-2-13b-gptqINT8 GPTQ serving variant. Fully open training data, code and checkpoints. The reproducibility baseline.
- Context
- 4k
- Speed
- 165 tok/s
- Input
- $0.038 / 1M tokens
- Output
- $0.091 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-1bTII's efficient dense family.
- Context
- 32k
- Speed
- 280 tok/s
- Input
- $0.024 / 1M tokens
- Output
- $0.058 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-3bTII's efficient dense family.
- Context
- 32k
- Speed
- 224 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-7bTII's efficient dense family.
- Context
- 32k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-7b-fp8FP8 serving variant. TII's efficient dense family.
- Context
- 32k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-7b-awqINT4 AWQ serving variant. TII's efficient dense family.
- Context
- 32k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-7b-gptqINT8 GPTQ serving variant. TII's efficient dense family.
- Context
- 32k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-10bTII's efficient dense family.
- Context
- 32k
- Speed
- 131 tok/s
- Input
- $0.063 / 1M tokens
- Output
- $0.151 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-10b-fp8FP8 serving variant. TII's efficient dense family.
- Context
- 32k
- Speed
- 169 tok/s
- Input
- $0.039 / 1M tokens
- Output
- $0.094 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-10b-awqINT4 AWQ serving variant. TII's efficient dense family.
- Context
- 32k
- Speed
- 207 tok/s
- Input
- $0.024 / 1M tokens
- Output
- $0.057 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
falcon-3-10b-gptqINT8 GPTQ serving variant. TII's efficient dense family.
- Context
- 32k
- Speed
- 186 tok/s
- Input
- $0.032 / 1M tokens
- Output
- $0.076 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
command-r-r-35bRetrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 53 tok/s
- Input
- $0.17 / 1M tokens
- Output
- $0.408 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
command-r-r-35b-fp8FP8 serving variant. Retrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 78 tok/s
- Input
- $0.105 / 1M tokens
- Output
- $0.253 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
command-r-r-35b-awqINT4 AWQ serving variant. Retrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 110 tok/s
- Input
- $0.065 / 1M tokens
- Output
- $0.155 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
command-r-r-35b-gptqINT8 GPTQ serving variant. Retrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 91 tok/s
- Input
- $0.085 / 1M tokens
- Output
- $0.204 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
command-r-r-104bRetrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 20 tok/s
- Input
- $0.467 / 1M tokens
- Output
- $1.121 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
command-r-r-104b-fp8FP8 serving variant. Retrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 31 tok/s
- Input
- $0.29 / 1M tokens
- Output
- $0.695 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
command-r-r-104b-awqINT4 AWQ serving variant. Retrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 48 tok/s
- Input
- $0.177 / 1M tokens
- Output
- $0.426 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
command-r-r-104b-gptqINT8 GPTQ serving variant. Retrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 37 tok/s
- Input
- $0.234 / 1M tokens
- Output
- $0.56 / 1M tokens
- Dedicated
- 4 × H100 · $8.56/hr
command-r-r7bRetrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
command-r-r7b-fp8FP8 serving variant. Retrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
command-r-r7b-awqINT4 AWQ serving variant. Retrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
command-r-r7b-gptqINT8 GPTQ serving variant. Retrieval-augmented generation and citation-grounded answering.
- Context
- 128k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
aya-expanse-8bMassively multilingual coverage across 23 languages.
- Context
- 128k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
aya-expanse-8b-fp8FP8 serving variant. Massively multilingual coverage across 23 languages.
- Context
- 128k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
aya-expanse-8b-awqINT4 AWQ serving variant. Massively multilingual coverage across 23 languages.
- Context
- 128k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
aya-expanse-8b-gptqINT8 GPTQ serving variant. Massively multilingual coverage across 23 languages.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
aya-expanse-32bMassively multilingual coverage across 23 languages.
- Context
- 128k
- Speed
- 57 tok/s
- Input
- $0.158 / 1M tokens
- Output
- $0.379 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
aya-expanse-32b-fp8FP8 serving variant. Massively multilingual coverage across 23 languages.
- Context
- 128k
- Speed
- 83 tok/s
- Input
- $0.098 / 1M tokens
- Output
- $0.235 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
aya-expanse-32b-awqINT4 AWQ serving variant. Massively multilingual coverage across 23 languages.
- Context
- 128k
- Speed
- 116 tok/s
- Input
- $0.06 / 1M tokens
- Output
- $0.144 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
aya-expanse-32b-gptqINT8 GPTQ serving variant. Massively multilingual coverage across 23 languages.
- Context
- 128k
- Speed
- 97 tok/s
- Input
- $0.079 / 1M tokens
- Output
- $0.19 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
granite-3-2bIBM's enterprise family with an explicit indemnified license.
- Context
- 128k
- Speed
- 248 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.07 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
granite-3-8bIBM's enterprise family with an explicit indemnified license.
- Context
- 128k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
granite-3-8b-fp8FP8 serving variant. IBM's enterprise family with an explicit indemnified license.
- Context
- 128k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
granite-3-8b-awqINT4 AWQ serving variant. IBM's enterprise family with an explicit indemnified license.
- Context
- 128k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
granite-3-8b-gptqINT8 GPTQ serving variant. IBM's enterprise family with an explicit indemnified license.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
granite-3-1b-a400mIBM's enterprise family with an explicit indemnified license.
- Context
- 128k
- Speed
- 302 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.053 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
granite-3-3b-a800mIBM's enterprise family with an explicit indemnified license.
- Context
- 128k
- Speed
- 287 tok/s
- Input
- $0.023 / 1M tokens
- Output
- $0.055 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
nemotron-70bNVIDIA's reward-tuned Llama derivatives.
- Context
- 128k
- Speed
- 29 tok/s
- Input
- $0.321 / 1M tokens
- Output
- $0.77 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
nemotron-70b-fp8FP8 serving variant. NVIDIA's reward-tuned Llama derivatives.
- Context
- 128k
- Speed
- 44 tok/s
- Input
- $0.199 / 1M tokens
- Output
- $0.478 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
nemotron-70b-awqINT4 AWQ serving variant. NVIDIA's reward-tuned Llama derivatives.
- Context
- 128k
- Speed
- 66 tok/s
- Input
- $0.122 / 1M tokens
- Output
- $0.293 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
nemotron-70b-gptqINT8 GPTQ serving variant. NVIDIA's reward-tuned Llama derivatives.
- Context
- 128k
- Speed
- 53 tok/s
- Input
- $0.161 / 1M tokens
- Output
- $0.385 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
nemotron-51bNVIDIA's reward-tuned Llama derivatives.
- Context
- 128k
- Speed
- 38 tok/s
- Input
- $0.239 / 1M tokens
- Output
- $0.574 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
nemotron-51b-fp8FP8 serving variant. NVIDIA's reward-tuned Llama derivatives.
- Context
- 128k
- Speed
- 58 tok/s
- Input
- $0.148 / 1M tokens
- Output
- $0.356 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
nemotron-51b-awqINT4 AWQ serving variant. NVIDIA's reward-tuned Llama derivatives.
- Context
- 128k
- Speed
- 84 tok/s
- Input
- $0.091 / 1M tokens
- Output
- $0.218 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
nemotron-51b-gptqINT8 GPTQ serving variant. NVIDIA's reward-tuned Llama derivatives.
- Context
- 128k
- Speed
- 68 tok/s
- Input
- $0.119 / 1M tokens
- Output
- $0.287 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
nemotron-mini-4bNVIDIA's reward-tuned Llama derivatives.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.037 / 1M tokens
- Output
- $0.089 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
smollm2-135mTiny models that still follow instructions. The bottom of the size curve.
- Context
- 8k
- Speed
- 313 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
smollm2-360mTiny models that still follow instructions. The bottom of the size curve.
- Context
- 8k
- Speed
- 304 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.053 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
smollm2-1-7bTiny models that still follow instructions. The bottom of the size curve.
- Context
- 8k
- Speed
- 257 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
starcoder2-3bPermissively trained code models over 600+ languages.
- Context
- 16k
- Speed
- 224 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
starcoder2-7bPermissively trained code models over 600+ languages.
- Context
- 16k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
starcoder2-7b-fp8FP8 serving variant. Permissively trained code models over 600+ languages.
- Context
- 16k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
starcoder2-7b-awqINT4 AWQ serving variant. Permissively trained code models over 600+ languages.
- Context
- 16k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
starcoder2-7b-gptqINT8 GPTQ serving variant. Permissively trained code models over 600+ languages.
- Context
- 16k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
starcoder2-15b-fp8FP8 serving variant. Permissively trained code models over 600+ languages.
- Context
- 16k
- Speed
- 137 tok/s
- Input
- $0.053 / 1M tokens
- Output
- $0.126 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
starcoder2-15b-awqINT4 AWQ serving variant. Permissively trained code models over 600+ languages.
- Context
- 16k
- Speed
- 176 tok/s
- Input
- $0.032 / 1M tokens
- Output
- $0.078 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
starcoder2-15b-gptqINT8 GPTQ serving variant. Permissively trained code models over 600+ languages.
- Context
- 16k
- Speed
- 154 tok/s
- Input
- $0.043 / 1M tokens
- Output
- $0.102 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-7bLong-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-7b-fp8FP8 serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-7b-awqINT4 AWQ serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-7b-gptqINT8 GPTQ serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-13bLong-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 112 tok/s
- Input
- $0.076 / 1M tokens
- Output
- $0.182 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-13b-fp8FP8 serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 148 tok/s
- Input
- $0.047 / 1M tokens
- Output
- $0.113 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-13b-awqINT4 AWQ serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 187 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.069 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-13b-gptqINT8 GPTQ serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 165 tok/s
- Input
- $0.038 / 1M tokens
- Output
- $0.091 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-34bLong-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 54 tok/s
- Input
- $0.166 / 1M tokens
- Output
- $0.398 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-34b-fp8FP8 serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 79 tok/s
- Input
- $0.103 / 1M tokens
- Output
- $0.247 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-34b-awqINT4 AWQ serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 112 tok/s
- Input
- $0.063 / 1M tokens
- Output
- $0.151 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-34b-gptqINT8 GPTQ serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 93 tok/s
- Input
- $0.083 / 1M tokens
- Output
- $0.199 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-70bLong-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 29 tok/s
- Input
- $0.321 / 1M tokens
- Output
- $0.77 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-70b-fp8FP8 serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 44 tok/s
- Input
- $0.199 / 1M tokens
- Output
- $0.478 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-70b-awqINT4 AWQ serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 66 tok/s
- Input
- $0.122 / 1M tokens
- Output
- $0.293 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
codellama-70b-gptqINT8 GPTQ serving variant. Long-standing code baseline with infilling and Python specialisation.
- Context
- 16k
- Speed
- 53 tok/s
- Input
- $0.161 / 1M tokens
- Output
- $0.385 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-1bVision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 280 tok/s
- Input
- $0.024 / 1M tokens
- Output
- $0.058 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-2bVision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 248 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.07 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-4bVision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 203 tok/s
- Input
- $0.037 / 1M tokens
- Output
- $0.089 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-8bVision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-8b-fp8FP8 serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-8b-awqINT4 AWQ serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-8b-gptqINT8 GPTQ serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-26bVision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 67 tok/s
- Input
- $0.132 / 1M tokens
- Output
- $0.317 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-26b-fp8FP8 serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 96 tok/s
- Input
- $0.082 / 1M tokens
- Output
- $0.196 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-26b-awqINT4 AWQ serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 132 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-26b-gptqINT8 GPTQ serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 112 tok/s
- Input
- $0.066 / 1M tokens
- Output
- $0.158 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
internvl2-5-38bVision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 49 tok/s
- Input
- $0.183 / 1M tokens
- Output
- $0.439 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
internvl2-5-38b-fp8FP8 serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 73 tok/s
- Input
- $0.113 / 1M tokens
- Output
- $0.272 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
internvl2-5-38b-awqINT4 AWQ serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 104 tok/s
- Input
- $0.07 / 1M tokens
- Output
- $0.167 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
internvl2-5-38b-gptqINT8 GPTQ serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 86 tok/s
- Input
- $0.091 / 1M tokens
- Output
- $0.22 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
internvl2-5-78bVision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 26 tok/s
- Input
- $0.355 / 1M tokens
- Output
- $0.852 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
internvl2-5-78b-fp8FP8 serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 40 tok/s
- Input
- $0.22 / 1M tokens
- Output
- $0.528 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
internvl2-5-78b-awqINT4 AWQ serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 61 tok/s
- Input
- $0.135 / 1M tokens
- Output
- $0.324 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
internvl2-5-78b-gptqINT8 GPTQ serving variant. Vision-language models with strong chart, table and document reading.
- Context
- 32k
- Speed
- 48 tok/s
- Input
- $0.177 / 1M tokens
- Output
- $0.426 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
llava-1-6-mistral-7bThe open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-mistral-7b-fp8FP8 serving variant. The open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-mistral-7b-awqINT4 AWQ serving variant. The open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-mistral-7b-gptqINT8 GPTQ serving variant. The open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-vicuna-13bThe open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 112 tok/s
- Input
- $0.076 / 1M tokens
- Output
- $0.182 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-vicuna-13b-fp8FP8 serving variant. The open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 148 tok/s
- Input
- $0.047 / 1M tokens
- Output
- $0.113 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-vicuna-13b-awqINT4 AWQ serving variant. The open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 187 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.069 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-vicuna-13b-gptqINT8 GPTQ serving variant. The open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 165 tok/s
- Input
- $0.038 / 1M tokens
- Output
- $0.091 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-34bThe open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 54 tok/s
- Input
- $0.166 / 1M tokens
- Output
- $0.398 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-34b-fp8FP8 serving variant. The open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 79 tok/s
- Input
- $0.103 / 1M tokens
- Output
- $0.247 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-34b-awqINT4 AWQ serving variant. The open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 112 tok/s
- Input
- $0.063 / 1M tokens
- Output
- $0.151 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llava-1-6-34b-gptqINT8 GPTQ serving variant. The open vision-language baseline. Broad ecosystem support.
- Context
- 32k
- Speed
- 93 tok/s
- Input
- $0.083 / 1M tokens
- Output
- $0.199 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
minicpm-v-2-6Small vision-language models that run on a single consumer GPU.
- Context
- 32k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
minicpm-v-2-6-fp8FP8 serving variant. Small vision-language models that run on a single consumer GPU.
- Context
- 32k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
minicpm-v-2-6-awqINT4 AWQ serving variant. Small vision-language models that run on a single consumer GPU.
- Context
- 32k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
minicpm-v-2-6-gptqINT8 GPTQ serving variant. Small vision-language models that run on a single consumer GPU.
- Context
- 32k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
pixtral-12b-fp8FP8 serving variant. Mistral's multimodal model, native variable-resolution image input.
- Context
- 128k
- Speed
- 155 tok/s
- Input
- $0.045 / 1M tokens
- Output
- $0.107 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
pixtral-12b-awqINT4 AWQ serving variant. Mistral's multimodal model, native variable-resolution image input.
- Context
- 128k
- Speed
- 193 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.066 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
pixtral-12b-gptqINT8 GPTQ serving variant. Mistral's multimodal model, native variable-resolution image input.
- Context
- 128k
- Speed
- 172 tok/s
- Input
- $0.036 / 1M tokens
- Output
- $0.086 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
molmo-7b-dOpen vision-language models with pointing and grounding supervision.
- Context
- 32k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
molmo-7b-d-fp8FP8 serving variant. Open vision-language models with pointing and grounding supervision.
- Context
- 32k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
molmo-7b-d-awqINT4 AWQ serving variant. Open vision-language models with pointing and grounding supervision.
- Context
- 32k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
molmo-7b-d-gptqINT8 GPTQ serving variant. Open vision-language models with pointing and grounding supervision.
- Context
- 32k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
molmo-72bOpen vision-language models with pointing and grounding supervision.
- Context
- 32k
- Speed
- 28 tok/s
- Input
- $0.33 / 1M tokens
- Output
- $0.792 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
molmo-72b-fp8FP8 serving variant. Open vision-language models with pointing and grounding supervision.
- Context
- 32k
- Speed
- 43 tok/s
- Input
- $0.205 / 1M tokens
- Output
- $0.491 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
molmo-72b-awqINT4 AWQ serving variant. Open vision-language models with pointing and grounding supervision.
- Context
- 32k
- Speed
- 65 tok/s
- Input
- $0.125 / 1M tokens
- Output
- $0.301 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
molmo-72b-gptqINT8 GPTQ serving variant. Open vision-language models with pointing and grounding supervision.
- Context
- 32k
- Speed
- 52 tok/s
- Input
- $0.165 / 1M tokens
- Output
- $0.396 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
idefics3-8bDocument-heavy multimodal model built on Llama.
- Context
- 16k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
idefics3-8b-fp8FP8 serving variant. Document-heavy multimodal model built on Llama.
- Context
- 16k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
idefics3-8b-awqINT4 AWQ serving variant. Document-heavy multimodal model built on Llama.
- Context
- 16k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
idefics3-8b-gptqINT8 GPTQ serving variant. Document-heavy multimodal model built on Llama.
- Context
- 16k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
got-ocr2-0-6bEnd-to-end OCR covering formulas, tables, sheet music and charts.
- Context
- 8k
- Speed
- 294 tok/s
- Input
- $0.023 / page
- Output
- $0.055 / page
- Dedicated
- 1 × H100 · $2.14/hr
florence-2-baseCompact unified vision model — captioning, detection and segmentation.
- Context
- 4k
- Speed
- 309 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
florence-2-largeCompact unified vision model — captioning, detection and segmentation.
- Context
- 4k
- Speed
- 288 tok/s
- Input
- $0.023 / 1M tokens
- Output
- $0.055 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
bge-m3-m3Multilingual, multi-granularity retrieval embeddings.
- Context
- 8k
- Speed
- 295 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.053 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
bge-smallThe retrieval workhorse. Dense embeddings across three sizes.
- Context
- 512
- Speed
- 318 tok/s
- Input
- $0.02 / 1M tokens
- Output
- $0.048 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
bge-baseThe retrieval workhorse. Dense embeddings across three sizes.
- Context
- 512
- Speed
- 315 tok/s
- Input
- $0.02 / 1M tokens
- Output
- $0.048 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
bge-largeThe retrieval workhorse. Dense embeddings across three sizes.
- Context
- 512
- Speed
- 305 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
bge-reranker-baseCross-encoder reranking — the expensive half of a good retrieval stack.
- Context
- 8k
- Speed
- 307 tok/s
- Input
- $0.021 / 1K queries
- Output
- $0.05 / 1K queries
- Dedicated
- 1 × H100 · $2.14/hr
bge-reranker-largeCross-encoder reranking — the expensive half of a good retrieval stack.
- Context
- 8k
- Speed
- 296 tok/s
- Input
- $0.022 / 1K queries
- Output
- $0.053 / 1K queries
- Dedicated
- 1 × H100 · $2.14/hr
e5-smallInstruction-tuned embeddings with strong zero-shot retrieval.
- Context
- 512
- Speed
- 314 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
e5-baseInstruction-tuned embeddings with strong zero-shot retrieval.
- Context
- 512
- Speed
- 307 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
e5-largeInstruction-tuned embeddings with strong zero-shot retrieval.
- Context
- 512
- Speed
- 296 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.053 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gte-baseLong-context general text embeddings.
- Context
- 8k
- Speed
- 313 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gte-largeLong-context general text embeddings.
- Context
- 8k
- Speed
- 301 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.053 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
gte-qwen2-7bLong-context general text embeddings.
- Context
- 8k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
nomic-embed-v1-5Fully open embeddings with reproducible training data.
- Context
- 8k
- Speed
- 313 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
jina-embeddings-v3Bilingual and code-aware embedding models.
- Context
- 8k
- Speed
- 295 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.053 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
jina-embeddings-v2-base-codeBilingual and code-aware embedding models.
- Context
- 8k
- Speed
- 312 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
mxbai-large-v1Compact embeddings tuned for Matryoshka truncation.
- Context
- 512
- Speed
- 305 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
stella-400m-v5High-ranking compact retrieval embeddings.
- Context
- 8k
- Speed
- 302 tok/s
- Input
- $0.022 / 1M tokens
- Output
- $0.053 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
stella-1-5b-v5High-ranking compact retrieval embeddings.
- Context
- 8k
- Speed
- 263 tok/s
- Input
- $0.026 / 1M tokens
- Output
- $0.062 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
whisper-tinyThe transcription baseline. 99 languages with timestamped output.
- Context
- 30s window
- Speed
- 318 tok/s
- Input
- $0.02 / audio min
- Output
- $0.048 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
whisper-baseThe transcription baseline. 99 languages with timestamped output.
- Context
- 30s window
- Speed
- 316 tok/s
- Input
- $0.02 / audio min
- Output
- $0.048 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
whisper-smallThe transcription baseline. 99 languages with timestamped output.
- Context
- 30s window
- Speed
- 309 tok/s
- Input
- $0.021 / audio min
- Output
- $0.05 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
whisper-mediumThe transcription baseline. 99 languages with timestamped output.
- Context
- 30s window
- Speed
- 288 tok/s
- Input
- $0.023 / audio min
- Output
- $0.055 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
whisper-large-v3-turboThe transcription baseline. 99 languages with timestamped output.
- Context
- 30s window
- Speed
- 286 tok/s
- Input
- $0.023 / audio min
- Output
- $0.055 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
distil-whisper-small-enDistilled Whisper — most of the accuracy at a fraction of the decode cost.
- Context
- 30s window
- Speed
- 312 tok/s
- Input
- $0.021 / audio min
- Output
- $0.05 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
distil-whisper-medium-enDistilled Whisper — most of the accuracy at a fraction of the decode cost.
- Context
- 30s window
- Speed
- 302 tok/s
- Input
- $0.022 / audio min
- Output
- $0.053 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
distil-whisper-large-v3Distilled Whisper — most of the accuracy at a fraction of the decode cost.
- Context
- 30s window
- Speed
- 288 tok/s
- Input
- $0.023 / audio min
- Output
- $0.055 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
parakeet-tdt-0-6bStreaming ASR built for realtime transcription.
- Context
- streaming
- Speed
- 294 tok/s
- Input
- $0.023 / audio min
- Output
- $0.055 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
parakeet-ctc-1-1bStreaming ASR built for realtime transcription.
- Context
- streaming
- Speed
- 276 tok/s
- Input
- $0.025 / audio min
- Output
- $0.06 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
canary-1bMultilingual ASR with speech translation in the same pass.
- Context
- streaming
- Speed
- 280 tok/s
- Input
- $0.024 / audio min
- Output
- $0.058 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
canary-180m-flashMultilingual ASR with speech translation in the same pass.
- Context
- streaming
- Speed
- 311 tok/s
- Input
- $0.021 / audio min
- Output
- $0.05 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
seamlessm4t-v2-largeSpeech-to-speech and speech-to-text translation across 100 languages.
- Context
- streaming
- Speed
- 240 tok/s
- Input
- $0.03 / audio min
- Output
- $0.072 / audio min
- Dedicated
- 8 × B200 · $54.40/hr
mms-1b-asrMassively multilingual speech — ASR and TTS across 1,000+ languages.
- Context
- 30s window
- Speed
- 280 tok/s
- Input
- $0.024 / audio min
- Output
- $0.058 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
mms-ttsMassively multilingual speech — ASR and TTS across 1,000+ languages.
- Context
- 30s window
- Speed
- 304 tok/s
- Input
- $0.022 / audio min
- Output
- $0.053 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
kokoro-82mVery small TTS with natural prosody. Cheap enough for realtime at scale.
- Context
- streaming
- Speed
- 316 tok/s
- Input
- $0.02 / audio min
- Output
- $0.048 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
bark-smallGenerative audio — speech, sound effects and non-verbal sounds.
- Context
- 14s window
- Speed
- 302 tok/s
- Input
- $0.022 / audio min
- Output
- $0.053 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
bark-baseGenerative audio — speech, sound effects and non-verbal sounds.
- Context
- 14s window
- Speed
- 283 tok/s
- Input
- $0.024 / audio min
- Output
- $0.058 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
musicgen-smallText-conditioned music generation with optional melody conditioning.
- Context
- 30s window
- Speed
- 306 tok/s
- Input
- $0.021 / audio min
- Output
- $0.05 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
musicgen-mediumText-conditioned music generation with optional melody conditioning.
- Context
- 30s window
- Speed
- 263 tok/s
- Input
- $0.026 / audio min
- Output
- $0.062 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
stable-audio-open-1-0Text-to-audio for sound design and loops.
- Context
- 47s window
- Speed
- 276 tok/s
- Input
- $0.025 / audio min
- Output
- $0.06 / audio min
- Dedicated
- 1 × H100 · $2.14/hr
stable-diffusion-1-5The open image baseline with the deepest fine-tune ecosystem.
- Context
- —
- Speed
- 284 tok/s
- Input
- $0.024 / image
- Output
- $0.058 / image
- Dedicated
- 1 × H100 · $2.14/hr
stable-diffusion-2-1The open image baseline with the deepest fine-tune ecosystem.
- Context
- —
- Speed
- 284 tok/s
- Input
- $0.024 / image
- Output
- $0.058 / image
- Dedicated
- 1 × H100 · $2.14/hr
stable-diffusion-xlThe open image baseline with the deepest fine-tune ecosystem.
- Context
- —
- Speed
- 213 tok/s
- Input
- $0.035 / image
- Output
- $0.084 / image
- Dedicated
- 1 × H100 · $2.14/hr
stable-diffusion-3-mediumThe open image baseline with the deepest fine-tune ecosystem.
- Context
- —
- Speed
- 248 tok/s
- Input
- $0.029 / image
- Output
- $0.07 / image
- Dedicated
- 1 × H100 · $2.14/hr
stable-diffusion-3-5-largeThe open image baseline with the deepest fine-tune ecosystem.
- Context
- —
- Speed
- 149 tok/s
- Input
- $0.054 / image
- Output
- $0.13 / image
- Dedicated
- 1 × H100 · $2.14/hr
stable-diffusion-3-5-large-turboThe open image baseline with the deepest fine-tune ecosystem.
- Context
- —
- Speed
- 149 tok/s
- Input
- $0.054 / image
- Output
- $0.13 / image
- Dedicated
- 1 × H100 · $2.14/hr
flux-1-schnellRectified-flow transformer. Current open quality leader for text rendering.
- Context
- —
- Speed
- 117 tok/s
- Input
- $0.072 / image
- Output
- $0.173 / image
- Dedicated
- 1 × H100 · $2.14/hr
sana-600mLinear-attention diffusion — 4K images on modest hardware.
- Context
- —
- Speed
- 294 tok/s
- Input
- $0.023 / image
- Output
- $0.055 / image
- Dedicated
- 1 × H100 · $2.14/hr
sana-1-6bLinear-attention diffusion — 4K images on modest hardware.
- Context
- —
- Speed
- 260 tok/s
- Input
- $0.027 / image
- Output
- $0.065 / image
- Dedicated
- 1 × H100 · $2.14/hr
kandinsky-2-2Latent diffusion with a strong prior model.
- Context
- —
- Speed
- 193 tok/s
- Input
- $0.04 / image
- Output
- $0.096 / image
- Dedicated
- 1 × H100 · $2.14/hr
kandinsky-3Latent diffusion with a strong prior model.
- Context
- —
- Speed
- 118 tok/s
- Input
- $0.071 / image
- Output
- $0.17 / image
- Dedicated
- 1 × H100 · $2.14/hr
playground-v2-5Aesthetic-tuned SDXL derivative.
- Context
- —
- Speed
- 213 tok/s
- Input
- $0.035 / image
- Output
- $0.084 / image
- Dedicated
- 1 × H100 · $2.14/hr
kolors-1-0Bilingual text-to-image with strong Chinese prompt handling.
- Context
- —
- Speed
- 233 tok/s
- Input
- $0.031 / image
- Output
- $0.074 / image
- Dedicated
- 1 × H100 · $2.14/hr
auraflow-0-3Fully open flow-based text-to-image.
- Context
- —
- Speed
- 162 tok/s
- Input
- $0.049 / image
- Output
- $0.118 / image
- Dedicated
- 1 × H100 · $2.14/hr
stable-video-diffusion-xtImage-to-video with camera motion control.
- Context
- 25 frames
- Speed
- 263 tok/s
- Input
- $0.026 / second
- Output
- $0.062 / second
- Dedicated
- 1 × H100 · $2.14/hr
cogvideox-2bText-to-video diffusion transformer with usable motion coherence.
- Context
- 6s
- Speed
- 248 tok/s
- Input
- $0.029 / second
- Output
- $0.07 / second
- Dedicated
- 1 × H100 · $2.14/hr
cogvideox-5bText-to-video diffusion transformer with usable motion coherence.
- Context
- 6s
- Speed
- 186 tok/s
- Input
- $0.041 / second
- Output
- $0.098 / second
- Dedicated
- 1 × H100 · $2.14/hr
mochi-1-previewOpen video generation with strong prompt adherence.
- Context
- 5s
- Speed
- 131 tok/s
- Input
- $0.063 / second
- Output
- $0.151 / second
- Dedicated
- 1 × H100 · $2.14/hr
ltx-video-2bRealtime-class video generation — faster than playback on a single GPU.
- Context
- 5s
- Speed
- 248 tok/s
- Input
- $0.029 / second
- Output
- $0.07 / second
- Dedicated
- 1 × H100 · $2.14/hr
hunyuanvideo-13bLarge open video model with cinematic motion.
- Context
- 5s
- Speed
- 112 tok/s
- Input
- $0.076 / second
- Output
- $0.182 / second
- Dedicated
- 1 × H100 · $2.14/hr
wan-2-1-1-3bText and image conditioned video with a small variant for consumer GPUs.
- Context
- 5s
- Speed
- 269 tok/s
- Input
- $0.026 / second
- Output
- $0.062 / second
- Dedicated
- 1 × H100 · $2.14/hr
wan-2-1-14bText and image conditioned video with a small variant for consumer GPUs.
- Context
- 5s
- Speed
- 106 tok/s
- Input
- $0.08 / second
- Output
- $0.192 / second
- Dedicated
- 1 × H100 · $2.14/hr
triposr-baseSingle-image to 3D mesh in under a second.
- Context
- single image
- Speed
- 298 tok/s
- Input
- $0.022 / asset
- Output
- $0.053 / asset
- Dedicated
- 1 × H100 · $2.14/hr
hunyuan3d-2-0Image-to-3D with texture synthesis.
- Context
- single image
- Speed
- 276 tok/s
- Input
- $0.025 / asset
- Output
- $0.06 / asset
- Dedicated
- 1 × H100 · $2.14/hr
instantmesh-baseFeed-forward sparse-view reconstruction to mesh.
- Context
- single image
- Speed
- 294 tok/s
- Input
- $0.023 / asset
- Output
- $0.055 / asset
- Dedicated
- 1 × H100 · $2.14/hr
shap-e-baseText-to-3D implicit function generation.
- Context
- text prompt
- Speed
- 306 tok/s
- Input
- $0.021 / asset
- Output
- $0.05 / asset
- Dedicated
- 1 × H100 · $2.14/hr
genie-style-world-model-baseLearned environment model for policy training and evaluation.
- Context
- rollout
- Speed
- 280 tok/s
- Input
- $0.024 / rollout-hr
- Output
- $0.058 / rollout-hr
- Dedicated
- 1 × H100 · $2.14/hr
llama-guard-3-1bInput and output safety classification against a configurable taxonomy.
- Context
- 8k
- Speed
- 280 tok/s
- Input
- $0.024 / 1M tokens
- Output
- $0.058 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
llama-guard-4-12bInput and output safety classification against a configurable taxonomy.
- Context
- 8k
- Speed
- 117 tok/s
- Input
- $0.072 / 1M tokens
- Output
- $0.173 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
shieldgemma-2bGemma-based content safety classifiers.
- Context
- 8k
- Speed
- 248 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.07 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
shieldgemma-27bGemma-based content safety classifiers.
- Context
- 8k
- Speed
- 65 tok/s
- Input
- $0.136 / 1M tokens
- Output
- $0.326 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
granite-guardian-3-1-2bEnterprise risk detection including groundedness and jailbreak checks.
- Context
- 8k
- Speed
- 248 tok/s
- Input
- $0.029 / 1M tokens
- Output
- $0.07 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
granite-guardian-3-1-8bEnterprise risk detection including groundedness and jailbreak checks.
- Context
- 8k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
esm-2-8mProtein language model — structure and function prediction from sequence.
- Context
- 1k residues
- Speed
- 319 tok/s
- Input
- $0.02 / 1M tokens
- Output
- $0.048 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
esm-2-35mProtein language model — structure and function prediction from sequence.
- Context
- 1k residues
- Speed
- 318 tok/s
- Input
- $0.02 / 1M tokens
- Output
- $0.048 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
esm-2-150mProtein language model — structure and function prediction from sequence.
- Context
- 1k residues
- Speed
- 313 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.05 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
esm-2-650mProtein language model — structure and function prediction from sequence.
- Context
- 1k residues
- Speed
- 292 tok/s
- Input
- $0.023 / 1M tokens
- Output
- $0.055 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
esm-2-3bProtein language model — structure and function prediction from sequence.
- Context
- 1k residues
- Speed
- 224 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
chemberta-77m-mtrMolecular property prediction from SMILES.
- Context
- 512
- Speed
- 316 tok/s
- Input
- $0.02 / 1M tokens
- Output
- $0.048 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
chronos-tinyPretrained forecasting on tokenised time series. Zero-shot on new series.
- Context
- 512 ctx
- Speed
- 319 tok/s
- Input
- $0.02 / 1M rows
- Output
- $0.048 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
chronos-miniPretrained forecasting on tokenised time series. Zero-shot on new series.
- Context
- 512 ctx
- Speed
- 319 tok/s
- Input
- $0.02 / 1M rows
- Output
- $0.048 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
chronos-smallPretrained forecasting on tokenised time series. Zero-shot on new series.
- Context
- 512 ctx
- Speed
- 317 tok/s
- Input
- $0.02 / 1M rows
- Output
- $0.048 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
chronos-basePretrained forecasting on tokenised time series. Zero-shot on new series.
- Context
- 512 ctx
- Speed
- 311 tok/s
- Input
- $0.021 / 1M rows
- Output
- $0.05 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
chronos-largePretrained forecasting on tokenised time series. Zero-shot on new series.
- Context
- 512 ctx
- Speed
- 290 tok/s
- Input
- $0.023 / 1M rows
- Output
- $0.055 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
chronos-bolt-tinyFaster Chronos variant with direct multi-step decoding.
- Context
- 2k ctx
- Speed
- 319 tok/s
- Input
- $0.02 / 1M rows
- Output
- $0.048 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
chronos-bolt-miniFaster Chronos variant with direct multi-step decoding.
- Context
- 2k ctx
- Speed
- 319 tok/s
- Input
- $0.02 / 1M rows
- Output
- $0.048 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
chronos-bolt-smallFaster Chronos variant with direct multi-step decoding.
- Context
- 2k ctx
- Speed
- 317 tok/s
- Input
- $0.02 / 1M rows
- Output
- $0.048 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
chronos-bolt-baseFaster Chronos variant with direct multi-step decoding.
- Context
- 2k ctx
- Speed
- 310 tok/s
- Input
- $0.021 / 1M rows
- Output
- $0.05 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
moirai-smallUniversal forecasting across frequencies and domains.
- Context
- 5k ctx
- Speed
- 319 tok/s
- Input
- $0.02 / 1M rows
- Output
- $0.048 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
moirai-baseUniversal forecasting across frequencies and domains.
- Context
- 5k ctx
- Speed
- 315 tok/s
- Input
- $0.02 / 1M rows
- Output
- $0.048 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
moirai-largeUniversal forecasting across frequencies and domains.
- Context
- 5k ctx
- Speed
- 306 tok/s
- Input
- $0.021 / 1M rows
- Output
- $0.05 / 1M rows
- Dedicated
- 1 × H100 · $2.14/hr
qwq-32b-fp8FP8 serving variant. Reasoning-first Qwen that thinks before answering.
- Context
- 128k
- Speed
- 83 tok/s
- Input
- $0.098 / 1M tokens
- Output
- $0.235 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwq-32b-awqINT4 AWQ serving variant. Reasoning-first Qwen that thinks before answering.
- Context
- 128k
- Speed
- 116 tok/s
- Input
- $0.06 / 1M tokens
- Output
- $0.144 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
qwq-32b-gptqINT8 GPTQ serving variant. Reasoning-first Qwen that thinks before answering.
- Context
- 128k
- Speed
- 97 tok/s
- Input
- $0.079 / 1M tokens
- Output
- $0.19 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
hermes-3-8bCommunity-tuned Llama with steerable system prompts.
- Context
- 128k
- Speed
- 149 tok/s
- Input
- $0.054 / 1M tokens
- Output
- $0.13 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
hermes-3-8b-fp8FP8 serving variant. Community-tuned Llama with steerable system prompts.
- Context
- 128k
- Speed
- 187 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.08 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
hermes-3-8b-awqINT4 AWQ serving variant. Community-tuned Llama with steerable system prompts.
- Context
- 128k
- Speed
- 223 tok/s
- Input
- $0.021 / 1M tokens
- Output
- $0.049 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
hermes-3-8b-gptqINT8 GPTQ serving variant. Community-tuned Llama with steerable system prompts.
- Context
- 128k
- Speed
- 203 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
hermes-3-70b-fp8FP8 serving variant. Community-tuned Llama with steerable system prompts.
- Context
- 128k
- Speed
- 44 tok/s
- Input
- $0.199 / 1M tokens
- Output
- $0.478 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
hermes-3-70b-awqINT4 AWQ serving variant. Community-tuned Llama with steerable system prompts.
- Context
- 128k
- Speed
- 66 tok/s
- Input
- $0.122 / 1M tokens
- Output
- $0.293 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
hermes-3-70b-gptqINT8 GPTQ serving variant. Community-tuned Llama with steerable system prompts.
- Context
- 128k
- Speed
- 53 tok/s
- Input
- $0.161 / 1M tokens
- Output
- $0.385 / 1M tokens
- Dedicated
- 2 × H100 · $4.28/hr
zephyr-7b-betaDPO-tuned Mistral. The open alignment reference point.
- Context
- 32k
- Speed
- 160 tok/s
- Input
- $0.05 / 1M tokens
- Output
- $0.12 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
zephyr-7b-beta-fp8FP8 serving variant. DPO-tuned Mistral. The open alignment reference point.
- Context
- 32k
- Speed
- 197 tok/s
- Input
- $0.031 / 1M tokens
- Output
- $0.074 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
zephyr-7b-beta-awqINT4 AWQ serving variant. DPO-tuned Mistral. The open alignment reference point.
- Context
- 32k
- Speed
- 231 tok/s
- Input
- $0.019 / 1M tokens
- Output
- $0.046 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
zephyr-7b-beta-gptqINT8 GPTQ serving variant. DPO-tuned Mistral. The open alignment reference point.
- Context
- 32k
- Speed
- 213 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
solar-10-7bDepth-upscaled 10.7B that outperforms its parameter count.
- Context
- 4k
- Speed
- 126 tok/s
- Input
- $0.066 / 1M tokens
- Output
- $0.158 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
solar-10-7b-fp8FP8 serving variant. Depth-upscaled 10.7B that outperforms its parameter count.
- Context
- 4k
- Speed
- 164 tok/s
- Input
- $0.041 / 1M tokens
- Output
- $0.098 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
solar-10-7b-awqINT4 AWQ serving variant. Depth-upscaled 10.7B that outperforms its parameter count.
- Context
- 4k
- Speed
- 202 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
solar-10-7b-gptqINT8 GPTQ serving variant. Depth-upscaled 10.7B that outperforms its parameter count.
- Context
- 4k
- Speed
- 181 tok/s
- Input
- $0.033 / 1M tokens
- Output
- $0.079 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
stablelm-2-1-6bCompact multilingual models trained on open data.
- Context
- 16k
- Speed
- 260 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.065 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
stablelm-2-12bCompact multilingual models trained on open data.
- Context
- 16k
- Speed
- 117 tok/s
- Input
- $0.072 / 1M tokens
- Output
- $0.173 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
stablelm-2-12b-fp8FP8 serving variant. Compact multilingual models trained on open data.
- Context
- 16k
- Speed
- 155 tok/s
- Input
- $0.045 / 1M tokens
- Output
- $0.107 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
stablelm-2-12b-awqINT4 AWQ serving variant. Compact multilingual models trained on open data.
- Context
- 16k
- Speed
- 193 tok/s
- Input
- $0.027 / 1M tokens
- Output
- $0.066 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
stablelm-2-12b-gptqINT8 GPTQ serving variant. Compact multilingual models trained on open data.
- Context
- 16k
- Speed
- 172 tok/s
- Input
- $0.036 / 1M tokens
- Output
- $0.086 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
tinyllama-1-1b1.1B chat model — the floor of the usable size range.
- Context
- 2k
- Speed
- 276 tok/s
- Input
- $0.025 / 1M tokens
- Output
- $0.06 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
phi-4-mini-3-8bThe smallest Phi-4. Reasoning quality at edge cost.
- Context
- 128k
- Speed
- 207 tok/s
- Input
- $0.036 / 1M tokens
- Output
- $0.086 / 1M tokens
- Dedicated
- 1 × H100 · $2.14/hr
deepseek-vl2-tinyMixture-of-experts vision-language for dense documents.
- Context
- 4k
- Speed
- 280 tok/s
- Input
- $0.024 / 1M tokens
- Output
- $0.058 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
deepseek-vl2-smallMixture-of-experts vision-language for dense documents.
- Context
- 4k
- Speed
- 228 tok/s
- Input
- $0.032 / 1M tokens
- Output
- $0.077 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr
deepseek-vl2-baseMixture-of-experts vision-language for dense documents.
- Context
- 4k
- Speed
- 194 tok/s
- Input
- $0.039 / 1M tokens
- Output
- $0.094 / 1M tokens
- Dedicated
- 8 × H200 · $23.84/hr