Models

Serverless prices are set live by the sourcing engine — always at or below the reference platform's list price, never below our cost.

GLM 5.2
allais/glm-5-2
text

Frontier open-weights MoE model with strong agentic tool-calling and long-context reasoning.

MoEreasoningtool-callingnew
Params
744B (40B active)
Context
202K
$/1M in·out
$0.72 · $2
reference: $0.9 · $2.5 /1M — you save 20%
Kimi K2 Instruct
allais/kimi-k2-instruct
text

Trillion-parameter MoE tuned for agentic workloads and coding.

MoEagentictool-calling
Params
1T (32B active)
Context
131K
$/1M in·out
$0.48 · $2
reference: $0.6 · $2.5 /1M — you save 20%
DeepSeek V3.2
allais/deepseek-v3-2
text

Efficient MoE with hybrid reasoning modes and strong math/code performance.

MoEreasoning
Params
685B (37B active)
Context
164K
$/1M in·out
$0.448 · $1.344
reference: $0.56 · $1.68 /1M — you save 20%
Qwen3 235B-A22B
allais/qwen3-235b-a22b
text

Hybrid thinking-mode MoE, excellent multilingual and reasoning trade-off.

MoEmultilingualreasoning
Params
235B (22B active)
Context
131K
$/1M in·out
$0.176 · $0.704
reference: $0.22 · $0.88 /1M — you save 20%
Llama 4 Maverick
allais/llama-4-maverick
vision

Natively multimodal MoE with 1M-token context window.

MoEvisionlong-context
Params
400B (17B active)
Context
1000K
$/1M in·out
$0.176 · $0.704
reference: $0.22 · $0.88 /1M — you save 20%
Llama 3.3 70B Instruct
allais/llama-3-3-70b
text

Workhorse dense instruct model; great quality-per-dollar for fine-tuning.

densegeneral
Params
70B dense
Context
131K
$/1M in·out
$0.72 · $0.72
reference: $0.9 · $0.9 /1M — you save 20%
Qwen3 Coder 32B
allais/qwen3-coder-32b
text

Code-specialized model with strong repo-level editing and FIM support.

codingdense
Params
32B dense
Context
131K
$/1M in·out
$0.5831 · $1.28
reference: $0.4 · $1.6 /1M — you save -46%
Llama 3.1 8B Instruct
allais/llama-3-1-8b
text

Small, fast, cheap — ideal for classification and high-volume tasks.

densefastcheap
Params
8B dense
Context
131K
$/1M in·out
$0.3077 · $0.3077
reference: $0.2 · $0.2 /1M — you save -54%
FLUX.1 [dev]
allais/flux-1-dev
image-gen

State-of-the-art open image generation, $0.0014/step reference pricing.

imagediffusion
Params
12B
Context
$/1M in·out
usage-based
Whisper V3 Large
allais/whisper-v3-large
audio

Fast speech transcription, $0.0015/audio-minute reference pricing.

speech-to-text
Params
1.5B
Context
$/1M in·out
usage-based
All Embed Large v3
allais/embed-large-v3
embedding

High-recall embedding model for RAG and semantic search.

embeddingsretrieval
Params
7B
Context
33K
$/1M in·out
$0.1753 · $0
reference: $0.08 · $0 /1M — you save -119%