Serverless

Pay-per-token inference with zero infrastructure. Prices are set live by the sourcing engine — at or below the reference platform's list price, never below our cost.

Open playground
ModelModalityContext$/1M input$/1M outputReference
GLM 5.2
allais/glm-5-2
text202K$0.72$2$0.9 · $2.5Try
Kimi K2 Instruct
allais/kimi-k2-instruct
text131K$0.48$2$0.6 · $2.5Try
DeepSeek V3.2
allais/deepseek-v3-2
text164K$0.448$1.344$0.56 · $1.68Try
Qwen3 235B-A22B
allais/qwen3-235b-a22b
text131K$0.176$0.704$0.22 · $0.88Try
Llama 4 Maverick
allais/llama-4-maverick
vision1000K$0.176$0.704$0.22 · $0.88Try
Llama 3.3 70B Instruct
allais/llama-3-3-70b
text131K$0.72$0.72$0.9 · $0.9Try
Qwen3 Coder 32B
allais/qwen3-coder-32b
text131K$0.5831$1.28$0.4 · $1.6Try
Llama 3.1 8B Instruct
allais/llama-3-1-8b
text131K$0.3077$0.3077$0.2 · $0.2Try
FLUX.1 [dev]
allais/flux-1-dev
image-genusage-basedTry
Whisper V3 Large
allais/whisper-v3-large
audiousage-basedTry
All Embed Large v3
allais/embed-large-v3
embedding33K$0.1753$0.08 · $0Try

Prices refresh every 10 minutes from live GPU-market quotes across six providers.