Open-model inference with a guaranteed rate.
Serverless, dedicated and batch inference for Kimi K3, DeepSeek, GLM and Qwen. Every node is tested before it serves, and every token is metered.
Frontier open models, on tested hardware.
List prices from each model maker as of September 2026, per million tokens, input / output.
Serverless, dedicated or batch.
Serverless
At listCall any model and pay per token. No commitment.
Dedicated
Up to ~40% below list*A guaranteed rate for one model: sustained tokens per second and a per-stream floor, measured every minute.
Batch
60% below listResults within 12 hours, run on our fleet's spare hours.
from openai import OpenAI
client = OpenAI(
base_url="https://api.standardcompute.sh/v1",
api_key=SC_API_KEY,
)
resp = client.chat.completions.create(
model="moonshotai/kimi-k3",
messages=[{"role": "user", "content": prompt}],
stream=True,
)
A rate you can plan around.
You reserve a sustained rate for one model. Capacity stays above it every minute, whatever your traffic does, and you see what was delivered against what you reserved. No rate limits and no shared queue.
- Available to you
- Your traffic
- Your rate
* Example: Kimi K3 at 400 tok/s sustained and 30+ tok/s per stream, used fully on agentic traffic (≈130k-token context, 93% cache hits, 1k-token responses), against list prices of $3 in / $15 out per M tokens, $0.30 cached. Indicative early-access pricing.
Every node is tested before it serves.
Our acceptance tests run on every machine before it joins the fleet, and on a schedule after. It's how we can put a number on a dedicated rate.
- Identity and healthpass
- Memory and NVLink bandwidthpass
- Interconnect (all-reduce across nodes)pass
- Storage and network under loadpass
- Sustained load, no clock sagpass
- Serving throughput and output paritypass
Tell us what you're running.
Early access is open. Tell us which models you use and the throughput you need.