Early access

Open-model inference with a guaranteed rate.

Serverless, dedicated and batch inference for Kimi K3, DeepSeek, GLM and Qwen. Every node is tested before it serves, and every token is metered.

Models

Frontier open models, on tested hardware.

List prices from each model maker as of September 2026, per million tokens, input / output.

  • Moonshot AI

    Kimi K3

    1M ctx · $3.00 / $15.00
  • Qwen

    Qwen3.8 2.4T

    1M ctx · $2.00 / $6.00
  • DeepSeek

    DeepSeek V4 Pro

    1M ctx · $1.32 / $3.96
  • Z.ai

    GLM-5.3

    1M ctx · $1.40 / $4.40
  • DeepSeek

    DeepSeek V4.1 Flash

    1M ctx · $0.30 / $1.20
  • Z.ai

    GLM-5.3 Flash

    1M ctx · $0.15 / $0.50
  • MiniMax

    MiniMax M3

    1M ctx · $0.30 / $1.20
  • Moonshot AI

    Kimi K2.7 Code

    262K ctx · $0.67 / $3.35
Ways to run

Serverless, dedicated or batch.

  • Serverless

    At list

    Call any model and pay per token. No commitment.

  • Dedicated

    Up to ~40% below list*

    A guaranteed rate for one model: sustained tokens per second and a per-stream floor, measured every minute.

  • Batch

    60% below list

    Results within 12 hours, run on our fleet's spare hours.

pythonOpenAI-compatible
from openai import OpenAI

client = OpenAI(
    base_url="https://api.standardcompute.sh/v1",
    api_key=SC_API_KEY,
)

resp = client.chat.completions.create(
    model="moonshotai/kimi-k3",
    messages=[{"role": "user", "content": prompt}],
    stream=True,
)
response Sure. The repo defines three FastAPI routes.
Dedicated

A rate you can plan around.

You reserve a sustained rate for one model. Capacity stays above it every minute, whatever your traffic does, and you see what was delivered against what you reserved. No rate limits and no shared queue.

One dedicated rate over a day
  • Available to you
  • Your traffic
  • Your rate

* Example: Kimi K3 at 400 tok/s sustained and 30+ tok/s per stream, used fully on agentic traffic (≈130k-token context, 93% cache hits, 1k-token responses), against list prices of $3 in / $15 out per M tokens, $0.30 cached. Indicative early-access pricing.

Testing

Every node is tested before it serves.

Our acceptance tests run on every machine before it joins the fleet, and on a schedule after. It's how we can put a number on a dedicated rate.

acceptance reportexample
  1. Identity and healthpass
  2. Memory and NVLink bandwidthpass
  3. Interconnect (all-reduce across nodes)pass
  4. Storage and network under loadpass
  5. Sustained load, no clock sagpass
  6. Serving throughput and output paritypass
6 of 6 checks passedready for traffic

Tell us what you're running.

Early access is open. Tell us which models you use and the throughput you need.

Get in touch