Real-time AI is
breaking your budget.
For every job beyond need-it-now.
Choose your timeline. Save on tokens.
POST /v1/batches
illustrative
{
"model": "glm-5.2",
"completion_window": "6h",
"messages": [
{
"role": "user",
"content": "Summarize these 40,000 support tickets."
}
]
}
A drop-in API for leading open source models. One new field: your deadline.
Time turns into efficiency.
Your timeline gives us room to optimize. Whether it's tagging a dataset, summarizing tickets, or clearing an eval queue, we can help you save money. Surplus hosts open source models to maximize high-quality throughput and passes the savings on. Same models, same performance. Pay for work, not gaps.
Run now~55% utilized
Given time~98% utilized
Efficiency means savings.
The more time you give, the less you pay. Just say when. e.g. Escalation triage that used to run in seconds can run in hours — same routing, a fraction of the cost.
today's market
surplus
Join the beta.
You're on the list.
We'll reach out about beta access soon.
We're onboarding a small group of teams & developers now.