Frontier-Level Results...at a fraction of the cost
Point every AI workload at SmartRoute.
RandiRouter selects and serves the best-fit model for each task, balancing quality, reliability, speed, and cost.
One intelligent default.Control when you need it.
Use SmartRoute and let RandiRouter decide, or select an inference level for workloads that need an explicit policy.
RandiRouter chooses the right inference level and the best current model for the task.
randirouter/smartroute
Maximum capability
For demanding work where depth, judgement, and advanced reasoning matter most.
Everyday intelligence
Strong general-purpose capability for everyday reasoning, writing, and synthesis.
Efficient throughput
Fast, capable inference for clear, repeatable, and high-volume work.
Keep the simplicity.Lose the frontier-only bill.
Sending every task to a premium model is simple, but wasteful. SmartRoute keeps frontier capability for work that needs it and uses more efficient inference for everything else.
The model market changes every week. Your integration should not.
3.0M tokens × $16/M
Save $18.3638% less per month
Illustrative only. The comparison assumes every request would otherwise use a premium frontier model. Actual rates and savings will vary by workload and route mix.
One request.The right model.
No model catalogue to manage. No routing rules embedded across your products.
Send every workload
Keep your OpenAI-compatible client and set the model to randirouter/smartroute.
SmartRoute decides
RandiRouter weighs task fit, quality, reliability, speed, and cost, then serves the best-fit current model.
Audit every decision
See the inference level, routing reason, token usage, latency, rate, and total cost.
Change the base URL.Not your architecture.
Use SmartRoute for automatic selection, or choose Frontier, Balanced, or Efficient when a workload needs an explicit level.
Request early access →from openai import OpenAI
client = OpenAI(
api_key="rr_live_••••••••",
base_url="https://api.randirouter.com/v1"
)
response = client.chat.completions.create(
model="randirouter/smartroute",
messages=[{"role": "user", "content": prompt}]
)
Separate the access.See every cost.
Create a key for each application, team, user, or agent. Meter usage independently, attribute spend clearly, set limits, and revoke one key without disrupting everyone else.
- 01Organise accessSeparate keys for every workload or owner
- 02Attribute usageTrack tokens and spend against the right key
- 03Enforce limitsSet per-key limits and revoke access independently
rr_live_••••7A2Frr_live_••••19BCrr_live_••••E410Trust the route.Verify the receipt.
SmartRoute removes the model decision from your teams without hiding the outcome. Inspect the inference level, routing reason, and exact cost.
- 01 Inference level and routing reason
- 02 Tokens, latency, reliability, and fallback
- 03 Exact cost and frontier-default comparison
Scale on your terms.
Start with metered inference. Move to recurring capacity when your workloads become consistent.
Usage
Each inference level has clear input and output rates. SmartRoute bills at the level selected.
by inference level
- One OpenAI-compatible endpoint
- SmartRoute plus three explicit levels
- Multiple keys with per-key metering
- Decision-level cost ledger
Capacity
A recurring weighted allowance shared across Frontier, Balanced, and Efficient inference.
long usage windows
- Everything in Usage
- Weighted inference allowance
- Shared across all three levels
- Per-key budgets and hard limits
The useful details.
Still deciding whether the router fits your stack? Start here.
Do I choose which model handles a request?
Not by default. Send requests to randirouter/smartroute and RandiRouter selects the level and best-fit underlying model. You can choose Frontier, Balanced, or Efficient explicitly when a workload needs tighter control.
Why not just use a frontier model for everything?
Because many production tasks do not need frontier-level reasoning. Routing simple and everyday requests to lower-cost models can materially reduce your blended spend while keeping premium intelligence available for hard work.
What does the decision ledger show?
The decision ledger shows the inference level, routing reason, token usage, latency, rate, and exact request cost. RandiRouter manages the underlying model selection so your teams do not have to manage the model catalogue.
Can I set a predictable monthly budget?
Yes. Usage pricing suits variable workloads. Capacity plans provide recurring weighted inference allowances, plus team-level budgets and hard limits.
Can I separate usage by application or team?
Yes. Create separate API keys for applications, teams, users, or individual agents. Each key can be metered independently, given its own usage limit, and revoked without affecting your other workloads.
Does RandiRouter use customer data to train models?
RandiRouter does not train the underlying language models on customer traffic. Optional, privacy-scrubbed outcome evidence may be used to improve the routing algorithm, and organisations can opt out.
Stop choosing models.Start routing outcomes.
Join the early-access programme and help shape RandiRouter for production AI teams.