Frontier-Level Results...at a fraction of the cost

Point every AI workload at SmartRoute.

RandiRouter selects and serves the best-fit model for each task, balancing quality, reliability, speed, and cost.

One OpenAI-compatible endpoint Automatic by default Control when you need it
Live routing Optimizing
model randirouter/smartroute
Refactor async payment flow 14,862 tokens · complex reasoning
Frontier$0.238
Summarize customer feedback 8,204 tokens · synthesis
Balanced$0.074
Extract invoice fields 5,180 tokens · structured task
Efficient$0.031
This session28.2K tokens
Routed cost$0.34
Saved vs frontier49%
Built for teams shipping with
OpenAI SDKLangChainVercel AI SDKcURLLiteLLM
Automatic by default

One intelligent default.Control when you need it.

Use SmartRoute and let RandiRouter decide, or select an inference level for workloads that need an explicit policy.

AutoSmartRoute

RandiRouter chooses the right inference level and the best current model for the task.

randirouter/smartroute
Frontier01

Maximum capability

For demanding work where depth, judgement, and advanced reasoning matter most.

ArchitectureHard reasoningAgent planning
Balanced02

Everyday intelligence

Strong general-purpose capability for everyday reasoning, writing, and synthesis.

SummariesDraftingClassification
Efficient03

Efficient throughput

Fast, capable inference for clear, repeatable, and high-volume work.

ExtractionFormattingTagging
Stop paying for the default

Keep the simplicity.Lose the frontier-only bill.

Sending every task to a premium model is simple, but wasteful. SmartRoute keeps frontier capability for work that needs it and uses more efficient inference for everything else.

The model market changes every week. Your integration should not.

Monthly token volume3.0 million
Illustrative rates
Example workload mixChoose a profile
Every request on FrontierWithout SmartRoute
$48.00

3.0M tokens × $16/M

SmartRouteSame workload, intelligently routed
$29.64

Save $18.3638% less per month

Frontier 34%$16.32
Balanced 16%$4.32
Efficient 50%$9.00

Illustrative only. The comparison assumes every request would otherwise use a premium frontier model. Actual rates and savings will vary by workload and route mix.

How SmartRoute works

One request.The right model.

No model catalogue to manage. No routing rules embedded across your products.

01

Send every workload

Keep your OpenAI-compatible client and set the model to randirouter/smartroute.

02

SmartRoute decides

RandiRouter weighs task fit, quality, reliability, speed, and cost, then serves the best-fit current model.

03

Audit every decision

See the inference level, routing reason, token usage, latency, rate, and total cost.

One integration

Change the base URL.Not your architecture.

Use SmartRoute for automatic selection, or choose Frontier, Balanced, or Efficient when a workload needs an explicit level.

Request early access
app.py
from openai import OpenAI

client = OpenAI(
    api_key="rr_live_••••••••",
    base_url="https://api.randirouter.com/v1"
)

response = client.chat.completions.create(
    model="randirouter/smartroute",
    messages=[{"role": "user", "content": prompt}]
)
route.level"balanced"824ms
Keys that match your organisation

Separate the access.See every cost.

Create a key for each application, team, user, or agent. Meter usage independently, attribute spend clearly, set limits, and revoke one key without disrupting everyone else.

  • 01
    Organise accessSeparate keys for every workload or owner
  • 02
    Attribute usageTrack tokens and spend against the right key
  • 03
    Enforce limitsSet per-key limits and revoke access independently
API key controlsIllustrative view
Production apprr_live_••••7A2F
72% of limit
1.8M tokens this monthLimit enabled
Engineering agentsrr_live_••••19BC
54% of limit
3.2M tokens this monthLimit enabled
Research teamrr_live_••••E410
31% of limit
620K tokens this monthLimit enabled
3 active keys+ New API key
Decision ledgerIllustrative view
Frontier 32%Balanced 28%Efficient 40%
Total routed4.82M tokensVs frontier default$21.43 lower
Transparent by decision

Trust the route.Verify the receipt.

SmartRoute removes the model decision from your teams without hiding the outcome. Inspect the inference level, routing reason, and exact cost.

  • 01 Inference level and routing reason
  • 02 Tokens, latency, reliability, and fallback
  • 03 Exact cost and frontier-default comparison
Simple ways to pay

Scale on your terms.

Start with metered inference. Move to recurring capacity when your workloads become consistent.

Pay as you goFor shipping & testing

Usage

Each inference level has clear input and output rates. SmartRoute bills at the level selected.

Meteredstable rates
by inference level
Request early access
  • One OpenAI-compatible endpoint
  • SmartRoute plus three explicit levels
  • Multiple keys with per-key metering
  • Decision-level cost ledger
Most predictable
SubscriptionFor steady workloads

Capacity

A recurring weighted allowance shared across Frontier, Balanced, and Efficient inference.

Capacityrecurring short and
long usage windows
Discuss capacity
  • Everything in Usage
  • Weighted inference allowance
  • Shared across all three levels
  • Per-key budgets and hard limits
Questions, routed

The useful details.

Still deciding whether the router fits your stack? Start here.

Do I choose which model handles a request?

Not by default. Send requests to randirouter/smartroute and RandiRouter selects the level and best-fit underlying model. You can choose Frontier, Balanced, or Efficient explicitly when a workload needs tighter control.

Why not just use a frontier model for everything?

Because many production tasks do not need frontier-level reasoning. Routing simple and everyday requests to lower-cost models can materially reduce your blended spend while keeping premium intelligence available for hard work.

What does the decision ledger show?

The decision ledger shows the inference level, routing reason, token usage, latency, rate, and exact request cost. RandiRouter manages the underlying model selection so your teams do not have to manage the model catalogue.

Can I set a predictable monthly budget?

Yes. Usage pricing suits variable workloads. Capacity plans provide recurring weighted inference allowances, plus team-level budgets and hard limits.

Can I separate usage by application or team?

Yes. Create separate API keys for applications, teams, users, or individual agents. Each key can be metered independently, given its own usage limit, and revoked without affecting your other workloads.

Does RandiRouter use customer data to train models?

RandiRouter does not train the underlying language models on customer traffic. Optional, privacy-scrubbed outcome evidence may be used to improve the routing algorithm, and organisations can opt out.

One intelligent default for every AI workload.

Stop choosing models.Start routing outcomes.

Join the early-access programme and help shape RandiRouter for production AI teams.