New Open source, Apache 2.0 · every number measured →

Decide faster.Spend smarter.

The open-source toll gate in front of your LLM. It blocks prompt injections, sends questions a cheap model can handle to the cheap model, and proves the savings on your own dashboard.

94% of cheap-routed answers good enough0.4% of normal requests blocked$0 to self-host
leanroute.online/playground
Playground
Input
Questions
Output

Sits in front of any OpenAI-compatible model

OpenAIAnthropicGeminiOllamaOpenRouterGroqLangChainn8n
Example decisions

Every quick call, made in milliseconds

Scam or safe, which team, hot lead or cold, cheap model or strong. If the answer is a choice, Laya can make it.

Features

An expensive expert, with a fast receptionist

Your LLM writes and reasons. Laya only picks from the options you give it, so it's fast, free to run, and can't invent an answer.

LLM Cost Cutter

One OpenAI-compatible endpoint. Set model="auto" and Leanroute's router predicts whether a cheap model's answer will be good enough; repeat questions can be answered from cache for $0.

Half of traffic sent to the cheap model: how many of its answers were good enough?

no router
86.5%
Laya score
89.4%
Leanroute
94.0%

10,000 held-out real prompts, answers graded by GPT-4. How we measured

Fast enough for every request

Small classifiers, no text generation.

~150ms

Guard and router together, measured on a laptop CPU. Faster on a GPU.

Guardrails that don't block your users

Prompt injections are stopped before a single token is billed. A first line of defence, measured honestly.

normal requests blocked0.4%
coding requests blocked0%
subtle injections caught36%

Proof, not promises

Turn on the quality check and a sample of cheap answers is re-asked to the strong model, which judges them. Your dashboard shows the pass rate, and the checks' cost comes out of your savings.

cheap answers checked5% sample
judged as goodpass rate
cost of checkingsubtracted

100+ languages

A built-in router picks the English or multilingual checkpoint per request.

Englishहिन्दीEspañol日本語DeutschالعربيةPortuguêsతెలుగుFrançais한국어

Your server, your data

Apache 2.0 all the way down. Run it on a laptop, a free cloud VM or your own GPU. Nothing leaves your infrastructure.

$ docker compose up -d ✓ guard + router loaded ✓ laya preloaded english · multilingual ✓ api + dashboard http://localhost:8000

Templates or your own questions

Seven ready-made templates, or write any yes/no, choice or level question in plain English.

Scam checkTicket routingEmail triageLead scoreGuardrailsModerationModel router+ Custom
Savings

See what you'd stop paying for

Move the sliders to match your traffic. Set prices to your provider's current rates.

Today, all strong model
With Leanroute
Saved per month

An estimate, not a guarantee. Apps that mostly sort, flag and route save a lot; apps that mostly write long answers save little. Server cost for self-hosted Laya (often $0 on a free-tier VM) isn't included.

Developers

Seamless integration.

Change one URL in the OpenAI client you already use, or call the decision API directly.

  • Drop-in compatibleWorks with any OpenAI-style SDK, agent framework or automation tool.
  • Zero vendor lock-inOpen weights, open code. Swap providers by changing one line of config.
  • Tunable thresholdsDecide exactly how sure Laya must be before it acts on its own.
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",  # your Leanroute server (was api.openai.com)
    api_key="lr_your_key",
)
r = client.chat.completions.create(
    model="auto",  # cheap vs strong, decided per request
    messages=[{"role": "user", "content": "Capital of Australia?"}],
)
# r.leanroute → {"route": "cheap", "saved_usd": 0.0021}
Pricing

Free and open source. Hosted plans coming soon.

Open source

$0 forever
  • Full source, Apache 2.0
  • Unlimited decisions
  • Docker, one command
  • Python SDK + OpenAI-compatible gateway
Get it on GitHub

Pro Coming soon

$19 /month
  • Hosted for you, no server to run
  • Savings dashboard
  • A model fine-tuned on your data
Watch on GitHub for launch

Team Coming soon

Custom
  • Guardrails at scale
  • Multiple custom models
  • Priority support
Get in touch

Stop paying an LLM
to say yes or no.

Try it in the Playground, or run it on your own machine in one command.