Decide faster.Spend smarter.
The open-source toll gate in front of your LLM. It blocks prompt injections, sends questions a cheap model can handle to the cheap model, and proves the savings on your own dashboard.
Sits in front of any OpenAI-compatible model
Every quick call, made in milliseconds
Scam or safe, which team, hot lead or cold, cheap model or strong. If the answer is a choice, Laya can make it.
An expensive expert, with a fast receptionist
Your LLM writes and reasons. Laya only picks from the options you give it, so it's fast, free to run, and can't invent an answer.
LLM Cost Cutter
One OpenAI-compatible endpoint. Set model="auto" and Leanroute's router predicts whether a cheap model's answer will be good enough; repeat questions can be answered from cache for $0.
Half of traffic sent to the cheap model: how many of its answers were good enough?
10,000 held-out real prompts, answers graded by GPT-4. How we measured
Fast enough for every request
Small classifiers, no text generation.
Guard and router together, measured on a laptop CPU. Faster on a GPU.
Guardrails that don't block your users
Prompt injections are stopped before a single token is billed. A first line of defence, measured honestly.
Proof, not promises
Turn on the quality check and a sample of cheap answers is re-asked to the strong model, which judges them. Your dashboard shows the pass rate, and the checks' cost comes out of your savings.
100+ languages
A built-in router picks the English or multilingual checkpoint per request.
Your server, your data
Apache 2.0 all the way down. Run it on a laptop, a free cloud VM or your own GPU. Nothing leaves your infrastructure.
Templates or your own questions
Seven ready-made templates, or write any yes/no, choice or level question in plain English.
See what you'd stop paying for
Move the sliders to match your traffic. Set prices to your provider's current rates.
An estimate, not a guarantee. Apps that mostly sort, flag and route save a lot; apps that mostly write long answers save little. Server cost for self-hosted Laya (often $0 on a free-tier VM) isn't included.
Seamless integration.
Change one URL in the OpenAI client you already use, or call the decision API directly.
- Drop-in compatibleWorks with any OpenAI-style SDK, agent framework or automation tool.
- Zero vendor lock-inOpen weights, open code. Swap providers by changing one line of config.
- Tunable thresholdsDecide exactly how sure Laya must be before it acts on its own.
from openai import OpenAI client = OpenAI( base_url="http://localhost:8000/v1", # your Leanroute server (was api.openai.com) api_key="lr_your_key", ) r = client.chat.completions.create( model="auto", # cheap vs strong, decided per request messages=[{"role": "user", "content": "Capital of Australia?"}], ) # r.leanroute → {"route": "cheap", "saved_usd": 0.0021}
curl http://localhost:8000/v1/decide \ -H "Authorization: Bearer lr_your_key" \ -H "Content-Type: application/json" \ -d '{ "text": "Refund the duplicate charge or we cancel.", "questions": { "churn_risk": {"type": "noul", "instructions": "Might the customer cancel?"} } }' # → {"answers": {"churn_risk": {"noul": 0.91}}}
git clone https://github.com/kdandu001-arch/leanroute cd leanroute/server && cp .env.example .env docker compose up -d # website + API → http://localhost:8000 # interactive docs → http://localhost:8000/docs
Free and open source. Hosted plans coming soon.
Open source
- Full source, Apache 2.0
- Unlimited decisions
- Docker, one command
- Python SDK + OpenAI-compatible gateway
Pro Coming soon
- Hosted for you, no server to run
- Savings dashboard
- A model fine-tuned on your data
Stop paying an LLM
to say yes or no.
Try it in the Playground, or run it on your own machine in one command.