← All toolsLive demo

Model Router

Watch Jev send each prompt to the cheapest model that can handle it — fast, balanced or powerful — and count what you'd save vs. always calling the frontier. No key, just press route.

Each prompt gets one Jev decision — the cheapest tier that can handle it — then we tally what you'd save vs. always calling the frontier model.
0/14routed
0.0selapsed
0/sdecisions/sec
—Jev cost
0FAST
0BALANCED
0POWERFUL
0%cheaper than all-frontier
PromptComplexityRouted to
What's the capital of France?…routing…
Translate 'good morning' into Japanese.…routing…
Fix the grammar: 'he go to the store yesterday'.…routing…
Classify the sentiment: 'I absolutely love this product'.…routing…
Summarize this in one line: our Q3 revenue grew 14% on strong enterprise renewals.…routing…
Write a Python function to merge two sorted lists.…routing…
Explain the difference between TCP and UDP.…routing…
Draft a short, friendly reminder email about an overdue invoice.…routing…
Debug this stack trace and explain the root cause: TypeError: cannot read 'id' of undefined at getUser (auth.js:42).…routing…
Design a distributed rate limiter for 1M requests/sec with fairness across tenants.…routing…
Refactor this 200-line React component to remove re-renders and memoize the tree.…routing…
Prove that the square root of 2 is irrational.…routing…
Write a full REST API in Go with JWT auth, tests, and OpenAPI docs.…routing…
What is 47 * 8?…routing…

A router is one Jev call in front of your models. Get an API key →

Model prices are illustrative ($0.15 / $0.60 / $3.00 per 1M tokens, 1200-token calls) to show the shape of the savings — plug in your own to size the real gain.

Get an API key →Ready-made APIs
Build this with Jev

The same demo is one Jev call

  1. Take each incoming prompt
  2. Ask Jev fast / balanced / powerful
  3. Route in parallel in front of your models
  4. Send each prompt to the tier Jev picked
one call per prompt · parallel · ~70% cheaper vs all-frontier
const results = await Promise.all(items.map((item) =>  // items = prompts
  fetch("https://jevtypesafeai.com/api/v1/decide", {
    method: "POST",
    headers: { Authorization: `Bearer ${process.env.JEV_API_KEY}` },
    body: JSON.stringify({
      state: item,
      questions: {
        tier: { type: "choice", instructions: "Cheapest model that can handle it?",
               criteria: { fast: "small", balanced: "mid", powerful: "frontier" } },
      },
    }),
  }).then((r) => r.json())
));
route(prompt, results[i].answers.tier.choice);   // → "fast"
See a real decision FAST · 198ms
State
Prompt: “What's the capital of France?”
Question choice · tier
fast95%
balanced4%
powerful1%
latency 198mscost $0.000012

How it works

Sending every prompt to a frontier model is the easy default and the expensive one. Most requests — lookups, formatting, classification, short code — are handled just as well by a small, cheap model. A router decides. Here each prompt is one Jev call that returns a typed tier — fast, balanced or powerful — plus a complexity score, all streamed in parallel, and we tally the routed cost against always calling the frontier model.

It's the same one-call-in-front-of-your-models pattern a production gateway uses, made visible: watch the split land and the savings counter climb.

Build it into your gateway

With a hosted key the same Jev call sits in front of your model pool: a typed routing decision per request at $0.25–$0.42/M input tokens, in milliseconds, returning a clean enum you can switch on directly — no parsing, no malformed output. Route the easy 80% to a cheap model, reserve the frontier for the hard 20%, and cut your inference bill without hurting quality on the requests that matter.

FAQ

Does Jev really route each prompt?

Yes — every prompt is an independent Jev call returning a typed tier and a complexity score. The batch runs in parallel and the routing decisions stream in live.

Are the cost savings real?

The shape is real; the exact numbers are illustrative. We use sample model prices ($0.15 / $0.60 / $3.00 per 1M tokens) and a fixed call size to show the saving — plug in your own prices to size the real gain.

Is it free? Can I route my own prompts?

Free, no key — a rate-limited call to the live Jev API. To route your real traffic, wire the same call into your gateway with a hosted key.

More: Context Compaction demo · PR Judge · All Jev tools · Jev API docs

Model Router — route prompts to the cheapest model, live · Jev by TypeSafe AI