Using Jev as a judge
LLM-as-a-judge usually means prompting a big model, parsing its reply, and trusting a score it made up. Jev as a judge returns exactly one of your grades, a calibrated probability for each, and a confidence value — in one round trip, with nothing to parse.
Jev is TypeSafe AI's first System One model, and the choice primitive turns it into a judge. You define the verdict space — pass / fail, a 1–5 rubric, which of two answers is better — send the thing being judged as state, and Jev returns one key locked to your options, a probability distribution across them, and a single confidence score. Because the output is a typed structured value rather than generated text, a Jev judge never invents a grade, never trails off, and never hands you prose you have to clean up.
The judge in code
// Grade an answer with Jev, then route on the calibrated confidence
const res = await fetch("https://jevtypesafeai.com/api/v1/decide", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.JEV_KEY}`, // jv_live_...
},
body: JSON.stringify({
state: { question, answer },
questions: {
verdict: { type: "choice", options: ["correct", "partial", "wrong"] },
quality: { type: "score", levels: ["unusable", "poor", "ok", "good", "excellent"] },
},
}),
});
const { verdict, quality } = await res.json();
if (verdict.confidence > 0.85) apply(verdict.key); // trust the confident calls
else queueForHumanReview({ question, answer, verdict }); // send the uncertain tail to a humanOne call grades on two axes at once (a pass/partial/wrong verdict and a 1–5 quality score), each with its own calibrated confidence. verdict.key is locked to your three options, so there is no invalid grade to handle — you branch straight on it.
Why calibration is the whole point of a judge
A judge is only useful if its confidence tracks how often it is right. Jev is trained with RLCD — Reinforcement Learning for Calibrated Decisions — which optimises for epistemically honest probabilities instead of fluent writing. So a 0.9 verdict really does hold up about nine times in ten. That is what lets you route on the score.
- One verdict, locked to the grades you defined — no invalid or made-up labels
- Calibrated probabilities across every grade
- Confidence you can threshold on for review routing
- Roughly 70–500ms per verdict — fast enough to grade in the request path
Jev judge vs prompting an LLM to grade
| Prompting an LLM | Jev as a judge | |
|---|---|---|
| Output | Free text to parse | One typed grade |
| Invalid grades | Possible | Impossible (schema-locked) |
| Calibration | Uncalibrated by default | RLCD-calibrated |
| Latency | Seconds | ~70–500ms |
| Cost | Frontier-LLM pricing | Far cheaper (see pricing) |
On TypeSafe's own four-workflow benchmark Jev reaches roughly 67.8% agreement with reference answers — on par with GPT-5.6 Terra (67.9%) while running about 25x faster (0.4s vs 10.1s) — and posts a 0% structured-output error rate where general LLMs land anywhere from 0.58% to 45.5%. For a judge that runs on every item in a queue, matching a frontier model's agreement while returning in 0.4s is the difference between grading everything and sampling a handful.
See it running
PR Judge scores a pull request, and Live Post Judge grades a social post in real time — both are Jev as a judge over the hosted endpoint. Point your own evaluations at POST https://jevtypesafeai.com/api/v1/decide with a jv_live_ key (or the official POST https://api.typesafe.ai/v1/systemone) and you get the same typed, calibrated verdict.
FAQ
Is Jev-as-a-judge as accurate as GPT or Claude?
On TypeSafe's four-workflow benchmark Jev sits around 67.8% agreement, level with GPT-5.6 Terra and a few points under GPT-5.6 Sol and Claude Opus 5 — but it returns in about 0.4 seconds versus ~10, and never emits an invalid grade. For high-volume evaluation that trade is usually worth it.
How is this different from LLM-as-a-judge?
LLM-as-a-judge generates text you parse and re-interpret. Jev emits a typed grade directly, locked to your rubric, with a calibrated probability — so there is no parsing, no invalid output, and the confidence is meaningful.
Can I trust the confidence number?
Treat it as calibrated but not a certificate. RLCD trains the probabilities to match real outcomes, so you can threshold on them — but you should still calibrate your cut-off against your own labelled data.
See also: PR Judge · Live Post Judge · How RLCD works · Jev classifier
Run Jev as a judge
Grab a jv_live_ key and return a typed, calibrated verdict in one call.