← Use cases

Using Jev as a judge

LLM-as-a-judge usually means prompting a big model, parsing its reply, and trusting a score it made up. Jev as a judge returns exactly one of your grades, a calibrated probability for each, and a confidence value — in one round trip, with nothing to parse.

Jev is TypeSafe AI's first System One model, and the choice primitive turns it into a judge. You define the verdict space — pass / fail, a 1–5 rubric, which of two answers is better — send the thing being judged as state, and Jev returns one key locked to your options, a probability distribution across them, and a single confidence score. Because the output is a typed structured value rather than generated text, a Jev judge never invents a grade, never trails off, and never hands you prose you have to clean up.

The judge in code

// Grade an answer with Jev, then route on the calibrated confidence
const res = await fetch("https://jevtypesafeai.com/api/v1/decide", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    Authorization: `Bearer ${process.env.JEV_KEY}`, // jv_live_...
  },
  body: JSON.stringify({
    state: { question, answer },
    questions: {
      verdict: { type: "choice", options: ["correct", "partial", "wrong"] },
      quality: { type: "score", levels: ["unusable", "poor", "ok", "good", "excellent"] },
    },
  }),
});
const { verdict, quality } = await res.json();

if (verdict.confidence > 0.85) apply(verdict.key);       // trust the confident calls
else queueForHumanReview({ question, answer, verdict }); // send the uncertain tail to a human

One call grades on two axes at once (a pass/partial/wrong verdict and a 1–5 quality score), each with its own calibrated confidence. verdict.key is locked to your three options, so there is no invalid grade to handle — you branch straight on it.

Why calibration is the whole point of a judge

A judge is only useful if its confidence tracks how often it is right. Jev is trained with RLCD — Reinforcement Learning for Calibrated Decisions — which optimises for epistemically honest probabilities instead of fluent writing. So a 0.9 verdict really does hold up about nine times in ten. That is what lets you route on the score.

Jev verdict confidence? calibrated high → auto-accept low → human review
Because every verdict carries calibrated confidence, you auto-accept the confident ones and send only the uncertain tail to a human.

Jev judge vs prompting an LLM to grade

Prompting an LLMJev as a judge
OutputFree text to parseOne typed grade
Invalid gradesPossibleImpossible (schema-locked)
CalibrationUncalibrated by defaultRLCD-calibrated
LatencySeconds~70–500ms
CostFrontier-LLM pricingFar cheaper (see pricing)

On TypeSafe's own four-workflow benchmark Jev reaches roughly 67.8% agreement with reference answers — on par with GPT-5.6 Terra (67.9%) while running about 25x faster (0.4s vs 10.1s) — and posts a 0% structured-output error rate where general LLMs land anywhere from 0.58% to 45.5%. For a judge that runs on every item in a queue, matching a frontier model's agreement while returning in 0.4s is the difference between grading everything and sampling a handful.

See it running

PR Judge scores a pull request, and Live Post Judge grades a social post in real time — both are Jev as a judge over the hosted endpoint. Point your own evaluations at POST https://jevtypesafeai.com/api/v1/decide with a jv_live_ key (or the official POST https://api.typesafe.ai/v1/systemone) and you get the same typed, calibrated verdict.

FAQ

Is Jev-as-a-judge as accurate as GPT or Claude?

On TypeSafe's four-workflow benchmark Jev sits around 67.8% agreement, level with GPT-5.6 Terra and a few points under GPT-5.6 Sol and Claude Opus 5 — but it returns in about 0.4 seconds versus ~10, and never emits an invalid grade. For high-volume evaluation that trade is usually worth it.

How is this different from LLM-as-a-judge?

LLM-as-a-judge generates text you parse and re-interpret. Jev emits a typed grade directly, locked to your rubric, with a calibrated probability — so there is no parsing, no invalid output, and the confidence is meaningful.

Can I trust the confidence number?

Treat it as calibrated but not a certificate. RLCD trains the probabilities to match real outcomes, so you can threshold on them — but you should still calibrate your cut-off against your own labelled data.

See also: PR Judge · Live Post Judge · How RLCD works · Jev classifier

Run Jev as a judge

Grab a jv_live_ key and return a typed, calibrated verdict in one call.

▶ Try Jev freeGet an API key →
Jev as a judge — typed LLM-as-a-judge with calibrated confidence · Jev by TypeSafe AI