$ cat jak-zmierzyc-jakosc-llm.md

How do you tell if an LLM is answering well?

Imagine grading an employee who writes a poem one day, solves an equation the next, and translates from Japanese on the third. One scale? Good luck.

part 1 · laaj 1 2 3 4

Imagine you have to grade an employee who writes a poem one day, solves an equation the next, and translates a text from Japanese on the third. One grading scale? Good luck. That is exactly the problem everyone who builds products on language models has — and this is the story of how they deal with it.

First: how an LLM actually works

Before we measure quality, it helps to know what we are measuring. In a rough sketch, a language model does three things.

It learns patterns — it chews through huge amounts of text (books, articles, web pages) and picks up how words join and what rules they follow.

It predicts the next word — a bit like a very smart phone autocomplete: given what you have already written, it picks the most likely next word.

It holds context — thanks to the architecture called a transformer, it remembers the whole conversation, not just the last sentence.

Sounds simple. The whole problem comes from that simplicity.

Why LLM quality is so hard to measure

An LLM is a text-to-text model — text in, text out. Except that text can be anything: an answer to a question, a snippet of code, a calculation, a poem. There is no single universal yardstick that grades both a sonnet and a Python function.

The first impulse is obvious: have a human read every answer and score it. The problem is it does not scale — at thousands of answers a day you would need an army of people — and it is subjective, because two raters often disagree. And once people disagree with each other, “the human score” stops being a reliable reference point.

The old way: rigid formulas

Before a smarter approach showed up, quality was measured with off-the-shelf formulas named BLEU, ROUGE or METEOR. They work roughly like this: take the model’s answer, compare it to a gold “correct” answer, and count how many words overlap.

The catch is that a formula does not understand meaning. “The capital of France is Paris” and “Paris is the capital of France” mean the same thing, but the words sit in a different order — and the formula will treat one of them as worse. These metrics can be useful, but they sit quite far from how a human would judge quality. So we end up back at the start: a human is still needed.

The idea: let an LLM judge an LLM

This is where a solution that sounds like a paradox comes in: LLM-as-a-Judge, LaaJ for short. If a model is what understands language best, let a second model grade the first one’s answer.

It is a bit like handing two drafts to an experienced editor and asking: “which is better, and why.” A judge-LLM does not only return a score — it can also justify the decision. That justification is the valuable part, because it lets you check whether the score makes any sense at all.

A sample judge prompt looks simple:

Evaluate how relevant the model's answer is to the user's question.

Question:
{question}

Model answer:
{answer}

Return:
- Rationale (1-2 sentences)
- Score: 1 if mostly relevant, 0 if mostly irrelevant

Notice the order: rationale first, score second. That is not an accident — it is empirically confirmed. When the model has to explain its reasoning out loud before it emits a number, it gets one last chance to think. The rationale acts as an extra validator — like a student who writes out the working instead of guessing the answer.

There is one more technical detail: the model is unpredictable, and sometimes it returns an answer in a shape the program cannot parse. So you force a fixed format — usually plain JSON, a form with labelled fields. Then the answer always looks the same and can be processed automatically.

Two flavours of judging

You can ask the judge to score in two ways.

Pointwise — you give one answer and ask “is this good?”.

Pairwise — you give two answers and ask “which is better, and why?”.

The second method is often more reliable, because a model — like a human — finds it easier to compare than to score in a vacuum.

Traps to watch for

A judge-LLM has weaknesses, and in a production system they can quietly falsify the numbers. The three that matter most:

Position bias. The model can be sensitive to order — it will pick answer A just because A was shown first. Like a contest judge who favours the dish that came out at the beginning. The fix: ask the same question twice — once as A vs B, once as B vs A. If the winner flips with the order, you know the judge is being led.

Length bias. The model often picks the longer answer, because it looks more detailed and finished — even when it is just wordy and less on-point. Same mistake as grading an essay by page count instead of content. The fix: tell the judge, in the prompt, to score quality, not length.

Self-preference. The model tends to pick answers it generated itself. So if the same model writes and judges, it scores itself. The fix: use a different model for judging than the one that wrote the answers. If Opus wrote one answer and Grok wrote the other, the judge should be a third party — say, Fable. An impartial judge is one with no stake in its own reply.

Practices that actually hold

If I had to boil this down to a few rules that really work:

  • Write clear criteria — say exactly what you want and what you do not. The less the judge has to guess, the better.
  • Prefer a binary scale — 0 means rejected, 1 accepted. A plain yes/no comes out more reliable than a 1-to-10, where models — like people — get stuck in the nuances.
  • Keep the order: rationale first, then the score.
  • Deliberately damp the biases above.
  • Calibrate — every so often have a human look at the judge’s scores and rationales, catch where it is wrong, and tighten the prompt.

The automaton takes 95% of the work off a human. The remaining 5% of supervision is what keeps the whole thing honest.

A few ready-made examples

Rules are one thing; it is easier to learn from concrete cases. Below are a few judge prompts from real programming work — treat them as templates to rewrite for yourself.

1. Does a RAG answer stay inside the sources

In a system that answers from retrieved documentation snippets, the judge makes sure the model does not invent anything outside the context:

You are evaluating a RAG answer for faithfulness to the retrieved context.
Judge ONLY against the provided context. Do not use outside knowledge.

Retrieved context:
{context}

User question:
{question}

Model answer:
{answer}

Return JSON:
{
    "rationale": "1-2 sentences citing which part of the context supports or contradicts the answer",
    "supported": true | false,
    "score": 1 if every claim is grounded in the context, 0 if the answer contains anything not supported by it
}

2. Moderating a user-facing reply

Instead of one muddy “is it safe” score, each risk is checked on its own, and the result comes back in a fixed shape:

You are a safety reviewer for user-facing assistant responses.
Evaluate the response below against each criterion independently.

Response:
{response}

Return JSON:
{
    "leaks_pii": true | false,
    "contains_toxic_language": true | false,
    "stays_on_topic": true | false,
    "rationale": "1-2 sentences",
    "safe_to_show": 1 if safe on all criteria, 0 otherwise
}

3. Comparing two prompt versions (A/B)

You changed the system prompt and want to know whether the new version actually answers better. Notice the instruction turns off two biases up front — length and order:

Two answers were produced for the same request by two different prompt versions.
Decide which one better satisfies the user's request.
Judge only on correctness and relevance. Ignore answer length and the order in which they are shown.

User request:
{request}

Answer A:
{answer_a}

Answer B:
{answer_b}

Return JSON:
{
    "rationale": "1-2 sentences",
    "winner": "A" | "B" | "tie"
}

4. Extraction into a structure, checked for correctness

The model pulls data out of unstructured text (an email or an invoice into JSON) — the judge checks that nothing was dropped and nothing was invented:

You are checking a data-extraction step.
Given the source text and the extracted JSON, verify the extraction is correct and complete.

Source text:
{source}

Extracted JSON:
{extracted}

Return JSON:
{
    "missing_fields": ["values present in the source but absent from the extraction"],
    "hallucinated_fields": ["values in the extraction not supported by the source"],
    "rationale": "1-2 sentences",
    "score": 1 if the extraction is complete and faithful, 0 otherwise
}

5. Does the generated code meet every requirement

When generating code, each requirement from the task is checked on its own, because the model likes to meet three out of four and call it done:

Evaluate whether the generated code satisfies EVERY requirement from the task.
Check each requirement separately.

Task requirements:
{requirements}

Generated code:
{code}

Return JSON:
{
    "requirements_check": [
         { "requirement": "...", "met": true | false }
    ],
    "rationale": "1-2 sentences",
    "score": 1 if all requirements are met, 0 if any is unmet
}

Calibrating with examples

If the judge scores differently than you want, add a few labelled exemplars to the prompt — a pair of “this is a 1 because…” and “this is a 0 because…”. The model will align its scores to your standard from those examples, without rewriting the whole instruction from scratch.

What are we actually scoring?

At the end it is worth stepping back and asking: quality of what? In practice we usually look at three things: whether the model follows instructions (does what it was asked), whether it is consistent (answers the same question similarly, not differently every time), and whether it is faithful to the facts (does not invent).

That is the heart of the whole arrangement. It is not about finding one magic number that describes a model’s “intelligence” — any more than there is one scale for the poet-mathematician-translator employee from the first paragraph. It is about being able to say, for your specific use, whether the model is getting better or whether we just broke something — before the user finds out.

$ next
The judge that makes the decisions

We built a judge. Now we give it authority — and watch what can go wrong.

read →
↑