$ cat sedzia-ktory-podejmuje-decyzje.md

The judge that makes the decisions

We built a judge. Now we give it authority — and watch what can go wrong.

part 2 · laaj 1 2 3 4

A follow-up to How do you tell if an LLM is answering well?. We built a judge. Now we give it authority — and watch what can go wrong.

In the previous piece I left you with a judge-LLM that can compare two answers and say which is better, and why. Sounds like the end of the story. It is the beginning, because I quickly found something inconvenient: the score itself is almost worthless.

Picture the naive first version of such a system. Two AI agents, each in its own git worktree, independently implement the same task. The judge gets both diffs and scores them: version A — 7/10, version B — 6/10. A wins.

A careful reader of the previous piece is already protesting: we established that a 0/1 scale beats 1–10! Fair — and that is this judge’s first sin. There is a second, deeper one, and that is what this article is about.

So A won. And?

A seven is not code you want in production. It is code that is less bad than the other guy. The judge compared two limping runners and solemnly announced which one limps less. The contest happened, the medal was handed out, and the actual problem — the code is not ready — is still standing there.

A score that does not trigger an action is an expensive ritual. Value shows up only when the verdict starts directing traffic: this code goes to merge, this one goes back for a fix, this one lands on a human’s desk. The judge stops being a juror and becomes a router.

First “is it ready”, then “which is better”

The structural mistake in the scene with the seven is that we asked a comparison question before we asked a binary one. The fixed flow has two stages.

Stage one: readiness check. Each branch goes through a readiness check on its own — not “how good is this code”, but “does it meet every requirement from the task”. Binary, requirement by requirement:

Evaluate whether the implementation satisfies EVERY requirement.
Check each requirement separately.

Task requirements:
{requirements}

Diff:
{diff}

CI results:
{ci_summary}

Return JSON:
{
  "requirements_check": [
    { "requirement": "...", "met": true | false }
  ],
  "issues": [
    "one concrete, fixable instruction per unmet requirement or defect"
  ],
  "rationale": "1-2 sentences",
  "ready": true only if every requirement is met
}

Notice the issues field. That is not decoration — that is the whole point of this stage, more on that in a moment.

There is a trap I learned the hard way: a judge asked to check requirements tends to invent extra ones. Asked to verify an endpoint, it will decide that “no rate limiting” is an unmet requirement — one that was never in the task. The fix: generate the requirement list once, before the agents start, from the task description, and hand that same frozen list to everyone — to the agents as a spec, to the judge as a checklist. The judge ticks boxes. It does not invent.

Which puts all the weight on the quality of the task description — because a checklist generated from a junk description will tick the wrong things, only faster and with more confidence. In a well-written task that list does not need to be generated at all: it is already there, and it is called acceptance criteria — points you can check with a yes/no, written by the person who filed the task. How to write tasks from which that checklist falls out on its own is a separate template for writing tasks. I wrote it for people who assign me work, before agents started taking that work — and that may be the most interesting finding in this whole game: a task written so a human does not have to guess is exactly the same task an agent will not guess at.

Stage two: pairwise comparison — but only among the ready ones. A branch enters the contest only after it passes the readiness check. If both are ready, the judge compares (two passes, A/B and B/A — the position bias from the previous article has not gone anywhere). If only one is ready — it wins by walkover and no comparison happens. If neither — it goes to a human.

We compare candidates for merge, not candidates for a fix. A seven versus a six is not allowed to happen.

A reroll is a lottery. Feedback is a loop.

Fine — what about the branch that failed the readiness check? First impulse: generate again. Maybe this time it works.

That is a lottery at the full price of a ticket. A new generation from scratch does not know what was wrong with the last one — it starts from the same prompt, with the same chance of the same mistakes, and you pay for the whole run again and hope for a different roll.

That is why the issues field on the readiness check matters so much. A judge that returns a list of concrete gaps stops being a goalkeeper punching the ball away and becomes a reviewer doing “request changes”. And the agent does not generate from scratch — it gets a continuation of its own session with an extra instruction stapled on: “fix exactly these three things, do not touch anything else”. It has full context of what it already did and why. The fix costs a fraction of a full generation and — more importantly — it converges: it aims at named gaps instead of drawing a new hand.

After the fix we do not do a full review from zero. The re-check only walks the checklist. Ticked → ready. A loop, not a roulette wheel.

“The best version” does not exist

Which brings us to the nastiest trap in the whole setup. Once you have a repair loop, it is tempting to run it until the thing is really good. Until it is the best.

The problem is that the judge will always find something to fix. Always. A variable could be named more clearly. That condition could be simpler. One more test would help. A model asked for critique will deliver critique — that is its job. “The best version” as a stopping rule is a recipe for an infinite polish loop, where every turn costs real money and the code asymptotically approaches an ideal it never ships.

The definition of “done” has to be binary and closed: every requirement on the frozen list met + tests green = merge. Stop. That is the old rule from the previous article — 0/1 beats 1–10 — in a new, more expensive edition: there the model got stuck in the nuances, here the whole system gets stuck, and it is your tokens paying for it.

Plus two hard circuit breakers, because rules without enforcement are wishes: at most two repair rounds per branch (then escalate to a human, not a third try) and a per-task budget in dollars, after which the system fails loudly. An agent that is stuck should stop and ask — not press on on assumptions, burning the limit.

A log, or trust on paper

One last piece, the least flashy and maybe the most important. A system that scores, repairs and chooses on its own makes a dozen decisions per task in your name. If you cannot reconstruct why it made each of them — you have no basis to trust it.

The solution is boring and it works: an event log, append-only. Each step is one line: what happened, when, who (which agent, which judge pass), with what verdict, rationale and cost. Readiness check on branch A: not ready, two issues. Repair round one. Re-check: ready. Pairwise, pass 1: A wins. Pass 2 (reversed order): A wins. Merge.

From that log you generate a summary at the end — and here is the subtlety: the summary follows from recorded events, it is not written by the judge “from memory”. A model summarising its own decisions after the fact can tidy them up. A log cannot.

A bonus that pays back: every time you override a verdict — “the judge picked A, I would have picked B, because…” — is a labelled example for few-shot calibration of the judge, from the previous article. A system whose decisions you question produces the material for its own improvement. If it keeps a log.

What next

This whole text is theory built on a small sample — the system I am describing is crawling across my own tasks. Before I add the next floors (a Trello queue, a panel for resolving escalations, browser verification), I am running twenty real tasks through it and, on every verdict, writing down my own call: agree, or override.

That one metric — agreement between the judge-router and the human — will decide whether it is worth building further, or whether we have to calibrate first. In the next part I will show the numbers: how many verdicts held up, how many tasks needed repair rounds, what it all cost per task.

That is the real difference between part one and part two of this series. There we asked whether the model answers well. Here we ask whether you can let a model decide — and that question is not answered by conviction, only by a log and an override counter.

$ prev
How do you tell if an LLM is answering well?

Imagine grading an employee who writes a poem one day, solves an equation the next, and translates from Japanese on the third. One scale? Good luck.

← read
$ next
I caught my own AI pipeline cheating. Then I measured what LLM-as-a-Judge cannot see

Eight runs, two agents, a judge with a double pass and a human at the end. All tests green. Along the way — a worker copying solutions from neighbouring tasks, three contradictory readings of one sentence of spec, and a counter of decisions nobody made.

read →
↑