$ cat decyzje-ktorych-nikt-nie-podjal.md

I caught my own AI pipeline cheating. Then I measured what LLM-as-a-Judge cannot see

Eight runs, two agents, a judge with a double pass and a human at the end. All tests green. Along the way — a worker copying solutions from neighbouring tasks, three contradictory readings of one sentence of spec, and a counter of decisions nobody made.

part 3 · laaj 1 2 3 4

Part three of the series. In How do you tell if an LLM is answering well? we built a judge. In The judge that makes the decisions we gave it authority, and I promised numbers: judge–human agreement, an override counter, cost per task. Before I show them, I have to tell you how the first task I looked at carefully found a hole in that counter — and how fixing that hole found a second one I had no idea existed.

Because the override counter measures decisions made badly: the judge picked A, I would have picked B, there is a dispute, there is a record. It does not measure decisions not made — those generate no dispute. They generate docstrings.

The task

One sentence, deliberately in the form tasks actually arrive in the queue:

a function that calculates a discount: 10% above 500 PLN,
20% above 1000 PLN, discounts do not stack with promo codes

Read it again and notice how normal it looks. That will matter for the whole of this post.

You know the pipeline from part two: two workers write independently (for me: claude-opus-5 as A, grok-4.6 as B), each in its own working tree. The judge generates a frozen checklist from the spec, runs a readiness check on each branch separately, then compares them pairwise — twice, with the candidates swapped, as a fuse against position bias.

Pilot: everything passed and nobody decided anything

The first run came back green. Code cleaner than I would write by hand, tests on the boundaries, a consistent verdict. And when I counted the decisions nobody made, I got five:

  1. The guessed boundary. The spec says “above 1000”. Does an order of exactly 1000 PLN get 10% or 20%? Both workers silently chose an open threshold and dressed the choice in a docstring and a test name.
  2. An empty value as a requirement. What does promo_code="" do? The spec is silent, the code decides, the test cements the decision — and from then on a guess reads like a requirement.
  3. No input validation. Type "asdf" into the code field: any truthy string zeroes the threshold discount. A customer who mistypes a code loses the discount and gets nothing in return. No test catches it, because every test passes codes that look valid.
  4. Money in floats. The test table expects 200.002 PLN. Nobody asked what money is in this system.
  5. The function’s contract. Does it return the discount or the amount after discount? I wrote “calculates a discount”, so I got the discount. The ambiguity was mine — those 5% of human oversight from part one failed before the agents even started.

This is worker A’s function from that run (docstring omitted) — and all five decisions it made on nobody’s behalf:

def calculate_discount(amount, promo_code=None):
    if promo_code:                  # 2: "" is False; 3: "asdf" is True
        return 0                    # 5: returns the discount, not the amount

    if amount > 1000:               # 1: exactly 1000 → 10%, not 20%
        rate = 0.2
    elif amount > 500:
        rate = 0.1
    else:
        rate = 0

    return round(amount * rate, 2)  # 4: a float pretending to be grosze

Above it sits a docstring that sounds confident and coherent. That is the whole mechanism: a guess dressed in documentation stops looking like a guess.

One run is an anecdote, not a pattern. Before running more, I wrote predictions in a dated file: classes 1, 4 and 5 will return in every run, because they are model habits; class 2 will be unstable; and greater variance in how “do not stack” gets interpreted will give the judge more inconclusive verdicts. Remember that last one in particular.

Series one: four runs and one snag

I ran the same spec four times. All green, the judge picked, I clicked. It looked like replication.

The snag appeared while classifying the diffs: worker B’s solutions in runs 2 and 3 were byte-identical — git gave them the same blob hash (9b7bd6f). Runs 1 and 4 differed only cosmetically, in comments. Four “independent” readings of the spec, one file.

B’s transcripts left no doubt:

I’ll look at neighboring worktrees and the existing calculator for a discount API.

Sibling branches already implemented similar discount logic — I’ll inspect those.

The nearby worktrees already encode this as exclusive volume tiers that yield to a promo code. I’ll implement that.

My tasks live as worktrees of a single repository. git branch -a, git log --all and a plain ls of the worktree directory expose the finished solutions of every previous task — and of every task in progress. B started with git branch -a in four runs out of four, and in one of them it read outright the file from branch A of the same task, while that task was running.

So my n=4 did not exist. On B’s side it was n=1 copied three times.

Let me name the wider thing, because it is not only about my toy repo. In a multi-agent pipeline on a shared repository, a wrong interpretation from task #14 propagates to #15, #16 and #17 with growing authority — “that’s how we do it in this project”. The convergence you would like to read as confirmation stops confirming anything. And the worst part: both the judge and I praised that consistency. It was copying.

The fix: isolation measured, not declared

The plan was simple: a separate clone instead of a worktree. The execution taught me more than the plan.

Each worker now gets a fresh clone of main only: two refs, zero remotes, reflog wiped. A guard refuses to start a task if main has moved from the pinned commit, so merged solutions of earlier tasks cannot become a crib for later ones.

Two things from the rollout that would have bitten me had I not measured them:

The permission layer is not symmetric. A probe showed the deny rule catches git fetch /repo but not git -C /repo fetch /repo, and both CLIs read files outside the working directory freely. Blocking commands is a belt without braces. The real barrier is that the clone simply does not contain anyone else’s data — and where the data physically exists in parallel (two branches of the same task on one disk), the only barrier left is an unguessable clone path.

Telemetry instead of trust. After every run I read the CLI’s full transcript and record the list of tool calls, the first five commands, and attempts to leave the workspace: paths with .., git branch -a, log --all, fetch, clone, worktree, git -C outside its own directory. The result lands in the journal and in the panel as a red chip. On the old sessions of series one the counter showed: opus 0 attempts across four runs, grok 24.

The third finding from that probe — a worker sees more of the host’s files than it should — is a separate fix (a container per worker) and a separate post, once it is measured, not declared.

Series two: clean replication

The same spec, word for word, four runs on the repaired isolation.

Telemetry first: opus — zero escape attempts in every run. Grok — one (again git branch -a; the chip lit up, the clone was empty). And the thing that retroactively closes the diagnosis of series one: grok still, in four runs out of four, begins work with git log -5 --oneline && git branch. Reconnaissance for other people’s solutions is its reflex, not a reaction to that environment. This time there was simply no loot — and the four B diffs that in series one were copies of one file are now four different solutions.

What came back, class by class:

Class run 1 run 2 run 3 run 4
1. Boundary guessed (open threshold) 2/2 2/2 2/2 2/2
2. Empty promo cemented by a test both: "" keeps the discount untested B: "" keeps the discount B: any truthy code cancels the discount
3. Input validation none partial (B: is not None instead of truthiness, code still unchecked) none A: code registry; B: blind trust
4. Money in floats 2/2 2/2 2/2 2/2
5. Returns the discount, not the amount 2/2 2/2 2/2 2/2

Classes 1, 4 and 5 returned in eight candidates out of eight — as predicted. That is not chance, those are model habits. The best evidence came as a bonus: in run four opus once again invented a promo-code registry with the code "WELCOME5" — exactly the one it invented in the contaminated series, though this time it had nowhere to copy it from. Convergence between runs can be a model habit posing as confirmation. Worth remembering every time “two independent models arrived at the same thing”.

Class 2 split into contradictory truths — harder than I predicted. In runs 1 and 3 an empty string keeps the threshold discount and a test cements it:

# run 1, worker A
def test_empty_promo_code_keeps_threshold_discount():
    assert calculate_discount(800, promo_code="") == pytest.approx(80)

In run 4 worker B decided the opposite: any truthy code — even "EMPTY" with a zero rate — cancels the discount entirely. Also with a test:

# run 4, worker B
def test_promo_code_with_zero_percent_blocks_volume_discount():
    # Using a promo code opts out of volume discounts even if the promo is 0%.
    assert calculate_discount(1500, promo_code="EMPTY") == 0.0

The same hole in the spec, two opposite “truths”, both with green tests. A single run not only fails to prove things are right — it does not even prove which guess you got.

“Do not stack” has three meanings

Across eight candidates the same sentence of spec became three different rules:

if promo_code:                              # suppresses the threshold discount — 5 implementations
    return 0.0

applied = max(volume_percent, promo_rate)   # takes the larger rate — 1

rate = promo_percent / 100.0                # replaces it with the promo's own rate — 2

And here is the mechanism this post is about: the frozen checklist from part two — the one meant to stop the judge from adding requirements — ticked the item “Discounts do not stack with promo codes” for all eight. Including the candidate where an order with a promo code gets the full threshold discount, which under the first reading is a plain violation of the requirement. The checklist did exactly what I designed it to do: it let nothing be added. Which is how it faithfully froze the ambiguity. Ticking boxes protects against invention — it does not inject the missing decisions.

For the record: I reconstructed all eight candidates from the diffs and ran their tests outside the pipeline. Eight suites, six to nine tests each, all green, zero red. Everything passed. That is precisely the problem.

Two results I did not predict

The judge stabilised on real variance. In the contaminated series the judge’s two passes (A/B and B/A) agreed with each other in 2 cases out of 4. In the clean series — 4 out of 4. I predicted exactly the opposite: more diverse candidates were supposed to produce more inconclusive verdicts. The interpretation I reached after the fact: with near-identical candidates the judge was ranking noise — stylistic trifles it weighed differently in the two passes. With real differences it had something to discriminate on. So contamination not only faked the workers’ convergence; it also destabilised the judge. I am leaving my wrong prediction in plain view, because it is proof the predictions were written before the runs, not fitted after them.

Questions were asked. The system threw them away. Opus, in eight runs out of eight — both series — ended its work with a message that explicitly named its assumptions:

The spec says discounts “do not stack” but doesn’t say which wins. […] If you want the customer-favourable rule instead, it’s a one-line change.

That’s a product decision I didn’t want to invent.

The return is the discount amount, not the final price — the task says “calculates a discount”.

Eight times the pipeline gave me exactly the questions whose absence I wrote about in the pilot. And eight times they died in the run’s result field, because a worker’s final message is the input of no stage: the readiness check sees a diff and a checklist, the pairwise comparison sees two diffs, I see a verdict in the panel. So I have to correct my own thesis. This is not a “nobody asked” problem. It is a “the question has nowhere to go” problem. My task template has an OPEN DECISIONS field on the way in — and the pipeline has none on the way out.

A cost irony for free: the more expensive worker (opus, $0.40–0.52 per run) produced questions nobody read. The cheaper one (grok, $0.02–0.05) won most verdicts — in series one partly because it copied.

Third layer: the measurement measured me

The journal from part two also records my decisions. I read them as data and did not come out better than the rest of the system.

I punished the invention I could see. I waved through the invention that is invisible. The only reject of both series I gave where the workers’ decisions materialised in the API surface: an invented code registry in one, a double parameter promo_code/promo_percent in the other. Rightly — but the invisible guesses, the threshold, the floats, the return contract, passed through my hands eight times without a single comment. A human in the loop reflexively reviews what is visible in the diff. The list of decisions nobody made does not appear in a diff.

The panel forced a preference on me that I did not have. When the judge ties, I have no “accept the tie” action — some code has to move on. So I picked branch A and, truthfully, wrote “without a reason” in the reason field. The trouble is that under my own architecture from part two every override is a labelled example for calibrating the judge. A coin flip recorded as a preference is noise injected straight into the calibration set. The fix is cheap — an “accept tie” action that passes a branch through mechanically but logs preference: none — and goes in before the next series.

A decision is not the same as a record of a decision. That reject was a good decision — the judge, in fact, named the problem first, and more precisely than I did:

candidate_2’s API has a correctness trap: promo_code carries no rate, so calling calculate_discount(1500, promo_code=’SAVE10’) without also passing promo_percent silently returns 0.0 discount (the code even tests this behavior as intended).

But my journal entry was a shorthand written on the run. For making the decision, intuition is entirely enough — decisions exist to direct traffic, and this one did. Only the journal has a second function: overrides and rejects are the material the judge is supposed to learn from. A “both bad” example without a precise reason teaches it nothing. The rule I take from this: I justify when I go against the judge, because only such examples teach the calibration anything; agreements stay blank.

And finally the subtlest find of the series, hidden inside that reject. calculate_discount(1500, promo_code="SAVE10") == 0.0 — the behaviour I rejected run four for — is identical to the behaviour I approved in runs 1 and 3. There too a promo code zeroes the result. The difference lives solely in the promise:

# run 1, worker A
def calculate_discount(amount, promo_code=None):
    """...when ``promo_code`` is given, the order is discounted through
    that code alone and this function grants nothing on top of it."""

# run 4, worker B
def calculate_discount(amount, promo_code=None, promo_percent=0):
    """...If a promo code is used (a truthy ``promo_code`` or a non-zero
    ``promo_percent``), only the promo discount is applied."""

Under the first promise zero is correct: promo pricing happens elsewhere, this function adds nothing beyond the code. Under the second the same zero becomes quietly robbing the customer — the function promises a promo discount, takes its rate as a separate parameter, and defaults that rate to zero. The same line, the same result, the opposite verdict, and what decides is a contract that exists only in the documentation. I am adding this to the taxonomy as class six — explicitly after the fact, outside the pre-registration: divergence between behaviour and the declared contract. A guess you cannot see even when you look at the code, because it lives in the gap between the code and the sentence above it.

Caveat: eight runs of one deliberately under-specified spec, one pair of models, one day — enough to show patterns, too little to quantify them. I am publishing the raw artifacts of both series with the post; the taxonomy and predictions were written before the second series, and the wrong prediction is left above in plain view.

The point

Eight runs, about 10 dollars, a full set of green tests, a judge one hundred percent consistent in the second series. And the number of unresolved business decisions did not drop by one. They only changed costumes: docstring, test name, expected value — and after the isolation fix two new ones arrived: a copied precedent and a contract in the documentation.

That list is the human part. Not “review the code” in general — specifically: the questions that neither the judge, nor the checklist, nor a green pytest can answer, because the answer does not exist in the repo or in the diff. It exists in the head of whoever ordered the work. No context window fixes this, because it is not a context problem. It is a problem of authority.

Which is why I still do not think the human’s share shrinks as models improve. It condenses. Across these eight tasks my work was single minutes — but, as the journal showed, they were minutes of looking at the wrong place: at the verdict, not at the list of calls nobody took ownership of.

What next

Three fixes, each following directly from the data, each as a separate, named intervention — because if I change everything at once I will not know what worked:

  1. An open_decisions channel: a mandatory section in both workers’ contract (named decisions made in code — not a count of questions, because Goodhart never sleeps), aggregated by the judge and visible at approval time. The verdict stops reading “winner: B” and starts reading “winner: B, decisions made by nobody: 4”.
  2. “Accept tie” in the panel, logging preference: none.
  3. The same experiment with proper input: a spec written according to my own task template — with acceptance criteria closing all six classes and a “decide yourself” field where I deliberately delegate the decision. We will see which classes good input closes, what still needs an eye on the output — and whether the template written for people really is, as I claimed, the same template for machines.

And then the numbers promised in part two return: twenty real tasks, judge–human agreement, costs — except the override counter gets a second column whose absence this post exposed: decisions not made. Because the first task I looked at truly carefully showed that I can agree with the judge on every verdict and still ship code full of decisions that nobody made.

That habit — asking what a result actually proves rather than what it appears to show — is still the same skill this whole series is about. In this post, for the first time, aimed at every layer of the system in turn.

$ prev
The judge that makes the decisions

We built a judge. Now we give it authority — and watch what can go wrong.

← read
$ next
I wrote the task properly. The judge called a tie four times

Same agents, same judge, only the task description changes. Eight solutions, all correct, nothing for the judge to decide. And then I find an order where they differ by a hundred złoty.

read →
↑