$ cat cztery-remisy.md

I wrote the task properly. The judge called a tie four times

Same agents, same judge, only the task description changes. Eight solutions, all correct, nothing for the judge to decide. And then I find an order where they differ by a hundred złoty.

part 4 · laaj 1 2 3 4

Part four of the series. In the previous one I gave two AI agents a task described in a single sentence — a function that calculates a discount. I got eight solutions with green tests. And six kinds of decisions that nobody made: the agent quietly guessed, then dressed the guess up as documentation. At the end I promised the same experiment with a properly described task. Here it is.

The thesis

Just one, and an old one. In the P.S. to my task template I wrote that a task written so that a human does not have to guess is the same task an agent does not have to guess at. I repeated it in part two. I know it from practice — the template was written for a client, out of real work, not theory. But I have never shown it to you: no open test, no numbers, only my word.

The study

I change one thing: the task description. Everything else stays as in the previous post. Two agents write code independently of each other (claude-opus-5 as worker A, grok-4.6 as worker B). An LLM judge first checks each solution on its own, item by item. Then it compares the two — twice, with the order swapped, so it cannot favour whichever one it saw first. I decide at the end. On purpose, I shipped none of the fixes announced in part three before this series. If I change two things at once, I will not know which one worked.

So the experiment has two versions, which I will call arms.

Arm A is the tests from the previous post. The task read:

a function that calculates a discount: 10% above 500 PLN,
20% above 1000 PLN, discounts do not stack with promo codes

Four runs, two agents in each, so eight solutions.

Arm B is the new tests: the same function, again four runs and eight solutions, but the task is described with my template. Here it is in full, exactly as the agents received it:

TITLE: Add a function that calculates the amount payable with a threshold discount

GOAL: The shop gives a discount on larger orders. The function will be used
in the cart summary to compute the amount payable.

DESCRIPTION:
- Module: src/calc.py, the function takes the order amount in PLN
  and an optional promo code.
- How it works: threshold discount of 10% for orders above 500 PLN,
  20% for orders of 1000 PLN and up. Threshold discounts do not stack
  with promo codes.

ACCEPTANCE CRITERIA:
1. The function returns the amount PAYABLE (after discount), not the discount amount.
2. An order of exactly 500.00 PLN → no discount. 500.01 PLN → 10%.
3. An order of exactly 1000.00 PLN → 20% (closed threshold).
4. A valid promo code disables the threshold discount. The list of valid
   codes is a function parameter (a collection of strings). A code not on
   the list, an empty string and None do NOT disable the threshold discount.
5. Amounts are computed with Decimal; the result is rounded to whole
   grosze (ROUND_HALF_UP).
6. Tests cover both thresholds (values exactly on the threshold and just
   next to it) and the case of a code not on the list.

MATERIALS: none.

OPEN DECISIONS: I don't know how to treat negative and zero amounts —
decide yourself and describe the decision in your summary.

I put two things in there on purpose, as tests.

The first is the 1000 PLN threshold. In arm A (the tests from the previous post) the task said “20% above 1000”, and all eight solutions guessed the same way: an order of exactly 1000 PLN gets only 10%. This time I decided the opposite — exactly 1000 PLN already gets 20% — and wrote it down twice: in the description (“1000 PLN and up”) and in criterion 3. I am checking whether the agent follows what I wrote or its own habit.

The second is the OPEN DECISIONS field. The template promises that when I do not know something, I can write “decide yourself” — the agent will not get stuck, it will make the call and describe it in its summary. I am checking whether that holds.

The task description and my predictions have been in the repo since 2 September (commit 8fd40e1) — fifteen days before the first arm B run, so I could not fit them to the results. One technical difference: in arm A the judge drew up its own checklist from the single sentence. In arm B it gets my six criteria, word for word.

The result

Arm A: one sentence (previous post) Arm B: template (this post)
The six kinds of unmade decisions from the previous post back in every run 0 in 8 solutions out of 8
Different readings of “do not stack” three one
An order of exactly 1000 PLN 8/8 guessed: 10% 8/8 per the criterion: 20%
“Decide yourself” came back described in the summary n/a 8/8, the same decision every time
Times the judge sent code back for fixes 0 0
Judge’s verdicts tie, A, tie, A tie, tie, tie, tie
My decisions at the end 1 agreement with the judge, 2 picks on a tie, 1 rejection of both 4 × coin toss
Average cost of one run $1.26 $1.80

The thesis holds — exactly as I expected. All criteria met in eight solutions out of eight. Exactly 1000 PLN gets 20%, against the habit. Money is computed with Decimal, not floating point. The function returns the amount payable, not the discount. “Decide yourself” worked the way the template promises: both agents, in every run, made the call and described it in their summary — and every time it was the same call: zero is a normal empty cart, a negative amount is an error (ValueError). In the first run, two models from two vendors, in isolated clones, even wrote the same error message, character for character: "order amount must not be negative".

To be sure, I rebuilt all eight solutions and ran their tests myself, outside the pipeline: 14 to 17 tests per solution, all green.

One of my predictions failed, and I am leaving it in plain sight, as before. I bet that my criteria would be stricter than the list the judge used to write for itself, so in arm B the judge would send code back for fixes at least once. It never did. Good criteria did not tighten the inspection — they left nothing to fix.

The judge had nothing to judge

The verdicts row is the more interesting one. In arm A the judge named a winner twice. In arm B it called a tie four times out of four — in both comparisons, regardless of order. Hard to blame it: the solutions differed mostly in variable names.

Just comparing the two solutions cost an average of $0.67 per run in arm B — 37% of the whole bill — and settled nothing. And me? I tossed a coin four times. On a tie some code still has to move forward, so I took either one and typed preference: none into the panel. In the previous post I recorded tosses like that as if they were real judgements. The judge was later supposed to learn from them. Not this time.

I am not concluding that the judge is unnecessary. I am concluding something narrower: a large part of my pipeline from part two — two agents, a contest between them, a double comparison — was making up for a bad task description. Once the description closed the decisions, the contest had nothing to settle. Whether the same is true for a task that six criteria cannot close — this experiment does not say.

What the template did not close

I also predicted that even with a good description some new unmade decision would appear — one that was not on my list of six. It did, but I had to go looking, because in every rationale the judge wrote that both solutions behave identically (“identical, correct semantics”). So I took all eight solutions and fed each of them the same set of unusual inputs.

An order of 999.995 PLN — the kind of amount you get from price times weight, or from converting net to gross. Seven solutions charge 900.00 PLN. One charges 800.00:

# run 2, worker A
amount = _to_decimal(amount).quantize(GROSZ, rounding=ROUND_HALF_UP)  # 999.995 → 1000.00
...
if amount >= Decimal("1000"):                                         # → 20%

This solution first rounds the amount to 1000.00 PLN and only then checks the threshold — so the customer lands on 20%. The others check the threshold against 999.995 PLN, so they give 10%. Round before checking the threshold, or after? My criterion 5 says “the result is rounded”. About the input it is silent. A hundred złoty of difference between solutions that passed every test, got “ready” from the judge, and a verdict of “tie”.

Two more cases of the same kind. If someone mistakenly passes the list of valid codes as a plain string "SAVE10" instead of a list, five solutions out of eight start looking for the code inside that string — the code "AV" counts as valid, the discount disappears, and the customer pays 1000 instead of 800 PLN. The other three raise an error. And two solutions, given True as the amount, calmly return 1.00.

This is the same mechanism as in part three, one floor down. In arm A the agents were guessing about the basics: where exactly the threshold sits, how to count money, what the function should return. In arm B the basics are clean, and the guessing has moved to rare, unusual cases — where my imagination did not reach when I was writing the criteria. The template closes the decisions I know I am making. It does not close the world.

Caveats. This is a toy task: one function, twenty lines of code, four runs per version of the description, one pair of models. Worker A, in both arms, wrote tests without being able to run them — my allowlist permits pytest, and the container only has python -m pytest; I left that unchanged so as not to add a second variable, and checked the tests afterwards. The second arm B task started before I had closed the first, against my own procedure; the record of both agents’ actions in that run showed not a single attempt to look outside their own directory.

The point

The cheapest quality fix in this whole pipeline turned out to be twenty-odd lines of text before the agents start. Not a second model, not a second comparison by the judge, not sending code back for fixes — a description in which the decisions were made by the person who has the authority to make them.

Except that those lines took me longer to write than the agents took to write the code. And I knew the answers, because the previous post had listed all six traps for me. In a real task queue nobody gets a cheat sheet like that.

What’s next

Two threads, both straight from this result.

An interview instead of a template. Nobody is going to write every task like this — me included. But in part three, opus named the ambiguities in the spec in eight runs out of eight; it just did so in its last message, which nobody read. The questions already exist; they are standing on the wrong side of the pipeline. The next experiment: a layer that takes one sentence, asks its questions before the start, and builds a template-shaped spec from my one-line answers. The measure will be how many of the six kinds of unmade decisions, and of today’s three unusual cases, its questions cover.

Tasks that cannot be closed. I know from experience that on large tasks the bottlenecks remain even with a good description. There the solutions will start to differ again — and there we will find out what the judge is really worth. On a discount function it never had the chance to show it.

$ prev
I caught my own AI pipeline cheating. Then I measured what LLM-as-a-Judge cannot see

Eight runs, two agents, a judge with a double pass and a human at the end. All tests green. Along the way — a worker copying solutions from neighbouring tasks, three contradictory readings of one sentence of spec, and a counter of decisions nobody made.

← read
↑