From "is the model answering well" to a pipeline where two agents write code, a judge from another model picks, and I count the decisions nobody made along the way.
A series about one question: can a model be allowed to decide. It starts with how to measure an LLM’s answers at all, then hands a judge power over two agents’ code, and from part three on every post is an experiment with numbers: predictions written down before the run, a registry of runs, the result, including when the prediction fell.
Read in order. Each part assumes the previous one, and the pipeline it all runs on grows from post to post.
Imagine grading an employee who writes a poem one day, solves an equation the next, and translates from Japanese on the third. One scale? Good luck.
We built a judge. Now we give it authority — and watch what can go wrong.
Eight runs, two agents, a judge with a double pass and a human at the end. All tests green. Along the way — a worker copying solutions from neighbouring tasks, three contradictory readings of one sentence of spec, and a counter of decisions nobody made.
Same agents, same judge, only the task description changes. Eight solutions, all correct, nothing for the judge to decide. And then I find an order where they differ by a hundred złoty.