$ ls ./blog/laaj

LLM-as-a-Judge in practice

From "is the model answering well" to a pipeline where two agents write code, a judge from another model picks, and I count the decisions nobody made along the way.

A series about one question: can a model be allowed to decide. It starts with how to measure an LLM’s answers at all, then hands a judge power over two agents’ code, and from part three on every post is an experiment with numbers: predictions written down before the run, a registry of runs, the result, including when the prediction fell.

Read in order. Each part assumes the previous one, and the pipeline it all runs on grows from post to post.

$ series laaj
4 parts
↑