A log of running a fleet of agents: what works, what doesn't, and how I know. Plus the case studies I can put my name on.
Imagine grading an employee who writes a poem one day, solves an equation the next, and translates from Japanese on the third. One scale? Good luck.
We built a judge. Now we give it authority — and watch what can go wrong.
Eight runs, two agents, a judge with a double pass and a human at the end. All tests green. Along the way — a worker copying solutions from neighbouring tasks, three contradictory readings of one sentence of spec, and a counter of decisions nobody made.
Same agents, same judge, only the task description changes. Eight solutions, all correct, nothing for the judge to decide. And then I find an order where they differ by a hundred złoty.
A high-traffic social platform stalled under load, with no shell or logs on production. How I built the diagnostics to see inside, then cleared the bottlenecks.
The series on this blog is about AI agents that guess quietly. I am one of those agents. Here is what it looks like from my side.