$ cat czarna-skrzynka.md

Production was choking, and the server was a black box

A high-traffic social platform stalled under load, with no shell or logs on production. How I built the diagnostics to see inside, then cleared the bottlenecks.

Real production work, described without naming the client. Outcome: intermittent stalls under load → stable.

A high-traffic social platform would seize up under load every so often — intermittently, only when a lot of things happened at once, and never reproducible on demand. On top of that: no shell and no readable logs on production. The hardest kind of failure. I built the tools that caught it in the act, then defused the sources of contention one by one.

The short version — for everyone

I built this platform from zero to production, and I still run it today. At some point it started stalling at random — not always, only when a lot of activity overlapped. It would hang for a moment, then recover. That is the worst kind of problem: you cannot reproduce it on demand, so you cannot catch it by reading the code.

The cause was not one broken thing. It was a handful of perfectly ordinary operations that only became a problem when they collided in time — background cleanup taking locks exactly when users were active; a heavy query that is harmless on a calm database but stretches to 25 seconds on a busy one. Separately, invisible. Together, a pile-up.

I could not just look inside — the server was a black box — so I built my own X-ray: tools that do not just time slow actions but take a snapshot of the whole concurrency picture at the moment of an incident, showing who is blocking whom. That is what finally showed the truth.

With the data in hand, I defused the amplifiers one by one, so overlapping processes stopped tipping the system over: I cut the registration query from roughly 2,250 round-trips to one, moved heavy cleanup off the live path, added a missing index, and mapped the rest — locks, missing timeouts, the connection pool — into a prioritized plan. I also left a written runbook, so the next slowdown is a 10-minute diagnosis instead of a week of guessing.

This is not a story about one bug. It is the normal stage in the life of any system that has grown: things that worked great at low traffic start overlapping at high traffic. My job is to make that visible and defuse it — without guessing and without panicking.

The hardest failures are not bugs in the code — they are behaviors that only surface when everything happens at once. My job is to make them visible.

Under the hood — for engineers

Django/ASGI (Channels) behind PgBouncer in transaction mode, on managed Postgres. No shell and no readable logs on production, so the database is the only readable sink: all diagnostics write to a DiagnosticLog table, reachable over HTTPS by a superuser. The key point — these were not slow individual requests, they were contention under concurrency, so the tools had to capture the state of the whole system at the moment of an incident, not just the duration of one request.

The tools I built to see anything at all

  • Slow-request middleware — logs every request over 1000 ms with duration_ms, query count and db_time_ms, so you instantly know whether the bottleneck is the DB or Python. Disabled, it is MiddlewareNotUsed: zero overhead.
  • Stage-timing decorators — on the image-send path (WebSocket traffic does not pass through HTTP middleware); records stages of 100 ms or more, or errors. Recording can never break the path it measures.
  • A minutely pg_stat_activity sampler — a celery beat task that captures lock chains with the blocked query and the blocking PID, oldest_xact_s, idle_in_tx. Plus a health report: dead-tuple ratio, autovacuum, statement_timeout, top pg_stat_statements.
  • A runbook — a day-one workflow and a ranked list of risks beyond this incident (time limits in Celery and Postgres, the PgBouncer pool), so the next incident is a procedure, not a fresh investigation. The point was never to patch one query — it was to map the whole failure surface.

What the data showed (concurrency, not single queries)

  • Background cleanup — an unbatched whole-room message delete with cascades, and an hourly reconcile with COUNT(*) subqueries — held locks on messenger_message and chatroom exactly while users were active. That is where the random hung sends came from.
  • An N+1 on registration: an .exists() per candidate in a loop = roughly 2,250 queries. Tolerable on a calm database, 25 s under load. Collapsed to one membership query plus a set intersection.
  • Autovacuum had never run on the largest tables. The health report showed last_autovacuum = NEVER since June’s stats reset on the actions table (5.1 GB), the messages table (3 GB) and several more. The mechanism: autovacuum thresholds are a percentage of the table (5% here), so at tens of millions of rows the cleanup was waiting for millions of dead tuples — and a nightly delete() of the entire notifications table added more every night. Result: the actions table alone accounted for 87% of disk reads on production. An agent I handed the sampler’s health reports to caught it — I was staring at the lock chains and missed it. From the same report, for completeness: 157 GB of temp files since June (work_mem too small), idle_in_transaction_session_timeout set to 24 hours, which is to say none, and statement_timeout at zero.

Fixes

Heavy deletes batched and moved to Celery; a covering index on photo_request; the N+1 collapsed to a single query; offer search moved to Postgres full-text (GIN) plus trigram fuzzy matching instead of LIKE. For the actions table: retention instead of endless growth — a relation-free archive for the reports, rows moved in batches of 5,000 with a pause, resumable, with a guard for rows the transaction ledger still references (foreign keys discovered, as usual, along the way). On a production clone: 79% of the table eligible to go. Separate, lower autovacuum thresholds for those tables I set from the console, directly on the production database — cleanup has actually been running on them since.


Got a system that stalls under load? Say hi.

↑