Notes on backend engineering, infrastructure, and things I learned the hard way.
A look at when a simpler queue actually beats a distributed log, and the migration path that got us there without downtime.
The tools, dead ends, and eventual root cause behind a leak that never once reproduced locally.
Zero-downtime migration patterns for tables with hundreds of millions of rows, and when to just accept a maintenance window.
The design decisions that make 3am pages rare, and the ones that guarantee you'll see them every week.
A running collection of infrastructure-as-code building blocks that save a day of setup on every new client engagement.