polycratia

The matching job ran clean on my laptop, every single time. In production it left a pile of unmatched payments every morning, and someone cleared them by hand.

The service reconciled incoming transfers against expected obligations, grouped by settlement day. Locally everything matched. In production a thin slice of transactions landed on the wrong side of the day boundary and fell into the unmatched bucket.

Same code, same database dump. My laptop and the server just disagreed about when a day starts.

That disagreement was nowhere in the repository. Not in the config file either. It was ambient, inherited from the host, read silently at runtime, never written down anywhere a reviewer would look.

The bug ticket was the cheap part. The expensive part was the daily manual review, and what that does to trust: once a person starts hand fixing reconciliation output, they stop believing the matches that were correct too. You lose the automation long before you lose the data.

What I do differently now: I treat the environment as an input to the program, with the same seriousness as a request payload. Timezone, locale, encoding, decimal precision, all declared explicitly in the application, asserted at startup, and the process refuses to boot if the host disagrees. Then I run the test suite under a deliberately hostile environment (wrong timezone, wrong locale, different collation). If the suite still passes, the code is honest about what it depends on.

"Works on my machine" is rarely about the code. It is about the inputs your code reads without asking: clock, locale, encoding, filesystem case sensitivity, ordering guarantees. None of them show up in the diff.

I write up more of these production post-mortems on my blog, if that is your kind of reading: https://polycratia.com/c/62cf7fab

I'd guess you have one of these too, and it cost more than the ticket said: timezone, encoding, or something ordering-related that only surfaced under load...

react

$ new-project --brief

or email hey@polycratia.com