polycratia

Three seconds to respond, and mostly during the morning peak.

Profiling the handler turned up nothing worth fixing: a handful of queries, all indexed, all fast on their own. So I stopped asking what the request was doing and asked what it was waiting on.

It was waiting on a database connection.

The handler opened a transaction, wrote a verification record, then called the external verification provider over HTTP inside that same transaction. Most of the day the provider answered fast enough that nobody noticed. Under load it answered in seconds, and every one of those requests held a Postgres connection for the whole round trip. Once the pool was saturated, endpoints with nothing to do with verification started timing out too.

The slow handler was the symptom. The shared pool was the blast radius.

What fixed it was ordering, not optimization. Commit the intent first, with an explicit pending state. Make the provider call outside any transaction. Write the result back in a second short transaction, and hand a reconciliation path ownership of whatever never came back.

One rule has been non-negotiable in anything I build since: no third-party network call inside a transaction. Not a slow one, not a fast one. Someone else's p99 becomes your p50 the moment you hold a lock across it.

I write up more production post-mortems like this one on my blog: https://polycratia.com/c/31e1dc5f

Worth checking where your request path still holds a connection open while waiting on someone else's API...

react

$ new-project --brief

or email hey@polycratia.com