polycratia

Your database is fast because it is the only dependency you own. Everything else in the request path belongs to someone else.

That is how a 4ms query and a 900ms endpoint end up sitting happily in the same trace. In the systems I build (multi-carrier delivery, card and account-to-account rails, KYC providers) one request fans out to third parties, and your p99 turns into the sum of other people's p99s, not the max. Three sequential upstream calls at 300ms each: slow endpoint, fast database.

The mechanics that actually bite:

Sequential fan-out when the calls do not depend on each other.

A connection pool sized for the database instead of the slowest upstream, so requests that are waiting hold slots they are not using.

Timeouts longer than the caller's patience, so you keep doing work nobody will read.

Retries stacked on a degraded provider, multiplying load at the exact moment it is failing.

How I approach it: measure time per dependency inside the request, not just the total. Parallelize the independent calls. Set the budget per request, not per call. And move anything not needed to produce the response out of the path entirely, since delivery rates can be precomputed and verification can be asynchronous with a state the client polls.

The part worth keeping is that an endpoint is only as fast as the slowest thing you do not control, and your real lever is deciding what stays in the path at all, not optimization.

I write about this kind of trade-off in more depth here: https://polycratia.com/c/67ce4be5

So your request budget goes somewhere: one slow provider, or five fast ones in a row. Probably worth knowing which before the trace tells you...

react

$ new-project --brief

or email hey@polycratia.com