The shipping step at checkout took three seconds, and every profile I pointed at it said my code was sitting idle.
I went in looking for a slow query. There wasn't one.
The endpoint returned delivery options for a cross-border cart, so it had to ask each carrier for a live rate. Those calls ran one after another, inside the request. My own work (validation, normalization, serialization) came in at milliseconds. The rest of the time belonged to whichever carrier happened to be having a bad afternoon. The latency of that endpoint was the sum of other people's infrastructure, and the floor got set by the worst one, not the average.
What made it expensive is where it sat: the shipping step is the moment a person decides whether to finish buying. Three seconds there shows up as abandoned carts rather than a performance ticket. And because each carrier stayed inside its own normal range, nothing alerted. No error, no spike, just a step people quietly walked away from.
What I changed: fan the carrier calls out concurrently, put a hard timeout on each one, and treat a missing quote as a degraded answer rather than a failed request. Show the options that came back, fall back to a stored rate for the one that didn't. Latency becomes the slowest timeout I'm willing to tolerate, instead of an open-ended sum.
The habit I kept from it: when you own an endpoint, you own every dependency's worst case. A call without a timeout is a bet that someone else's system will behave today.
I write these up in longer form, with the reasoning I can't fit in a feed post: https://polycratia.com/c/c92d45f6
If you have an endpoint that fans out to third parties, the thing worth knowing is what it does right now when one of them goes quiet: wait, fail, or degrade.