polycratia

Cutting an API from 3 seconds to 200ms is rarely optimization. It's a decision about what the endpoint is allowed to wait for.

The 3 seconds rarely belong to your code. In a delivery-heavy marketplace endpoint most of it was fan-out to external providers: carrier rate quotes, one call after another, each on someone else's latency and none of them under my control. Profile the handler and you see almost nothing, because the handler is mostly sitting there waiting.

Three things changed it.

First, the calls stopped being sequential. Rate quotes don't depend on each other, so they run concurrently under one budget for the whole request.

Second, the request got a deadline instead of a timeout per call. When the budget runs out the response returns what it has, and the missing carriers get marked unavailable rather than silently dropped.

Third, everything stable moved out of the request path. Anything that doesn't change per user gets computed ahead of time and read, not derived.

So 200ms is not a faster version of 3 seconds. It's a narrower promise: a bounded answer that may be incomplete, instead of a complete answer with no bound. That trade has to be visible in the response contract, or the client will treat a partial result as the whole truth and quietly show a customer the wrong set of options.

I write up more of these trade-offs on my blog: https://polycratia.com/c/04714806

When you cut latency like this you gave something up, probably completeness, freshness or accuracy...

react

$ new-project --brief

or email hey@polycratia.com