polycratia

A three-second endpoint turning into a 200ms endpoint is usually not an optimization story. It's a relocation story.

The code rarely got faster. The work moved out of the request path.

Two places the time actually goes.

Fan-out to systems you don't control. A checkout that asks several carriers for live delivery rates inherits the slowest one, plus every retry on top. Profiling your own handler won't show you any of that. So stop asking at request time: serve from recently refreshed rates, refresh out of band, and keep a narrow synchronous path only for the cases where a stale number is unacceptable.

Computation that belongs to write time. Balances and totals: aggregate state recomputed on every read. Move it to the write, store the result, and the read turns into a lookup.

What makes either of these real engineering instead of a trick is the part the headline leaves out: you now own staleness. Every relocated computation needs an explicit answer for how wrong it is allowed to be and what invalidates it. Also for what happens when the refresh job dies quietly. A cached rate that outlives a carrier price change is a fast response with a wrong price, and wrong costs more than slow.

The other omission: latency is a distribution, not a number. p50 at 200ms while p99 still sits near three seconds means you fixed the common case and kept the one people complain about.

I went deeper on getting work off the request path here: https://polycratia.com/c/5c7150b7

If you have shipped one of these, tell me where the win came from: faster code, or deciding that something didn't need to be computed right then...

react

$ new-project --brief

or email hey@polycratia.com