polycratia

A worker timed out partway through a batch of product descriptions. The task retried. The catalog came out with two versions of the same item.

Nothing threw. Both runs produced good output.

The pipeline adapts a retail catalog for another locale: source fields in, localized fields out, at catalog scale. In testing I checked the step item by item, feeding a source field, reading the result, judging it. Correct every time. What testing never exercised was the second attempt.

In production that step lives inside a queue. Timeouts, redeploys, a broker hiccup: a retry is normal operations there, not an incident. And the model is not a pure function. Re-adapt the same source field and you get wording that is close but not identical.

Downstream I was keying on the adapted text itself, for dedup and cache lookups, and for change detection on top of those. Close but not identical is a different key. So one retried batch looked like new products and changed products at the same time.

The cost was hours of reconciling a live catalog by hand, plus a stretch where I could not trust my own change feed.

What I do differently now: I treat a model call like a payment request. The output gets persisted before the step completes, keyed by source hash plus prompt version plus target locale, and the step reads through that store. A retry replays the stored result instead of generating a new one. Regeneration becomes an explicit decision, a prompt version bump, and not an accident of infrastructure.

Testing proves the step works. Production asks what happens when it runs twice.

I write up more of these testing-versus-production gaps here: https://polycratia.com/c/b3e44660

I'd still like to know where you pin non-determinism in your pipelines: at the model call, or at the write...

react

$ new-project --brief

or email hey@polycratia.com