An agent that passes its tests has proven one thing: it can act on a world that stopped moving.
That is the entire gap. A fixture is frozen, so the listing is still available at step 7 because nothing could take it between step 2 and step 7. Production does not hold still. Between the read that informed the plan and the write that carries it out, the balance got spent, the price changed, the order was filled, the document was countersigned by the other party.
The agent does not crash on this. It does something worse: it executes a plan that was correct a few seconds ago, confidently, with a clean log.
Two mechanics make it expensive.
Preconditions get checked once, at planning time, where they are cheap to evaluate and meaningless to guarantee.
And side effects get scattered across the steps, so there is no single place where the world is touched, and so no single place to re-check it before touching.
How I build around it: the planning stage reads and decides, and does nothing else. Every write lands in one executing step that re-validates the preconditions it actually depends on at the moment it acts, against the state version the plan was built on, and it refuses instead of improvising when that version has moved. A refusal you can see beats an adaptation you cannot.
This is not an LLM problem. It is the same discipline payments and marketplace backends have needed for years: conditional writes, explicit state versions, and one owner of side effects. Agents just made it easy to skip, probably because the reasoning step sounds authoritative enough that nobody re-asks the database.
Staging proves correctness under a frozen world. Production asks for correctness under a moving one. Those are not the same test.
I wrote up the mechanics in more depth here: https://polycratia.com/c/1ebc5697
For those of you running agents with real side effects: where did yours first act on state that had already moved, balances, inventory, or document state?