Release a withdrawal hold only when the rail says no
A withdrawal reserves funds before they leave, and something has to let that reservation go when the send does not work out. The trouble is that "did not work out" packs two unrelated facts into one phrase: a rail that refused the transaction, and a rail that never answered. Only the first one tells you the money is still yours. Code that treats both the same way will, sooner or later, hand a balance back to a user whose transfer is already confirming on-chain.
You do not discover this late. It is the first design decision in the submission path, and it shows up as a single except clause.
The hold is a claim on money that has not moved#
Between the moment you decide to send and the moment you know the send happened, the user's spendable balance is wrong unless you reserve against it. That reservation is the hold. In the pipeline in github.com/polycratia/withdrawals it sits in a fixed position, after the routing decision and before the send:
from withdrawals import Hold, Submissions, broadcast, ensure_approved, route
submissions = Submissions()
def withdraw(request, key, custody, holds, policy, rails):
ensure_approved(request, policy)
decision = route(request, custody)
holds.place(Hold.on(request))
send = rails[decision.rail]
return submissions.submit(
key, request, lambda pending: broadcast(pending, send, holds)
)The order matters in one direction only: the hold has to exist before the send, because a refusal needs something to give back. Place it after the send and a refused withdrawal has nothing to reverse. Place it before and every failure path has to answer the question this article is about.
Look at what route buys you here. A destination that custody recognises is an internal move: a ledger entry, no fee, nothing to confirm. There is no transaction whose fate can be unknown. The whole problem below lives on the external rail and nowhere else, which is a good reason for the routing decision to be a recorded value with a reason attached rather than an if buried in the sender.
Two failures that produce the same stack trace#
From inside a try block, these look alike:
- the node rejects the transaction: bad nonce, insufficient hot-wallet balance, destination not permitted, fee below the floor;
- the request times out, the connection resets, the load balancer returns a gateway error, or the process is killed between the socket write and the response.
In the first case the rail has made a statement: it did not accept this transaction. In the second the rail has said nothing at all. The transaction may be in a mempool. It may be mined. It may have never left your side of the socket. Nothing in the exception tells those outcomes apart, because the information does not exist yet.
So the type has to carry the distinction, at the boundary where it is still knowable. That is the one job of a rail adapter:
from withdrawals import BroadcastRejected
def send_external(pending):
try:
return node.broadcast_raw(
pending.destination, pending.amount, pending.currency
)
except node.Refused as exc:
raise BroadcastRejected(str(exc)) from excEverything else the node can raise propagates untranslated. That is deliberate, and it is the most consequential line of the integration: if you widen the except, catching the provider's base error class or matching on a message substring because the vendor's error codes are inconsistent, you have quietly declared that timeouts are refusals. The package cannot detect that mistake. It can only honour the claim the adapter makes.
The release belongs in an except clause, not a finally#
broadcast wraps the send with exactly one asymmetry. When BroadcastRejected comes out, it fails the request and releases the hold in the same call: one rail statement, both consequences, no window where the request is dead while the funds are still reserved. Any other exception leaves the hold where it is and travels on to the caller.
The instinctive version of this code is a finally that releases the hold on every exit path, because that reads like resource hygiene: open, use, clean up. A hold is not a resource. It is a statement about where money is. Releasing it credits a spendable balance, and on an external rail that credit does not come back: a confirmed on-chain transfer will not reverse because your service decided to clean up after itself. Holding too long costs a support ticket and one reconciliation cycle. Releasing too early costs the amount. Asymmetric costs make an asymmetric default, and the default has to be "keep holding".
The same reasoning explains why an unknown outcome does not get a tidy terminal state. The error propagates and the request stays where the submission left it: on its way. No transition would be true, so none is recorded. A request sitting in a non-terminal state with its hold intact is not an inconsistency waiting to be swept up, it is the system telling you, accurately, that it does not know yet.
The state chart will not let you call it a rejection#
The request is an immutable value, and the transitions encode the same rule. Rejection is reachable only while the money has not moved, from requested and from approved. Once the request is sending, the only way out other than success is failed:
sending = request.approve().start_sending()
failed = sending.fail('rail refused: destination not permitted')
assert failed.state is WithdrawalState.FAILED
assert failed.is_terminalCalling reject there raises InvalidTransition instead of producing a plausible-looking record. This is not bookkeeping pedantry: "rejected" means nothing was ever attempted, "failed" means something was attempted and did not complete. Those two words drive different conversations with a user, a compliance reviewer and an auditor, and the chart is the only place that difference gets enforced rather than remembered.
What the state does not say is where the money is. That is the hold's job. Two facts, two places: a request can be failed with its hold already released, or sending with its hold intact, and both are well-formed.
Holding is only safe because the send is idempotent#
Keeping a hold after an unknown failure is tolerable only if the client's retry cannot create a second transaction. Otherwise you are picking between duplicate payouts and premature releases, which is not a choice. submissions.submit(key, request, send) runs the send once per request key however many times the caller retries, so the retry after a timeout returns the original outcome instead of broadcasting again.
The settlement side needs the same property, because confirmation callbacks arrive twice:
sent = sending.mark_sent('0xdeadbeef')
assert sent.mark_sent('0xdeadbeef') is sentReplaying a transition that already happened returns the same value, so a duplicate callback costs nothing. A replay carrying a different reference raises InvalidTransition rather than overwriting the first one, which is what you want: two references for one request key is a real problem, and silently keeping the newer one destroys the evidence.
Who resolves the unknown#
Not the request path. The caller that got the timeout has nothing new to offer, and waiting on the socket longer only swaps one unknown for another. Resolution happens later, from the outside: query the rail by reference, or scan for a transfer matching the destination and amount, then either confirm the request or fail it and release the hold with a reason.
For transactions that were accepted and then stop moving, review does the time arithmetic against a StuckPolicy (bump_after and cancel_after) and returns the decision together with the reason behind it, as a value. It does not bump, cancel or log on your behalf. That separation is what makes the behaviour testable without a clock and auditable afterwards: the reason a withdrawal was bumped at a particular moment is a value you stored, not an inference from a log line.
What I would do differently#
In earlier payment work I let the service layer classify failures by reading provider error messages, because the providers' own error codes disagreed with each other and the strings were sitting right there. That works until a vendor rewords a message, and the failure mode runs in the expensive direction: an unrecognised refusal becomes an unknown, which is harmless, but an unrecognised timeout phrased like a rejection becomes a release. I would now put the classification in the adapter, where the vendor's semantics are still in view, and give it the narrowest surface I can: one exception type that means refused, and nothing else.
Two smaller things. I would make the release carry the reason from the start, so a released hold explains itself without a join against logs. And I would resist the single generic PaymentError: one exception class with a kind field feels economical, and it is probably the exact shape that lets a careless except swallow the distinction the whole design rests on.
Money that does not leave has to come back. It only has to come back when you know it did not leave, and the one place that knowledge exists is the boundary where the rail spoke or stayed silent.