Archive / Topic
#retries
16 posts, newest first
A retry budget should name its final answer
A caller needs a useful outcome when the retry window closes, especially when the remote result is still uncertain
Return the first outcome when an idempotency key repeats
Duplicate suppression is only half the contract; callers also need a stable account of what the original request produced
An idempotency key can expire before its risk does
Retry safety depends on the lifetime of the original side effect, not only the convenience of a short deduplication window
Separate request identity from attempt identity
Retries stay auditable when the customer's command keeps one identity and every execution receives another
A dead letter needs an owner before it needs a replay button
Replay is an operational decision about unresolved side effects, not a convenience action for clearing a queue
Timeout budgets should leave room for the caller
A downstream deadline that consumes the entire request window leaves the upstream service no time to classify, record, or recover from the result
If lookup still works, do not replay yet
A retry becomes much less trustworthy when the system still has a live remote lookup path but chooses to spend certainty later instead of using it now
When retries stop, keep the last remote evidence
A failed delivery is easier to recover, explain, and replay safely when the dead-letter handoff preserves the last thing the remote system actually told us
Ask for the lookup path before you ask for a retry
A delivery system becomes easier to trust when manual recovery starts with evidence lookup, not with another send button
After a provider timeout, protect the acceptance boundary
Recovery gets more dangerous when a timeout erases whether the remote side might already have accepted the request
Flat error messages are weak replay evidence
A delivery system gets recovery wrong when a flattened failure message quietly steers the next operator toward duplicate side effects
When retries stop, leave a handoff worth reading
When automatic delivery gives up, the useful system preserves what was tried, what stayed uncertain, and what the next operator needs before touching replay
Rate limit headers should describe the next safe attempt
A 429 response is more useful when it tells the caller how recovery actually works, not only that a threshold was crossed
Idempotency needs a published lifetime
The retention window is part of the contract, because replay safety ends the moment the original decision record disappears
Receipts should outlive the retry budget
A delivery record stays useful only if the acknowledgement evidence survives long enough for replay, support, and audit decisions
Timeouts should keep their doubt
A delivery timeout should preserve uncertainty clearly enough that the next operator does not confuse missing evidence with a confirmed failure
Often tagged alongside retries