Archive / Topic

#retries

16 posts, newest first

A retry budget should name its final answer

A caller needs a useful outcome when the retry window closes, especially when the remote result is still uncertain

API designretriestimeouts

Return the first outcome when an idempotency key repeats

Duplicate suppression is only half the contract; callers also need a stable account of what the original request produced

idempotencyAPI designretries

An idempotency key can expire before its risk does

Retry safety depends on the lifetime of the original side effect, not only the convenience of a short deduplication window

idempotencyretriesdistributed systems

Separate request identity from attempt identity

Retries stay auditable when the customer's command keeps one identity and every execution receives another

retriesidempotencyobservability

A dead letter needs an owner before it needs a replay button

Replay is an operational decision about unresolved side effects, not a convenience action for clearing a queue

dead lettersretriesqueues

Timeout budgets should leave room for the caller

A downstream deadline that consumes the entire request window leaves the upstream service no time to classify, record, or recover from the result

timeoutsdeadlinesretries

If lookup still works, do not replay yet

A retry becomes much less trustworthy when the system still has a live remote lookup path but chooses to spend certainty later instead of using it now

retrieslookupsreceipts

When retries stop, keep the last remote evidence

A failed delivery is easier to recover, explain, and replay safely when the dead-letter handoff preserves the last thing the remote system actually told us

retriesdead-letter queuesreceipts

Ask for the lookup path before you ask for a retry

A delivery system becomes easier to trust when manual recovery starts with evidence lookup, not with another send button

reliabilityoperationsretries

After a provider timeout, protect the acceptance boundary

Recovery gets more dangerous when a timeout erases whether the remote side might already have accepted the request

timeoutsdelivery semanticsretries

Flat error messages are weak replay evidence

A delivery system gets recovery wrong when a flattened failure message quietly steers the next operator toward duplicate side effects

incident responseretriesidempotency

When retries stop, leave a handoff worth reading

When automatic delivery gives up, the useful system preserves what was tried, what stayed uncertain, and what the next operator needs before touching replay

retriesreliabilityoperations

Rate limit headers should describe the next safe attempt

A 429 response is more useful when it tells the caller how recovery actually works, not only that a threshold was crossed

rate limitingapi designretries

Idempotency needs a published lifetime

The retention window is part of the contract, because replay safety ends the moment the original decision record disappears

idempotencyapi designretries

Receipts should outlive the retry budget

A delivery record stays useful only if the acknowledgement evidence survives long enough for replay, support, and audit decisions

reliabilityreceiptsretries

Timeouts should keep their doubt

A delivery timeout should preserve uncertainty clearly enough that the next operator does not confuse missing evidence with a confirmed failure

reliabilityretriesdelivery