Guide · Agent orchestration
Retry Agent Tasks Without Sending or Creating Things Twice
Separate safe retries from uncertain external actions, and preserve completed work when a multi-agent step fails.
A timeout tells you that a response did not arrive in time. It does not always tell you whether the external action happened. If an agent retries a customer message or a record creation blindly, a temporary failure can become a duplicate that someone has to clean up later.
Classify the failed step
Start by identifying whether the step was a read, a local computation, or a write to another system. A failed document lookup often has different retry consequences from a failed invoice creation. Keep the action type and its inputs in the task record.
For a hypothetical onboarding workflow, reading a customer's existing setup can usually be attempted again within the service's limits. Creating the customer's workspace needs more care. If the create request timed out after the service accepted it, a second request may create another workspace.
Give each logical action an identity
Where the external service supports idempotency, use one stable key for the same logical action across retries. A new key on every attempt defeats that protection. Check the service's documented behavior and retention window before relying on it.
Also record the external identifier returned by a successful action. If the response is uncertain, look up the action's status through the service when possible. A worker should be able to report "the result is unknown" without the coordinator interpreting that as permission to start again.
Keep this logic in the tool or application layer. The model can explain what it is trying to do, but a durable action record should survive a worker restart, a model switch, or an interrupted conversation.
Retry only the unfinished work
Suppose research and drafting succeeded, but saving the approved result failed. Retry the save using the accepted draft. Do not restart the whole workflow and produce a different document unless the input or requirement changed.
Set an attempt limit and distinguish failures that may recover from failures that need a new decision. An unavailable service may justify another attempt later. Missing permission or invalid input usually needs correction before retrying. Keep the original error available so the next worker can see what actually happened.
Return a recoverable stop
When the retry limit is reached, preserve the accepted artifacts, the uncertain action, and the evidence needed to inspect its status. State whether continuing could duplicate a write. That gives an operator a specific next step instead of an unhelpful "task failed" message.
Rehearse this with a controlled timeout in a test environment. Check that the application does not create a second logical action, that completed work remains available, and that a resumed worker receives the correct state. The failure path deserves attention because it is exactly where a confident agent can make the wrong next move.