Skip to content
DevLaunch home

Guide · AI employees

An AI Task Failed. What Should Happen Next?

Design a recovery path that preserves useful work, checks what already happened, and gives the next person a clear starting point.

By DevLaunchPublished

A workflow can fail after doing something useful. It might finish the research and lose the connection while saving a draft. It might create a customer record but fail before recording that success. Hitting "run again" without checking the current state can turn one failure into a duplicate or a confusing mess.

Find the last confirmed result

Start by separating what you know happened from what the interface merely attempted. Look for a saved artifact, a destination record, or a provider receipt. A red status beside the overall task doesn't prove that every step failed.

For a hypothetical proposal workflow, the outline might be saved while the document export is missing. Keep the outline and retry the export. If the failure occurred during delivery, check the destination before sending again. Uncertainty about whether an external action happened deserves inspection, not an automatic second attempt.

Keep the inputs attached

A recovery note should include the task's original inputs, the version of the instructions, and the error from the failing step. Without those, a teammate may restart the job using a different brief and produce a different result without realizing it.

Record enough detail to reproduce the problem while keeping credentials and unnecessary customer information out of logs. A reference to the protected source record may be more appropriate than copying its contents into an activity feed everyone can read. Give reviewers access through the existing permission system.

Distinguish retryable failures from bad inputs

A temporary connection problem and a missing customer address call for different responses. Waiting and retrying might help the connection. It won't supply the missing address. Repeatedly running the same incomplete task just creates more noise.

Use a small set of recovery states a person can understand: retry available, waiting for information, needs review, and canceled. The exact names matter less than the action each one calls for. Include a reason and the next responsible person rather than leaving the task in a permanent "error" bucket.

Put a limit on automatic retries

Define which operations may be retried and how often. Repeating a read is different from repeating a message send or a record creation. For actions that must not happen twice, the application needs a way to recognize the same request and inspect any existing result.

You don't need an elaborate incident system for a first workflow. A bounded retry count, a clear stopping point, and a useful explanation can be enough to start. Test those rules with the actual providers and data path you're using. A retry policy written in a document doesn't enforce itself.

Leave a recovery handoff

When a task needs a person, give them a short account of the situation. "Export failed after the outline was saved; no delivery attempted" tells them much more than a generic warning icon.

Include the last confirmed step, where its output lives, what remains uncertain, and the action you're recommending. After recovery, keep the original failure in the history and record what fixed it. That history helps you notice a recurring problem and decide whether to change the workflow rather than asking people to rescue it every week.

Sources & further reading

Keep building

View topic →
Workflow guide

Run Your First Campaign with AI Employees

Plan a small campaign with Morgan, hand approved work to Luca and Miles, and review the editable outputs before publishing. Includes a reusable brief.