Guide · Agent orchestration
Debug a Multi-Agent Run from the First Wrong Decision
Trace assignments, inputs, tool results, and accepted artifacts to find where a multi-agent workflow first went off course.
The final answer is often the worst place to start debugging an agent run. By then, a missing input may have passed through several summaries and become a confident conclusion. Work backward from the incorrect result until you find the first decision that no longer followed the evidence.
Keep a timeline you can reconstruct
Record the user request, assignment creation, tool calls, worker results, and acceptance decisions. Include identifiers that connect a worker's result to the task and artifact revision it belongs to. Timestamps alone are not enough when several jobs finish close together.
For a hypothetical inventory report, the timeline might show an export request, a changed date range, a second export, and two returning files. If the report used the first file, you have a state-management problem to investigate. The final model's writing quality has little to do with the failure.
Keep sensitive data out of routine logs where it is not needed. Record artifact references and the minimum useful context, with access appropriate to the project. A debugging record should help reproduce behavior without becoming an uncontrolled copy of customer records.
Inspect boundaries before changing prompts
Check what the worker actually received. Was the current date range in its assignment? Could it access the file? Did the tool return an error that the wrapper converted into an empty result? Was a completed task accepted after it had become obsolete?
These questions separate an instruction problem from a tool or application problem. Rewriting the coordinator prompt will not repair a wrapper that labels a failed request as success. Fix the earliest incorrect state and replay the affected path.
Save the configuration with the run
Record the model identifier, effort settings, prompt revision, tool definitions, and relevant application version. This matters when comparing GPT-6 Astra and Claude Fable 5.1 or updating either configuration. Without it, two runs that appear identical may have used different behavior or permissions.
Keep the original failed example as a regression case. Remove unnecessary private details while preserving the conditions that caused the failure. A smaller reproducible case is easier to rerun and easier for another developer to understand than a complete production transcript.
Confirm the fix at the same boundary
After a change, replay the failure and check the specific invariant that was broken. In the inventory example, only the export matching the current requested range should be accepted. Also check the normal case so the fix does not block legitimate results.
Write a short incident note with the observed failure, its cause, the change, and the evidence from the replay. If the original failure cannot be reproduced, say what remains uncertain. That record gives the next debugging session a reliable starting point instead of another round of speculation.