Module 5.3 · Topic 4
Error Recovery and Resilience
Bottom Line Up Front: Agent failures fall into predictable patterns. Diagnosing root cause informs whether you need new instructions, corrected context, or a different approach. Design workflows with error tolerance so…
4.1 Common Agent Failure Modes
Agents fail in consistent patterns. Early recognition helps you diagnose root causes accurately.
- Misunderstanding objectives: Agent understands parts but misses intent. Asked for strategic analysis, returns tactical procurement. Logic is sound given misinterpretation; problem is interpretation.
- Hallucination: Agent states false facts confidently. "Revenue $2.3B" when actually $230M. Sounds authoritative but is wrong.
- Incomplete execution: Agent completes 80% and stops. Specified 10 points; provides 7. Signals capability limit reached.
- Format violation: Asked for table; returns prose. Requested 3 options; lists 8. Signals constraint parsing failure.
- Reasoning error: Logic chain breaks. Correctly identifies trends but misapplies them. Reasoning is visible but flawed.
4.2 Diagnosing What Went Wrong
When agents fail, diagnose before demanding a redo. Understanding why failure occurred determines whether corrected input fixes it or a different approach is needed.
Re-read your original prompt. "Analyze customer segment" is vague; the agent might analyze profitability, demographics, or loyalty. Check facts: hallucination, outdated data, or missing spec? Look at scope: did the agent exceed it or undershoot? Different errors need different corrections.
4.3 Corrective Prompting and Re-Delegation
Corrective prompting is more efficient than starting over. Provide corrected information and regenerate.
Logic behind this approach:
Agents incorporate corrected input and adjust output while preserving valid prior work. Narrowing change to the failed element is faster than full restart.
Sample prompt:
What to expect in reply:
Agent adds missing two competitors, reorganizes by price, and preserves other analysis intact.
4.4 Designing Workflows That Tolerate Agent Errors
Design workflows that absorb agent errors. Use parallel paths, verification gates, and fallback options from the start.
- Parallel execution: Run agent work in parallel with human backup for high-stakes work.
- Verification gates: Evaluate deliverables before downstream delivery. Failed work goes back, not cascading.
- Bounded autonomy: Low-stakes tasks run autonomous; high-stakes have approval gates.
- Rollback: Stage output in sandbox. Revert if wrong without destroying dependent work.
- Escalation: Agent flags cases where it hits limits for human review.