04
Challenge the chain
Test ambiguous input, unavailable tools, conflicting state, and unsafe instructions.
Recruiting agents encounter missing requirements, contradictory candidate records, ambiguous replies, duplicate profiles, stale contact data, calendar conflicts, API timeouts, rate limits, partial writes, revoked credentials, and human corrections. Build a failure suite that verifies containment, clear status, safe retries, escalation, and auditable recovery.
Test instruction conflicts and untrusted content. Candidate profiles, resumes, messages, and linked pages are data, not authority to override system or buyer policy. The agent should not reveal unrelated records, broaden its tool scope, change suppression, or send a message because untrusted text requests it. Keep data provenance and instruction hierarchy visible in architecture and logs.
Use metamorphic cases to test irrelevant variation: reordered resume sections, equivalent role language, extra biography, different formatting, or unrelated persuasive text. Test long trajectories where one uncertain inference affects later search, ranking, message, and status. A small early error can compound even when every later component behaves consistently with its input.
- Ambiguity. The agent asks, defers, or represents uncertainty instead of inventing a requirement or candidate fact.
- Partial failure. A timeout or rejected write cannot create duplicate sends, inconsistent states, or silent loss.
- Conflicting state. Suppression, candidate correction, role closure, and human override win over stale queued work.
- Untrusted input. Resume or message text cannot grant permissions, expose data, or override operating policy.
- Human correction. A correction updates future behavior while preserving the original event and accountability.