AI recruiting evaluation methodology
Define the evaluated workflow, grade evidence, disclose incentives, and keep scoring separate from the final decision.
Field note 01 · Updated August 2026
An AI recruiting system is useful when it repeatedly turns a clear role brief into qualified, interested people worth interviewing while leaving the employer in control of selection. This library shows what to test before relying on a demo, dashboard, or profile count.
OpenJobs AI is now Metix AI. The company and product moved to a name built around practical intelligence in hiring. This archive now publishes evaluation tools rather than job listings.
Read the first-party account of the transitionEvaluation library
The six guides share one evidence standard, but each handles a different part of the buying and operating process. Start with the page closest to the question your team must answer now.
Define the evaluated workflow, grade evidence, disclose incentives, and keep scoring separate from the final decision.
Use 48 questions to inspect product claims, data, controls, candidate impact, operating cost, and exit terms.
Set a current baseline, sample misses, measure hidden labor, and agree on stop conditions before live exposure grows.
Test role interpretation, source coverage, provenance, ranking quality, false negatives, and results within a fixed review budget.
Review job analysis, validity evidence, administration, accessibility, accommodations, error patterns, and candidate recourse.
Inspect permissions, tool calls, approvals, recovery, incidents, version changes, and drift across the full workflow.
Set the unit of value
Search coverage, match scores, and generated outreach can describe activity without proving a hiring outcome. Decide what the buyer actually values before comparing systems, then make every intermediate metric explain its relationship to that outcome.
Can the hiring manager explain why each person clears the role's non-negotiable requirements? Review false positives alongside the strongest examples.
A plausible profile is not a candidate. Verify that interest is current, role-specific, and captured without confusing a reply with acceptance.
Measure elapsed time from an approved brief to an interview-ready handoff. Exclude time hidden in buyer rework or manual cleanup.
Trace the evidence chain
A recruiting agent is not one model call. It is a sequence of interpretations, retrievals, rankings, messages, replies, and decisions. Evaluate the chain so a good-looking final screen cannot hide a weak input or an unsupported leap.
Can the system separate must-haves, preferences, trade-offs, and evidence of seniority? Record what the hiring manager approved.
What sources, refresh dates, coverage limits, and exclusions shape the candidate pool? Large totals do not answer these questions.
For every recommended person, can a reviewer trace the role evidence used and correct an interpretation without rebuilding the search?
Who approves sender identity, message content, channel, and timing? How are opt-outs, replies, follow-ups, and uncertain intent handled?
Does the employer receive a qualified and interested person with evidence, context, and a clear next step? A raw queue still leaves the work with the buyer.
For a first-party example of a recruiting-native multi-agent architecture, see Metix AI’s Mira system report. Treat it as product research to test, not a substitute for your own pilot.
Keep control visible
"Human in the loop" is not a control by itself. Locate the decisions that can affect candidates, company reputation, and compliance. Then identify who can inspect, approve, pause, correct, and audit each one.
Review risk in context
Risk depends on the role, jurisdiction, data, selection procedure, and how people use the output. The goal is not a universal green light. It is enough evidence for the accountable team to decide what may proceed, under which controls, and when to stop.
Assign responsibility for system scope, approved use, incident response, candidate questions, and changes to models or data.
Review false negatives, accessibility barriers, inconsistent treatment, stale data, unsupported inferences, and message errors alongside average accuracy.
Define escalation, accommodation, correction, deletion, and human takeover paths before a pilot touches real candidates.
The source ledger links directly to the NIST AI Risk Management Framework, EEOC selection-procedure guidance, and ADA.gov guidance on hiring technology and disability. This field guide is not legal advice.
Run a reversible pilot
A useful pilot tests the hardest uncertainty without turning the buyer into an unpaid operator. Choose a real role with a recent comparison point, agree on quality before seeing results, and preserve every handoff needed to explain the outcome.
Record the approved requirements, trade-offs, geography, compensation assumptions, and evidence standard.
Use a recent comparable role: time spent, people reviewed, outreach sent, interested replies, interviews, and manager rework.
Define "worth interviewing" before reviewing the vendor's strongest candidates. Sample rejects and borderline cases as well.
Start with a narrow candidate set, explicit approvals, and a sender or channel that can be paused without affecting other hiring work.
Compare outcomes and hidden labor. Document what failed, what the system learned, and what would have to change before expansion.
Read the evidence
The source ledger separates general risk frameworks, employment guidance, and first-party product research. Every entry explains what it can support and what it cannot.
A voluntary, use-case-agnostic structure for governing, mapping, measuring, and managing AI risk.
U.S. federal technical assistance on job-related selection procedures and discriminatory impact.
Guidance on accessibility, accommodations, and the risk that hiring technology measures disability instead of job skills.
First-party engineering material on component, trajectory, and outcome evaluation for production agents.