Field note 01 · Updated August 2026

Evaluate AI recruiting through interview outcomes.

An AI recruiting system is useful when it repeatedly turns a clear role brief into qualified, interested people worth interviewing while leaving the employer in control of selection. This library shows what to test before relying on a demo, dashboard, or profile count.

A note on the name

OpenJobs AI is now Metix AI. The company and product moved to a name built around practical intelligence in hiring. This archive now publishes evaluation tools rather than job listings.

Read the first-party account of the transition

Evaluation library

Choose the guide for the decision in front of you.

The six guides share one evidence standard, but each handles a different part of the buying and operating process. Start with the page closest to the question your team must answer now.

Procurement

AI recruiting vendor checklist

Use 48 questions to inspect product claims, data, controls, candidate impact, operating cost, and exit terms.

Pilot

AI recruiting pilot design

Set a current baseline, sample misses, measure hidden labor, and agree on stop conditions before live exposure grows.

Sourcing

AI candidate sourcing and ranking

Test role interpretation, source coverage, provenance, ranking quality, false negatives, and results within a fixed review budget.

Screening

AI candidate screening

Review job analysis, validity evidence, administration, accessibility, accommodations, error patterns, and candidate recourse.

Set the unit of value

Measure progress toward an accepted interview.

Search coverage, match scores, and generated outreach can describe activity without proving a hiring outcome. Decide what the buyer actually values before comparing systems, then make every intermediate metric explain its relationship to that outcome.

A · QUALITY

Worth meeting

Can the hiring manager explain why each person clears the role's non-negotiable requirements? Review false positives alongside the strongest examples.

B · INTENT

Actually interested

A plausible profile is not a candidate. Verify that interest is current, role-specific, and captured without confusing a reply with acceptance.

C · TIME

Ready to schedule

Measure elapsed time from an approved brief to an interview-ready handoff. Exclude time hidden in buyer rework or manual cleanup.

Trace the evidence chain

Ask what changed at every handoff.

A recruiting agent is not one model call. It is a sequence of interpretations, retrievals, rankings, messages, replies, and decisions. Evaluate the chain so a good-looking final screen cannot hide a weak input or an unsupported leap.

1

Brief

Can the system separate must-haves, preferences, trade-offs, and evidence of seniority? Record what the hiring manager approved.

2

Search

What sources, refresh dates, coverage limits, and exclusions shape the candidate pool? Large totals do not answer these questions.

3

Match

For every recommended person, can a reviewer trace the role evidence used and correct an interpretation without rebuilding the search?

4

Engage

Who approves sender identity, message content, channel, and timing? How are opt-outs, replies, follow-ups, and uncertain intent handled?

5

Handoff

Does the employer receive a qualified and interested person with evidence, context, and a clear next step? A raw queue still leaves the work with the buyer.

For a first-party example of a recruiting-native multi-agent architecture, see Metix AI’s Mira system report. Treat it as product research to test, not a substitute for your own pilot.

Keep control visible

Find the person who can stop or correct each action.

"Human in the loop" is not a control by itself. Locate the decisions that can affect candidates, company reputation, and compliance. Then identify who can inspect, approve, pause, correct, and audit each one.

ApprovalNothing external happens before the right person can review it.
CorrectionA bad assumption can be fixed without restarting the workflow.
TraceabilityInputs, recommendations, edits, sends, replies, and handoffs remain distinguishable.
FallbackThe team can pause automation and finish the process manually.

Review risk in context

A vendor checklist cannot make the employer’s decision for them.

Risk depends on the role, jurisdiction, data, selection procedure, and how people use the output. The goal is not a universal green light. It is enough evidence for the accountable team to decide what may proceed, under which controls, and when to stop.

GOVERN

Name the owner

Assign responsibility for system scope, approved use, incident response, candidate questions, and changes to models or data.

MEASURE

Test real failure modes

Review false negatives, accessibility barriers, inconsistent treatment, stale data, unsupported inferences, and message errors alongside average accuracy.

MANAGE

Make reversal possible

Define escalation, accommodation, correction, deletion, and human takeover paths before a pilot touches real candidates.

The source ledger links directly to the NIST AI Risk Management Framework, EEOC selection-procedure guidance, and ADA.gov guidance on hiring technology and disability. This field guide is not legal advice.

Run a reversible pilot

Run one real role against a recent baseline and agreed stop conditions.

A useful pilot tests the hardest uncertainty without turning the buyer into an unpaid operator. Choose a real role with a recent comparison point, agree on quality before seeing results, and preserve every handoff needed to explain the outcome.

A

Freeze the brief

Record the approved requirements, trade-offs, geography, compensation assumptions, and evidence standard.

B

Set the baseline

Use a recent comparable role: time spent, people reviewed, outreach sent, interested replies, interviews, and manager rework.

C

Pre-score quality

Define "worth interviewing" before reviewing the vendor's strongest candidates. Sample rejects and borderline cases as well.

D

Limit exposure

Start with a narrow candidate set, explicit approvals, and a sender or channel that can be paused without affecting other hiring work.

E

Close the loop

Compare outcomes and hidden labor. Document what failed, what the system learned, and what would have to change before expansion.

Read the evidence

Check the source type before using a claim.

The source ledger separates general risk frameworks, employment guidance, and first-party product research. Every entry explains what it can support and what it cannot.

NIST

Artificial Intelligence Risk Management Framework

A voluntary, use-case-agnostic structure for governing, mapping, measuring, and managing AI risk.

Open source
EEOC

Employment Tests and Selection Procedures

U.S. federal technical assistance on job-related selection procedures and discriminatory impact.

Open source
ADA.gov

Algorithms, AI, and Disability Discrimination in Hiring

Guidance on accessibility, accommodations, and the risk that hiring technology measures disability instead of job skills.

Open source
Metix AI

Agent Evaluation, Done Right

First-party engineering material on component, trajectory, and outcome evaluation for production agents.

Open first-party source

Apply one standard

Compare systems on the same role and evidence request.

Use the methodology and vendor checklist to record what each system can demonstrate, what failed, and which questions remain unresolved. A score helps organize the review; it does not replace the accountable decision.