Someone at iTutorGroup configured the company's application software to
reject women aged 55 and over and men aged 60 and over the moment their
birthdate hit the form. The EEOC sued, and in September 2023 the company
paid $365,000
to settle. No neural network, no training data, nothing anyone would call
artificial intelligence, and that is exactly why it remains the clearest case
in the field: the discrimination was legible. You could read it off a config
screen and point at it in court. Nothing since has been that easy to read,
which is the actual problem with handing hiring to models.
The story most people carry around is the Amazon one. Train a screener on a
decade of your own hiring, find your own historical preferences reflected
back, quietly kill the project. It fits the intuition that machines launder
the past, and for a while the research agreed. Brookings researchers
simulated resume screening with large language models and found significant
gender and racial discrimination,
concentrated on Black men: in their setup, resumes with Black male names were
preferred over white male names zero percent of the time. Five leading models
tested through VoxDev showed the same
intersectional pattern,
where Black women, Black men and white women each get sorted differently, so
a fairness test built on race alone or gender alone reports nothing.
Then the sign changed. A paired-resume audit of fourteen models, modelled on
the classic correspondence experiments and running 24,024 matched job
postings per model, found that GPT-3.5-turbo reproduced the human pro-white
callback gap at 2.12 percentage points, and that
every model released in 2024 or later
showed either no gap or a statistically significant reversal favouring
Black-coded names, by as much as three points. The same flip appeared on the
gender axis. A separate study across twenty-nine models found that language
models now advantage female and Black candidates relative to comparable white
and male ones while
disadvantaging disabled candidates, the
effect of demographic identity worth somewhere between six months and a year
of extra education. Post-training alignment, not the pretraining corpus,
turned out to be the main driver.
I don't think that's good news. A screener that favours one group over
another fails the same test whichever way it leans, and a bias that reverses
between model generations is harder to govern than one that doesn't, because
it can't be certified. It moves with vintage, with provider, with whatever
the alignment team shipped last quarter. New York City's Local Law 144
requires an annual bias audit of automated employment tools. The EU AI Act
classifies hiring systems as high-risk. Both regimes were designed around the
Amazon story, where bias is durable, inherited, and points where you'd
predict. An annual audit of something whose sign changes between releases is
a photograph of a moving object, and the auditors are mostly reading figures
the vendor produced, a weakness familiar from
self-reported benchmarks.
The comparison that matters, though, isn't against a fair process. It's
against human recruiters, and the human baseline is dreadful. It has been
measured since the 2004 correspondence study in which Bertrand and
Mullainathan sent fake CVs to Boston and Chicago employers and watched
white-sounding names collect fifty percent more callbacks. In a field
experiment covering seventy thousand applicants,
AI-led interviews produced
twelve percent more job offers,
eighteen percent more people actually starting work, and better thirty-day
retention. Those are outcomes, on real hires, and anyone defending the status
quo has to account for them. The fairness number from the same study needs
more care than it usually gets: reported gender-based discrimination fell
from 5.98 percent to 3.30 percent, and "reported" is carrying the sentence.
That is what candidates said about their experience, not a count of who got
selected, and a machine interviewer can feel less prejudiced while sorting
people exactly as badly. The hiring and retention gains I'd take at face
value. The fairness gain I'd want measured a different way.
The catch is what happens when you put a person back in the loop, which is
what every compliance framework asks for. In a screening experiment with 528
participants, people deciding alone, or alongside an unbiased model, selected
candidates from all racial groups at roughly equal rates. Point them at a
biased model and their choices
tracked its preferences up to ninety percent of the time.
The humans were fine until the recommendation showed up. Human oversight, the
phrase carrying most of the load in AI hiring regulation, describes a person
whose judgement the system has already colonised.
Against that, a study on Denmark's largest job portal compared recruiters
searching manually, algorithmic matching, and the two together, and found the
hybrid produced the
fairest candidate lists of the three,
better than either alone. Oversight worked there. The difference I'd bet on
is accountability: the Danish recruiters were professionals on their own
platform filling real vacancies with their reputations riding on the
shortlist, while the experiment's participants were completing a task for a
stranger and had no stake in being right. If that's the mechanism, then
"human oversight" in a statute means nothing unless the human carries
consequences, and none of the current rules require that. They require a
person to be present.
A separate finding sits outside the demographic argument entirely. Maryland
researchers ran twenty-two hundred resumes through commercial and open-source
models and found
self-preference rates of 67 to 82 percent:
the models rank resumes generated by themselves above human-written ones of
equivalent quality. Whatever that is measuring, it isn't the candidate. It
rewards knowing which vendor screens your application, which is knowledge
distributed exactly as unevenly as you'd expect.
Applicants adjust too, and that quietly undermines everything above. An IZA
experiment found application rates falling 4.6 points for women and 3.2
points for men as AI involvement in the evaluation increased. If women
withdraw from AI-screened roles at a higher rate than men, every
callback-parity statistic in every study I've cited is computed on a pool the
screener already reshaped before it read a single CV. A tool can post clean
demographic parity across the applications it receives and still have skewed
the workforce, because the skew happened upstream, in who bothered to apply.
No audit regime I'm aware of looks there.
None of this makes the tools indefensible, and the scale argument does more
work for me than any individual study. A bad human recruiter damages a few
hundred careers across a working life, and eventually somebody notices the
pattern. A model licensed across the Fortune 500 applies one idiosyncratic
preference to millions of applications at once, and the only people
positioned to notice are the ones running it. Mobley v. Workday is grinding
through the Northern District of California on
an age-discrimination theory, and it matters more than its facts because of
the position it puts everyone in. The applicant can't see the screen. The
employer that licensed the tool often can't either. The vendor owes an
explanation to neither. A judgment against a vendor is the only lever anyone
has found that reaches inside the model, and it's a slow, expensive,
ten-year lever, aimed at an industry that is simultaneously
removing the entry-level rungs
those applications were pointed at.
Sources:
-
iTutorGroup to Pay $365,000 to Settle EEOC Discriminatory Hiring Suit — U.S. Equal Employment Opportunity Commission
-
Gender, race, and intersectional bias in AI resume screening via language model retrieval — Brookings
-
AI hiring tools exhibit complex gender and racial biases — VoxDev
-
Can LLMs Hire Fairly? Racial Bias in Resume Screening — arXiv
-
AI Alignment Amplifies the Role of Race, Gender, and Disability in Hiring Decisions — arXiv
-
Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — SSRN
-
No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy — AAAI/ACM Conference on AI, Ethics, and Society
-
Fairness of human, AI, and hybrid recruiting on a real-world platform — arXiv
-
AI Hiring Tools May Favor Their Own Work, Smith Study Finds — University of Maryland Robert H. Smith School of Business
-
Human–AI Evaluation and Gender Transparency: Application Decisions in Competitive Hiring — IZA Institute of Labor Economics
-
Bias in the Machine: How AI Hiring Tools Create Risk for Employers — Ice Miller via JD Supra