Someone at iTutorGroup configured the company's application software to reject women aged 55 and over and men aged 60 and over the moment their birthdate hit the form. The EEOC sued, and in September 2023 the company paid $365,000 to settle. No neural network, no training data, nothing anyone would call artificial intelligence, and that is exactly why it remains the clearest case in the field: the discrimination was legible. You could read it off a config screen and point at it in court. Nothing since has been that easy to read, which is the actual problem with handing hiring to models.

The story most people carry around is the Amazon one. Train a screener on a decade of your own hiring, find your own historical preferences reflected back, quietly kill the project. It fits the intuition that machines launder the past, and for a while the research agreed. Brookings researchers simulated resume screening with large language models and found significant gender and racial discrimination, concentrated on Black men: in their setup, resumes with Black male names were preferred over white male names zero percent of the time. Five leading models tested through VoxDev showed the same intersectional pattern, where Black women, Black men and white women each get sorted differently, so a fairness test built on race alone or gender alone reports nothing.

Then the sign changed. A paired-resume audit of fourteen models, modelled on the classic correspondence experiments and running 24,024 matched job postings per model, found that GPT-3.5-turbo reproduced the human pro-white callback gap at 2.12 percentage points, and that every model released in 2024 or later showed either no gap or a statistically significant reversal favouring Black-coded names, by as much as three points. The same flip appeared on the gender axis. A separate study across twenty-nine models found that language models now advantage female and Black candidates relative to comparable white and male ones while disadvantaging disabled candidates, the effect of demographic identity worth somewhere between six months and a year of extra education. Post-training alignment, not the pretraining corpus, turned out to be the main driver.

I don't think that's good news. A screener that favours one group over another fails the same test whichever way it leans, and a bias that reverses between model generations is harder to govern than one that doesn't, because it can't be certified. It moves with vintage, with provider, with whatever the alignment team shipped last quarter. New York City's Local Law 144 requires an annual bias audit of automated employment tools. The EU AI Act classifies hiring systems as high-risk. Both regimes were designed around the Amazon story, where bias is durable, inherited, and points where you'd predict. An annual audit of something whose sign changes between releases is a photograph of a moving object, and the auditors are mostly reading figures the vendor produced, a weakness familiar from self-reported benchmarks.

The comparison that matters, though, isn't against a fair process. It's against human recruiters, and the human baseline is dreadful. It has been measured since the 2004 correspondence study in which Bertrand and Mullainathan sent fake CVs to Boston and Chicago employers and watched white-sounding names collect fifty percent more callbacks. In a field experiment covering seventy thousand applicants, AI-led interviews produced twelve percent more job offers, eighteen percent more people actually starting work, and better thirty-day retention. Those are outcomes, on real hires, and anyone defending the status quo has to account for them. The fairness number from the same study needs more care than it usually gets: reported gender-based discrimination fell from 5.98 percent to 3.30 percent, and "reported" is carrying the sentence. That is what candidates said about their experience, not a count of who got selected, and a machine interviewer can feel less prejudiced while sorting people exactly as badly. The hiring and retention gains I'd take at face value. The fairness gain I'd want measured a different way.

The catch is what happens when you put a person back in the loop, which is what every compliance framework asks for. In a screening experiment with 528 participants, people deciding alone, or alongside an unbiased model, selected candidates from all racial groups at roughly equal rates. Point them at a biased model and their choices tracked its preferences up to ninety percent of the time. The humans were fine until the recommendation showed up. Human oversight, the phrase carrying most of the load in AI hiring regulation, describes a person whose judgement the system has already colonised.

Against that, a study on Denmark's largest job portal compared recruiters searching manually, algorithmic matching, and the two together, and found the hybrid produced the fairest candidate lists of the three, better than either alone. Oversight worked there. The difference I'd bet on is accountability: the Danish recruiters were professionals on their own platform filling real vacancies with their reputations riding on the shortlist, while the experiment's participants were completing a task for a stranger and had no stake in being right. If that's the mechanism, then "human oversight" in a statute means nothing unless the human carries consequences, and none of the current rules require that. They require a person to be present.

A separate finding sits outside the demographic argument entirely. Maryland researchers ran twenty-two hundred resumes through commercial and open-source models and found self-preference rates of 67 to 82 percent: the models rank resumes generated by themselves above human-written ones of equivalent quality. Whatever that is measuring, it isn't the candidate. It rewards knowing which vendor screens your application, which is knowledge distributed exactly as unevenly as you'd expect.

Applicants adjust too, and that quietly undermines everything above. An IZA experiment found application rates falling 4.6 points for women and 3.2 points for men as AI involvement in the evaluation increased. If women withdraw from AI-screened roles at a higher rate than men, every callback-parity statistic in every study I've cited is computed on a pool the screener already reshaped before it read a single CV. A tool can post clean demographic parity across the applications it receives and still have skewed the workforce, because the skew happened upstream, in who bothered to apply. No audit regime I'm aware of looks there.

None of this makes the tools indefensible, and the scale argument does more work for me than any individual study. A bad human recruiter damages a few hundred careers across a working life, and eventually somebody notices the pattern. A model licensed across the Fortune 500 applies one idiosyncratic preference to millions of applications at once, and the only people positioned to notice are the ones running it. Mobley v. Workday is grinding through the Northern District of California on an age-discrimination theory, and it matters more than its facts because of the position it puts everyone in. The applicant can't see the screen. The employer that licensed the tool often can't either. The vendor owes an explanation to neither. A judgment against a vendor is the only lever anyone has found that reaches inside the model, and it's a slow, expensive, ten-year lever, aimed at an industry that is simultaneously removing the entry-level rungs those applications were pointed at.

Sources: