Why Aggregate Metrics Miss Bias in AI Hiring

August 14th, 2026

An audit of an AI hiring system could report full compliance and still be missing most of the bias in it. That's not a hypothetical; it's what happened when Eticas.ai ran two different audits on the same system and compared the results. The system is the one behind our recent case study, Evaluation of an Algorithmic Hiring System: A Public-Sector Case Study — five years of operational data from Barcelona Activa's AI-assisted shortlisting pipeline, which uses the third-party TalentClue platform for candidate search and filtering. That piece covers what we found. This one is about how we found it, and why the method matters as much as the result.

The problem: a compliant-looking system that isn't
Run the simplest possible fairness check on this system — compare the share of women among registered candidates to the share among those eventually hired — and the numbers look fine. Women are 51.5% of registered candidates and 49.8% of those hired. The difference isn't statistically significant. A compliance audit scoped to that single comparison — the kind mandated under frameworks like NYC Local Law 144, for example — would sign off. That pattern isn't unique to Barcelona Activa. Recent research on Local Law 144's first year found that only 5% of covered employers publicly posted the required audit reports, and 96% of those that did reported passing impact ratios (Wright et al., 2024).

Almost everyone who checks, passes. That should be a signal about what the checking is measuring, instead of a clean bill of health for the industry.

The reason a single aggregate number can look this reassuring is structural, not accidental. Recruitment isn't one decision, but a sequence: an employer submits a vacancy, a Barcelona Activa analyst searches TalentClue, TalentClue's proprietary algorithm filters and ranks, the analyst reviews and shortlists, the employer decides. Bias that appears at one stage can be diluted, offset, or hidden by what happens at the next. Aggregate audits measure the beginning and the end. They don't see the pipeline in between.

What we built: a methodology that follows the whole pipeline

Eticas.ai's evaluation methodology treats the AI system as what it actually is: a sociotechnical pipeline, not a single technical component. For this audit, that meant mapping and testing all seven stages Barcelona Activa operates — from vacancy receipt through analyst keyword selection, TalentClue's platform filtering, and shortlist handoff — across three analytical layers:

  • Pre-processing — is the population entering the pipeline representative of Barcelona's labor market, benchmarked against external labor force survey data?

  • In-processing — at each stage, is any group being filtered out at a different rate, tested with Disparate Impact Ratios, intersectional analysis, and statistical significance testing rather than a single aggregate comparison?

  • Post-processing — are these patterns stable, or shifting over the five years of data available?

The approach is the same that underpins the Eticas AI Risk Taxonomy: a risk isn't real until it's been broken into a testable mechanism, measured, and graded.

Seeing it work: five disparities the aggregate check missed

Stratifying the same five years of data by pipeline stage, salary band, sector, age, and gender together rather than collapsing it into one registered-to-hired ratio, surfaced disparities the aggregate comparison couldn't have caught:

  1. Adverse impact for women specifically in mid-salary shortlisting, invisible in the overall hiring number.

  2. Salary-matching gaps that persist within individual sectors, not just across them — ruling out sector mix as the explanation.

  3. A shortlisting penalty concentrated in full-time roles specifically, where job stability and pay are highest.

  4. A compounded age-and-gender effect for women in their late career, distinct from the gender or age effects on their own.

  5. The complete, structural absence of workers aged 55 and over from every stage of the pipeline — not underrepresentation, but zero. None of these five findings would appear in a registered-to-hired comparison. All five came from testing the pipeline stage by stage instead of the outcome alone. (The full set of results, including the non-binary shortlisting findings and country-of-origin disparities, is in the companion case study.)

Why this matters beyond one agency

Three structural conditions produced this gap between what the aggregate showed and what the pipeline actually did, and none of them are specific to Barcelona Activa:

  • Vendor opacity. TalentClue's matching and ranking logic is proprietary. Neither Barcelona Activa nor Eticas.ai's audit team could see it directly — only its effects, reconstructed from outcomes.

  • Undocumented human discretion. Search keywords, filter choices, and shortlist judgment calls happen without a record, which means responsibility for any given disparity can't be cleanly attributed to the algorithm or the analyst.

  • Fairness that moves. The gender gap in shortlisting narrowed over the five-year window — but so did shortlisting rates for everyone, which is at least as consistent with a tightening, more competitive pipeline as with a fairer one. A point-in-time audit run in 2017 and one run in 2022 would have told two different stories about the same system, without either one being wrong for its moment.

That last point is the argument for continuous evaluation rather than periodic audits alone. Under the EU AI Act, Article 72 already requires post-market monitoring for high-risk AI systems, including those used in employment — the kind of ongoing oversight Eticas.ai's post-deployment monitoring service is built around. A monitoring plan built only around the metrics an aggregate audit would check is a monitoring plan that reproduces the same blind spot, on a recurring schedule.

The takeaway for anyone deploying AI hiring tools

Passing a fairness check that only looks at the aggregate outcome doesn't tell you the system is fair — it tells you where you didn't look. Evaluating the full pipeline, with the right stratification and over enough time to see whether patterns hold, is what turns "we checked" into evidence you can actually stand behind.

Read the full paper here: http://arxiv.org/abs/2608.13022

Next
Next

Our CEO at the Launch of the UK's AI Assurance Stakeholder Consortium