Why Aggregate Metrics Miss Bias in AI Hiring

August 14th, 2026

An audit of an AI hiring system could report full compliance and still miss most of the bias in it. That's not a hypothetical; it's what happened when Eticas.ai ran two different audits on the same system and compared the results. The system is the one behind our recent case study, Evaluation of an Algorithmic Hiring System: A Public-Sector Case Study — five years of operational data from Barcelona Activa's shortlisting pipeline, which uses the third-party TalentClue platform for candidate search and filtering. That piece covers what we found, while this one is about how we found it, and why the method matters as much as the result.

The problem: a compliant-looking system that isn't
Run the simplest possible fairness check on this system — compare the share of women among registered candidates to the share among those eventually hired — and the numbers look fine. Women are 51.5% of registered candidates and 49.8% of those hired. The difference isn't statistically significant. A compliance audit scoped to that single comparison — the kind mandated under frameworks like NYC Local Law 144, for example — would sign off. That pattern isn't unique to Barcelona Activa. Recent research on Local Law 144's first year found that only 5% of covered employers publicly posted the required audit reports, and 96% of those that did reported passing impact ratios (Wright et al., 2024).

Almost everyone who evaluates their system passes. This suggests, like previous emerging regulations (cite), that these evaluations have become more of a rubber-stamping routine and a “tick-off” exercise.

Recruitment is often a process that involves a sequence, and it is not different in the particular case we evaluated: an employer submits a vacancy, a Barcelona Activa analyst searches TalentClue, TalentClue's proprietary algorithm filters and ranks, the analyst reviews and shortlists, the employer decides. Bias that appears at one stage of the sequence can be diluted, offset, or hidden by what happens at the next. Aggregate audits measure the beginning and the end. They don't see the pipeline or how bias propagates to amplify or hide implicit bias.

What we built: a granular methodology that follows the whole pipeline

Eticas.ai's evaluation methodology treats the AI system as what it actually is: a sociotechnical pipeline, not a single technical component. For this audit, that meant mapping and testing all seven stages Barcelona Activa operates — from vacancy receipt through analyst keyword selection, TalentClue's platform filtering, and shortlist handoff:

  • We analyze incoming applications and how various decisions/design choices (e.g., access to platform, data guardrails etc) can introduce bias into the pipeline before an application is even screened. In this particular case, is the population entering the pipeline representative of Barcelona's labor market, benchmarked against external labor force survey data?

  • We evaluate the whole pipeline to map what the risks are and at what stage they may enter. At each stage, is any group being filtered out at a different rate, tested with Disparate Impact Ratios, intersectional analysis, and statistical significance testing rather than a single aggregate comparison? This involves diving deeper into the data at a much more granular level to surface insights that are typically hidden in aggregate metrics.

  • We evaluate overtime: Are these patterns stable? Is bias amplified or not with time?

The same approach underpinning the Eticas AI Risk Taxonomy applies here: a risk isn't real until it's been broken into a testable mechanism, measured, and graded.

Seeing it work: five disparities the aggregate check missed

Stratifying the same five years of data by pipeline stage, salary band, sector, age, and gender together rather than collapsing it into one registered-to-hired ratio, surfaced disparities the aggregate comparison couldn't have caught:

  1. Women are less likely to be hired in the mid-salary range, and this was invisible in the overall hiring number.

  2. There are shortlist imbalances within individual sectors: women’s representation falls to almost half in Real Estate and Architecture, for example.

  3. High-salary and full-time jobs are mostly going to men, even within the same sector.

  4. There is a compounded age-and-gender effect for women in their late career, distinct from the gender or age effects on their own.

  5. Workers aged 55 and over are completely absent from every stage of the pipeline while they represent 15.6% of the city’s active labor force.

None of these five findings would appear in a registered-to-hired comparison. All five came from testing the pipeline stage by stage instead of the outcome alone. (The full set of results, including the non-binary shortlisting findings and country-of-origin disparities, is in the companion case study).

Why this matters beyond one agency

Four structural conditions produced this gap between what the aggregate showed and what the pipeline actually did, and none of them are specific to Barcelona Activa:

  • Vendor opacity. TalentClue's matching and ranking logic is proprietary. Neither Barcelona Activa nor Eticas.ai's audit team could see it directly; we could only see its effects and reconstruct from outcomes.

  • Undocumented human discretion. Search keywords, filter choices, and shortlist judgment calls happen without a record, which means responsibility for any given disparity can't be cleanly attributed to the algorithm or the analyst. There is no audit trail available to identify and mitigate these concerns.

  • Fairness as a dynamic concept. The gender gap in shortlisting narrowed over the five-year window — but so did shortlisting rates for everyone, which is at least as consistent with a tightening, more competitive pipeline as with a fairer one. A point-in-time audit run in 2017 and one run in 2022 would have told two different stories about the same system, without either one being wrong for its moment.

  • Lack of ongoing monitoring. The previous point is the argument for ongoing monitoring rather than periodic audits alone. Under the EU AI Act, Article 72 already requires post-market monitoring for high-risk AI systems, including those used in employment — the kind of ongoing oversight Eticas.ai's post-deployment monitoring service is built around. Of course, a monitoring plan built only around aggregate metrics is a monitoring plan that reproduces the same blind spots.

The takeaway for anyone deploying AI hiring tools

Passing a fairness check that only looks at the aggregate outcome of the system doesn't tell you it is fair — it tells you where you didn't look. Evaluating the full pipeline, with the right stratification and granularity, and over enough time to see whether patterns hold, is what turns "we checked" into evidence you can actually stand behind.

Read the full paper Applied and Filtered: An End-to-End Algorithmic Fairness Audit of Public Employment Agency, on arXiv: http://arxiv.org/abs/2608.13022

Previous
Previous

Eticas.ai Joins PACT AI as Founding Member to Scale Trust in the AI Economy 

Next
Next

Our CEO at the Launch of the UK's AI Assurance Stakeholder Consortium