What We Learned Automating Bias Audits for NYC Local Law 144
December 13th, 2024
What We Learned Automating Bias Audits for NYC Local Law 144
Since 2023, New York City's Local Law 144 has required employers to run independent bias audits on any automated employment decision tool (AEDT) used in hiring. At Eticas.ai, we built ITACA_144 — a version of our broader bias-auditing platform, ITACA_OS — to help employers meet this requirement. Along the way, we surfaced several structural gaps in the law that limit how effective it can be, and we believe those lessons are relevant well beyond New York.
Data requirements are too loose. The law doesn't specify what data an audit must use — it could be historical, out-of-region, or unrelated to NYC hiring altogether. We recommend future legislation require recent (12-month) data drawn specifically from NYC-relevant hiring processes, so audits are comparable and meaningful.
The 2% exclusion rule undermines representation. Auditors can exclude any demographic group under 2% of the dataset from impact-ratio calculations. In NYC, that threshold quietly excludes American Indian, Alaska Native, Native Hawaiian, Pacific Islander, and multiracial populations — often the groups most vulnerable to algorithmic bias. We recommend removing this threshold and tightening the definition of "Some Other Race."
Impact ratio alone isn't fairness. Local Law 144 requires only one metric — impact ratio — which measures selection-rate differences between groups. Proportional outcomes don't guarantee unbiased treatment: proxy features (like language correlating with ethnicity) can still drive indirect discrimination. Real assessment requires counterfactual analysis, not just outcome parity.
Model metrics miss "effective bias." A hiring decision is a long pipeline — screening, testing, interviews, manager sign-off — and the AI model is only one link. Measuring the model in isolation is blind to bias introduced before and after it. Our audits track outlier performance across pre-, in-, and post-processing stages to capture bias where it actually happens.
Metrics need enforcement. The law references the 80/20 rule but doesn't act when audits fall outside it. We supplement this with representativity benchmarks drawn from US Census data, so audits don't just measure — they set a standard for what "acceptable" looks like.
Self-reported audit data needs oversight. Audits currently rely on data the audited organization itself provides. We recommend regulators commit to independent, in-depth spot checks — verifying systems in operation rather than only reviewing submitted reports — to protect the credibility of the audit process itself.
Local Law 144 is an important precedent — one of the first real-world tests of mandatory AI bias auditing, and a likely template for future US state and local legislation. But its current limitations show why the details of how audits are required matter just as much as whether they're required. We share these findings in the hope they help other jurisdictions build stronger frameworks — and help ensure that AI accountability regulation actually protects the people it's meant to protect.
Read the full paper here: https://arxiv.org/html/2501.10371v1