Eticas.ai has been evaluating AI systems in production since 2012. This is where we share what we’ve learned. You’ll find case studies from client work, practical guides on AI governance and evaluation, reports on emerging risks, interviews with practitioners, and a glossary of the concepts that matter most in the field. Whether you build AI, deploy it, or are accountable for its outcomes, this is the evidence base we work from.
Our CEO at the Launch of the UK's AI Assurance Stakeholder Consortium
Eticas.ai's Founding President, Dr. Gemma Galdon-Clavell, will join the launch of the UK's AI Assurance Stakeholder Consortium in London, convened by the Department for Science, Innovation and Technology (DSIT) and led by BCS, The Chartered Institute for IT.
Gemma speaks to ANFAIA community
Eticas.ai's Founder & CEO, Dr. Gemma Galdón-Clavell, recently sat down with a group of young students from the ANFAIA community for a conversation about her professional journey and the case for auditing artificial intelligence. The full session is in Spanish — for readers who don't speak the language, here's an English recap of the main themes she covered.
The Eticas AI Risk Taxonomy: Operationalizing AI Evaluations with Measurable Proof
Eticas.ai's open AI risk taxonomy turns accountability principles into measurable proof: a tested, graded method for operationalizing AI audits.Evaluation of an Algorithmic Hiring System: A Public-Sector Case Study
A public employment agency in Europe relies on a third-party algorithmic platform to shortlist candidates for job vacancies, a system classified as high-risk under the EU AI Act. Eticas.ai conducted an independent, post-deployment fairness and bias evaluation of that system across five years of operational data. The evaluation found systematic disparities in shortlisting outcomes by gender, age, education level, and national origin — including adverse impact for women in mid-salary roles and the near-total exclusion of candidates aged 55 and over. Eticas.ai delivered targeted recommendations to address the findings.
WIRED trusts Eticas’ expertise in UK investigation
WIRED trusted Eticas as an expert to support their rigorous investigative work, and the work has finally been made public! Avon and Somerset Police built at least 23 predictive models. Some scored close to half a million people. For years, almost no one outside the force knew they existed. A major WIRED and Liberty Investigates investigation — supported by the Bristol Cable and Lighthouse Reports — has documented what happened next. Eticas reviewed the performance data the force disclosed and found models operating at precision rates below 10 percent, metrics shifting in ways inconsistent with well-governed systems, and bias testing that measured average risk scores by ethnicity without testing for discriminatory outcomes. Two child exploitation models were scrapped. Their source code could not be found. What the Bristol case makes visible is the distance between deploying an AI system and understanding what it is actually doing. The data existed. The performance records existed. The gap was in who was looking at them, with what methodology, and with enough independence to say plainly what they showed. High-stakes AI does not fail loudly. It accumulates decisions. The people affected rarely know they were scored.
Frequencia CEO - Podcast (v.o. Spanish)
Eticas.ai isn't a pivot or a trend. In her own words: it's her legacy! Her way of contributing to a world where technology is built to actually work in it, not just in a lab. Thirteen years later, the market is catching up. Gemma tells the full story in a recent episode of Frecuencia CEO, the podcast produced by @Qonto.
IAPP AI Governance Global Europe
Gemma took the stage at the IAPP AI Governance Global Europe conference in Dublin as part of the AI System Audits and Assurance workshop.
CEO Interview by InnovEU – The EU Project Chronicles
Is your AI a financial risk? Most organizations don't know — because they've never actually measured it. Gemma joined InnovEU – The EU Project Chronicles with @Fernando AC Gaspar to talk about exactly that. The conversation covers algorithmic liability, why auditing AI is a hardcore engineering challenge (not a paper ethics exercise), and what it really takes for innovation to hold up in the environments where it's actually deployed. The episode is sharp, practical, and built for organizations that are moving beyond the hype and starting to ask the harder questions.
Equitable AI for Outcomes
The conversation the sector needs is happening: how do we move AI from a set of principles on paper to systems that actually perform safely and fairly in the environments where they are used? For nonprofits and social sector organisations, the stakes are high and the guardrails are often still weak.
CEO Inerview by our partner, Armilla.ai
We were lucky to catch up with Gemma Galdon Clavell recently and have shared our conversation below.
Auditing AI Enabled Career Advisory Platform
The independent audit by Eticas.ai of Career Scoops AI-enabled career advisory platform for students (ages 13+) was performed post-deployment and focused on the end users in real educational settings. The findings confirmed a safe, reliable and student-appropriate deployment of AI for career exploration, and pointed to a few recommendations for improvement.
Deep Learning for social services
The evaluation of Allegheny County’s homelessness risk tool examined performance and potential disparities across protected groups. The insights led to stronger monitoring, clearer procedures, and better guidance for teams using the system in practice.
Detecting bias in AI hiring systems
The FINDHR project introduced practical tools, guidelines, and auditing frameworks designed to reduce discrimination in AI-assisted hiring. These resources support more transparent, inclusive, and accountable recruitment practices across Europe.
Gemma Closes DLA Piper European Technology Summit
Our Founder and CEO delivered the closing keynote at the DLA Piper European Technology Summit, bridging into the announcement of DLA Piper's AI campaign. DLA Piper's testimonial: "Gemma was absolutely brilliant on the day. Her energy and delivery really captured the room, and she created a seamless bridge from her closing keynote into the announcement of our AI campaign. She's a dynamic and engaging speaker."
Gemma Galdon Clavell Speaks at UNESCO Expert Roundtable on AI Audits
Gemma Galdon Clavell speaks at the UNESCO Expert Roundtable II: Capacity Building for Competent Authorities on AI, in the session "Tools for Supervising Authorities: Understanding Algorithm Audits," addressing EU data protection and market surveillance authorities.
Our CEO at RightsCon 2025: Covering the Last Mile of AI Auditing
Gemma joined Shea Brown at RightsCon 2025 to present "Covering the Last Mile of Accountability and Oversight: AI Auditing," on the critical role of AI auditing in ensuring accountability and transparency.
Bridging Socio-Technical Gaps in Bias Detection
Artificial intelligence (AI) models are increasingly autonomous in decision making, making the pursuit of responsible AI more critical than ever. Responsible AI (RAI) is defined by its commitment to transparency, privacy, safety, inclusiveness, and fairness. But while the principles of RAI are transparent and shared, RAI practices and auditing mechanisms are still incipient. A key challenge is establishing metrics and benchmarks that define performance goals aligned with RAI principles. This paper presents how the ITACA AI auditing platform incorporates demographic benchmarking for AI recommender systems to identify and measure bias. We propose a Demographic Benchmarking Framework to measure populations potentially affected by specific models, set acceptable performance ranges, and guide policymakers and developers. Our approach integrates socio-demographic insights directly into AI systems, reducing bias while also improving overall performance. The main contributions of this study include: 1. Defining control datasets tailored to specific demographics so they can be used in model training to quantify sampling bias; 2. Comparing the overall population with those impacted by the deployed model to identify discrepancies and account for structural bias; and 3. Quantifying drift in different scenarios continuously and as a post-market monitoring of deployment bias.