Data Scientist - AI Evaluator/Auditor
📍 Global — Remote
Seniority level — Mid / Senior
Employment type — Full-time
AI systems are only as trustworthy as the evidence behind them. As an Evaluator (also called an Auditor) on the Eticas.ai team, you will be the one building that evidence — putting AI systems through rigorous, real-world testing so organizations can prove, not just claim, that their AI does what it was designed to do.
You will be responsible for executing end-to-end system evaluation projects: from defining the evaluation scope after clarifying hypotheses with clients, to executing quantitative and qualitative analyses, and finally delivering actionable findings that help our clients minimize risk while improving system performance.
This role is especially critical as Eticas.ai continues to grow its capabilities and its platform, empowering expert evaluators with technology to ensure AI systems are fair in their impact on the world and behave as designed — including the agentic systems where independent evaluation matters most, and is hardest to get right. You will help bring a growing and evolving product to market, working closely with internal teams that shape positioning, define use cases, and drive adoption with key clients.
Although your responsibilities are clearly defined, we are a small, extremely agile team, with plenty of opportunity to learn and contribute across many areas — both internally and externally — as you help shape Eticas.ai's methodologies and contribute to setting global standards for algorithm accountability.
How we work
We're a fully remote team, spread across time zones and continents — but we work very closely together. We stay aligned through regular team syncs and project check-ins, and we default to clear, proactive communication and organized documentation, so distance never becomes disconnection.
We give people real flexibility in how and when they work. In return, we expect real ownership: you're accountable for outcomes, not hours online. It's a team built on trust, where independence and collaboration reinforce each other.
Key responsibilities
Lead AI system evaluations from scoping to delivery: define evaluation goals, data needs, methods, and success criteria.
Design and implement methodologies to measure performance, bias, fairness, robustness, and explainability in AI systems, as defined in the Eticas.ai Risk Taxonomy.
Build evaluation pipelines (Python) for data ingestion, preprocessing, metrics, statistical testing, and visualization — ensuring reproducibility and proper documentation.
Requirements
Fluency in English (written & spoken).
Master's degree (or equivalent) in Data Science, Computer Science, Statistics, Mathematics, or a related field.
Experience in AI/ML model evaluation.
Proficiency in Python (pandas, numpy, scikit-learn; familiarity with PyTorch/TensorFlow a plus).
Statistical knowledge: hypothesis testing, confidence intervals, sampling, calibration, and drift detection.
Ethics-driven mindset: strong understanding of the societal impact of AI and commitment to fairness and accountability.
Who we are
We are a mission-driven company founded in 2012 as one of the first firms dedicated to AI auditing. We independently evaluate and monitor AI systems. Our socio-technical team — empowered by the Eticas.ai platform — evaluates AI systems end-to-end, in the production environments where they actually run, turning rigorous evidence into practical tools for responsible AI development and deployment.
Eticas.ai has completed more than 200 evaluations across 15+ industries, with clients across four continents, covering expert and predictive systems, LLM-based systems, and agentic systems. Our work follows a repeating cycle of independent evaluation, continuous monitoring, and periodic re-evaluation — giving organizations verifiable evidence that their AI is doing what it was designed to do, while remaining fair, safe, and accountable.
Eticas.ai is the sister organization of Eticas Foundation, a non-profit whose mission is to build public-interest auditing capacity through community-led audits and shared evidence on AI harms.
Learn More
Our public evaluation library on GitHub: github.com/eticasai/eticas-audit
Our latest research: "The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits" — arxiv.org/abs/2607.02201
How to apply
Send us an email to: careers@eticas.ai