Your AI system, evaluated by
independent experts
We independently assess how AI systems behave in the real world, giving you credible assurance that your AI is safe and dependable in operation.
Who this is for
Developers
You build AI systems.
Your clients trust that they work responsibly.
We give you the evidence to back that up.
Deployers
You deploy AI built by someone else.
You are still accountable for what it does in your operations and how it impacts your users, regardless of who built it.
AI System Risk Management Process
Step 0: Forward Deployment
Get our experts to advise your team at earlier stages in the design or development process, so you can minimize risks from the outset
Step 1: First Evaluation
Get a clear picture of the risks your system is exposed to in a language board members or investors can understand, and your team can act on risk mitigation.
Step 2: Automated Monitoring
After the evaluation, we track agreed metrics at defined intervals, depending on the risk profile of your system. We alert you of drifts and interpret what they mean.
Step 3: New Evaluations
An evaluation - or audit - is only valid for a specific period, as AI systems by nature evolve with time. We regularly provide a new full evaluation guided by our experts.
Step 0 (optional): Forward Deployment
Get advice during pre-production stages, so you can minimize risks from the outset
Our experts have extensive experience advising product and technical teams during the design and development stages. Our team can surface risks that are far cheaper to address at the design stage than after launch. Drawing on a socio-technical approach that pairs technical evaluation with regulatory and contextual expertise, we assess governance structures, data quality, and system requirements early, so that fairness, explainability, and compliance gaps are identified before they become embedded in a live system. This pre-production advisory work — from readiness diagnostics through vendor and procurement vetting — gives teams a clear risk baseline from the outset, reducing the cost and disruption of retrofitting fixes later and setting a stronger foundation for responsible deployment and future evaluations.
Get a clear picture of the risks your system is exposed to in a language board members or investors can understand, and your team can act on risk mitigation.
Step 1: First Evauation
Our experts listen to your concerns. Whether you are worried about bias or hallucination, ROI, liability, or curious about environmental impact, or just want insight to make sure your AI systems are performing as expected. Then, with our expertise on where the biggest risks lie for your particular system and your additional requirements, we build the risk matrix and a data schema.
With your input, we assign metrics, benchmarks and AI judges to scan your production data, access your AI systems and quantify risks, diagnose risk sources, and suggest mitigation strategies.
Most importantly, you get an Evaluation Brief with a standardized Score, which you can share with users, clients, investors, regulators and partners. Your Eticas.ai Badge publicly shows your commitment to AI accountability.
Risk picture in board language: What the system is doing, where it could fail and how serious the exposure is; translated from technical language
Risk dynamics dashboard: Visualization of risk levels across the metrics that matter for your specific system
Impact insight: How the system is affecting your business, your staff and your learners. Not just whether it works, but who it works for
Mitigation pathways: Prioritized recommendations for each risk surfaced, with clear actions the team can implement
Monitoring-ready benchmarks: Metrics and thresholds agreed during evaluation, ready to carry forward into continuous monitoring
The Evaluation Score and Badge: A standardized scorecard and badge shareable with clients, investors, regulators, and partners as evidence of AI accountability
Steps 3 & 4: Monitoring & New Evaluation
An evaluation – or audit - is only valid for a specific period, as AI systems by nature evolve with time.
We automatically monitor agreed metrics at defined intervals, depending on your system’s risk profile. We alert you of drifts, changes in model or user behavior and interpret what they mean.
After a year or a material system change, we revisit risks, metrics and benchmarks, updating the data schema as appropriate, with a new full evaluation (also called an audit).
You get an up-to-date picture of how your AI system is performing in your specific context, control over impacts on revenue, operators, users and compliance, as well as defensible assurance and safety controls.
Over time, you make better decisions, with the right AI, at the right price - and it shows.
Dashboard including last-read values and evolution
Alerts when metrics drift beyond agreed thresholds
Remediation guidance when issues surface: not just a flag, but a path forward
Sandbox to test potential impact of system changes
Evaluation trail with every new evaluation, to present as proof of audit
Data flows
Model & AI components
Business rules & UX
Human oversight
Organisational context
Real-world outcomes
Why Eticas
Independent
With no interest in the outcome, we objectively quantify any risk.
Socio-technical
We evaluate the full system: data, model, business rules, software, people and process.
Powered by Tech
Proprietary methodology and tech built over more than a decade of auditing systems in production.
Guided by Experts
Evaluating AI since 2012. We don’t stop at metrics. We give you the evidence and guidance to act.
FAQs
-
Forward Deployment involves Eticas experts advising your team during the early stages of design or development. This proactive approach helps identify and mitigate risks from the outset, before they're built into the system.
-
As early as possible in the design or development process. Early engagement allows for better risk management and ensures potential issues are addressed before they escalate.
-
The full AI system, not just the model.
That means data flows, preprocessing, business rules, user interface, human oversight, and real-world outcomes.
We scope every evaluation – also called an audit - around a structured set of risk dimensions, which we prioritize according to what matters most for your specific system. -
A typical evaluation takes between a few days and a couple of weeks, depending on scope and data access. That time is a feature, not a limitation: a genuine end-to-end audit of a production AI system requires understanding the specific system, the organization around it, and the context in which it operates.
Our methodology and tooling make the process significantly faster than a purely manual approach, but full automation is incompatible with what an evaluation needs to do. -
Both. We use a “Service as Software” model. We use structured methodology and software that we have been developing for over a decade, including open-source risk-assessment libraries, which makes us faster and more consistent than a purely manual approach. But every evaluation is adapted to the specific system, its data, and its organizational context. Technology is key for the execution of the audits, but the expertise of our human auditors is critical for guidance and certification.
-
We use our own Eticas Auditing Framework (risks, specific methodology, applicable metrics etc), taking into account the client’s needs and applicable jurisdictions. The client only provides specific inputs to help us scope the audit.
Some are listed below:
System context; what the system does, what decisions it makes, who it affects. This determines which parts of our framework apply.
Data access; what can be provided (labeled outcomes, logs, synthetic data) determines which metrics we can compute and depth of the audit/evaluation.
Regulatory context; which regime the system is subject to (NYC LL144, EU AI Act, etc.) determines which threshold set applies.
-
There's no fixed number, and it depends on a variety of factors, including regulations. For example, NYC LL 144 requires n ≥ 30 per subgroup-outcome cell minimum to report a point estimate at all. Secondly, N (the size of the sample) is itself a finding and therefore is not always prescribed.
-
While we always aim to go as deep as possible, we work with whatever access is available. Before we begin, we agree on the scope together: what data and documentation you can share, which system components we can test directly, and what we assess from outputs alone. The depth of the evaluation is proportional to the depth of access, and we are transparent about what each level of access allows us to certify.
-
AI systems evolve over time, so an initial evaluation is only valid for a specific period. Regular new evaluations reassess your system's risks, metrics and benchmarks as it becomes more complex.
-
A new evaluation gives you an up-to-date picture of how your AI system performs in its current context, helps you control impacts on revenue, operators, users, and compliance, and ensures defensible assurance and safety controls.
-
The evaluation is always the starting point and it needs to be performed regularly as AI systems naturally evolve with time. Evaluations always involve working with your team and they may involve changes in the prioritized risk or dimensions or the data schema, for example.
Monitoring is the continuation of any evaluation: we track the metrics and benchmarks identified and alert you when something changes.
Monitoring is not available as a standalone service: it depends on the evaluation to establish what to measure and what the thresholds should be.
-
Absolutely! Yes, and it goes further. GRC platforms help you record what processes you have in place and how they match the processes required by existing standards (like ISO42001 or SOC1/2) and regulations (like the EU Act). Independent evaluation tells you whether your AI system actually performs as those processes intend in production, with real data. The two are complementary, not competing.
-
Impact ratio and statistical parity, score distribution comparisons, error rate parity, counterfactual or perturbation testing and proxy analysis plus others, depending on system type (ADM vs. LLM based system), domain, and applicable regulation.
All findings are evaluated for statistical robustness including significance testing, repeated runs, and variance/confidence intervals, so we can distinguish a real effect from noise, especially on smaller subgroup samples where a disparity metric can look large just from sample size. Some examples are shared below:
For classification/ADM systems:
Disparate Impact ratio and Statistical parity difference;
Error-rate parity through calibration and Equalized-Ddds metrics when ground-truth labels exist;
Proxy analysis via da_inconsistency/da_informative, which flag features that act as stand-ins for a protected attribute.
For LLM systems:
Counterfactual/perturbation testing is our core method and involves paired queries that hold everything constant except a protected attribute (decision-swap, multi-attribute grid, and intersectional-grid probes), scored for allocation disparity, quality-of-service disparity, and intersectional disparity.
Score-distribution comparison is covered by a continuous-output metric (allocation_score_disparity) for systems that output a graded score rather than a categorical decision, rather than collapsing everything to a single threshold crossing.
Real clients. All verticals. Real impact.