AI audits have become a cornerstone of regulatory approaches worldwide, yet merely measuring fairness misses a critical dimension: whether the AI system's goals themselves reflect equitable and just choices. Experts call for integrating power tests in audit regimes to evaluate the social legitimacy of targeted outcomes, beyond just technical accuracy or bias metrics.
- AI audits currently emphasize fairness scores but often ignore system goals.
- Power tests assess legitimacy of AI objectives within social and institutional contexts.
- Existing frameworks partially address these concerns but lack uniform requirements.
What happened
A 2019 research case exposed that a US healthcare algorithm designed to predict future costs rather than medical necessity systemically disadvantaged Black patients. The model’s objective—cost prediction—reflected existing inequalities in healthcare spending, resulting in sicker Black patients being underserved despite receiving similar risk scores to white patients. This study exemplifies how AI systems can be technically accurate yet substantively flawed when their goals embed social disparities.
Governments and standards bodies worldwide have increasingly mandated AI audits to evaluate bias and risk. For example, New York City requires bias audits for automated hiring tools, the EU AI Act mandates conformity and fundamental rights assessments for high-risk AI, and NIST offers a voluntary framework for AI risk management. However, these frameworks primarily focus on technical performance and risk without consistently scrutinizing whether the system’s fundamental objectives are socially acceptable and justified.
Why it matters
Audits that certify AI fairness by measuring error disparities alone risk legitimizing policies that embed systemic discrimination through their objectives. Without evaluating who chooses the AI's goals and whether those goals are appropriate, there is a danger of 'bias laundering'—where political or institutional decisions about what to predict become perceived as neutral technical outcomes. This undermines accountability and can entrench social inequalities under algorithmic authority.
Some frameworks begin to address these gaps by recommending documentation of intended purposes, stakeholder consultations, and contextual impact assessments. Yet, participation from affected communities is often optional, and enforcement remains uneven. The EU’s delays in enforcing fundamental rights impact assessments further highlight a governance deficit. Bridging this divide is essential to move beyond compliance checklists toward meaningful justice and fairness in AI deployment.
What to watch next
Stakeholders should monitor the evolution of international AI governance, particularly how audit regimes integrate power tests that assess the legitimacy of system objectives and ensure affected groups have meaningful participation. Improvements to the EU AI Act’s implementation and NIST’s framework could set benchmarks for requiring transparent decision-making around AI system goals.
Advocates and policymakers may also push for unified standards that mandate explicit evaluation of AI objective selection, accountability for political and institutional influence, and clear mechanisms to challenge inequitable outcomes. These developments will be crucial for preventing AI from reinforcing entrenched social harms despite technical fairness compliance.