Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
Sam Relins, Daniel Birks
Why It Matters
What makes this one worth your time
Understanding the limitations of current AI governance frameworks is crucial for AI engineers and researchers to ensure safe and responsible deployment of general-purpose AI in public services.
The paper critiques current AI governance frameworks for failing to address the unique challenges posed by general-purpose AI.
Summary
The paper argues that existing public service AI governance frameworks are inadequate for general-purpose AI (GPAI) due to its generality, accessibility, and low deployment cost, which undermine traditional AI safety concepts. It uses policing as a case study to illustrate potential governance failures and suggests a taxonomic distinction between narrow and general-purpose AI, a pause on GPAI deployment in policing, and the establishment of a national safety infrastructure.
Key contributions
- Critique of existing AI governance frameworks in the context of general-purpose AI.
- Case study analysis of policing to illustrate governance challenges.
- Recommendations for policy changes and infrastructure development for GPAI safety.
Notable insights
- General-purpose AI challenges traditional safety concepts like accuracy and bias due to its unbounded output nature.
- Human-in-the-loop oversight and expert evaluation may be ineffective for GPAI, necessitating new governance strategies.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.25648v1 Announce Type: cross Abstract: Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressure has intensified with general-purpose AI (GPAI): AI built on large language models that can be directed by prompt alone to perform an effectively unbounded range of tasks. We argue that the properties that make these models attractive - their generality, accessibility, and low deployment cost - undermine the conditions under which AI safety has historically been pursued. The safety concepts that public service governance frameworks foreground - accuracy, bias, explainability, and accountability - were made tractable by narrow, purpose-built AI, and the mitigations that guidance documents prescribe presuppose exactly what GPAI removes. Accuracy cannot be quantified over unbounded outputs. Bias cannot be disaggregated when outputs are free-text judgements rather than categorical predictions. Explainability gives way to the appearance of explanation, and accountability erodes as outputs are optimized to persuade. We develop this through the case of policing, where the consequences of governance failure are most severe, and show why the same failure is likely to recur across other public services. The two mitigations that dominate policing AI strategy - expert evaluation and human-in-the-loop oversight - both rest on assumptions that GPAI violates. Safety assurance thus shifts from an intrinsic feature of building an AI tool to an optional add-on. We recommend a clear taxonomic distinction between narrow and general-purpose AI in governance documentation, a preference for technological parsimony, a pause on operational deployment of GPAI in policing until adequate evidence exists, and a coordinated national safety infrastructure with the authority to generate that evidence and determine when responsible deployment is achievable.