Back to today's list

HoloAegis: Frozen Representation, Topological Inference --- Minimally Parametric Safety Manifolds and Their Capability Boundaries for LLM Guardrails

Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee, Kin Chung Ho, Ping Shum, Michael K. Ng

Published Sep 16, 2026Featured #1In the daily list Aug 12, 2026
Daily score78.6
Editorial review8.2
Relevance0.466
Freshness0.722

Why It Matters

What makes this one worth your time

This work addresses critical challenges in LLM safety by providing a method that maintains representation integrity while ensuring efficient decision-making, which is crucial for real-world applications.

HoloAegis offers a novel geometric approach to LLM safety without fine-tuning.

Summary

The paper introduces HoloAegis, a minimally parametric framework for LLM safety that utilizes geometric reasoning over frozen representations, avoiding the pitfalls of fine-tuning and high inference costs.

Key contributions

  • Introduction of a minimally parametric topological inference framework for LLM safety.
  • Formalization of safety evaluation through geometric reasoning and Gibbs-Boltzmann Free Energy.
  • Empirical validation across multiple benchmarks demonstrating state-of-the-art performance with minimal latency.

Notable insights

  • The use of a Gibbs-Boltzmann Free Energy computation for safety evaluation is a novel approach that leverages topological properties.
  • The Topological Boundary Stability Conjecture suggests a new way to stabilize decision boundaries against perturbations, which could influence future research in representation learning.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2608.08485v2 Announce Type: replace Abstract: Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs. We ask a complementary question: how far can safety be achieved through pure geometric reasoning over frozen representations, and where does it fail? We present HoloAegis, a minimally parametric topological inference framework that decouples representation from reasoning: an un-fine-tuned encoder maps text to the unit sphere S^{d-1}, and all decisions reduce to Gibbs-Boltzmann free-energy differences over pre-computed anchor centroids. We contribute a boundary-mapping study rather than a leaderboard claim. On a frozen three-benchmark protocol, HoloAegis (3.2 MB) statistically matches WildGuard-7B (14 GB) on toxicity (0.96 vs. 0.96), exceeds it on harmful behaviors (0.99 vs. 0.79), and cedes oversafety detection (0.62 vs. 0.98) -- while ShieldGemma-2B fails on indirect harms (0.34). These failure modes are complementary and mechanistically traceable: potential-difference scoring senses manifold clustering, whereas policy-conditioned LLM judging requires explicit taxonomy matching. We restate our Topological Boundary Stability conjecture in ratio form and validate it via reference-set bootstrap: anchor banks reduce score variance 4-15x and boundary displacement to approximately 0.44 + 0.23 sqrt(k/K) of the full-space estimator. Per-domain analysis further reveals that geometric separability tracks within-domain semantic homogeneity. Our results chart where geometric guardrails substitute for, and where they must defer to, LLM judges.