HoloAegis: Frozen Representation, Topological Inference --- Minimally Parametric Safety Manifolds and Their Capability Boundaries for LLM Guardrails
Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee, Kin Chung Ho, Ping Shum, Michael K. Ng
Why It Matters
What makes this one worth your time
This work addresses critical challenges in LLM safety by providing a method that maintains representation integrity while ensuring efficient decision-making, which is crucial for real-world applications.
HoloAegis offers a novel geometric approach to LLM safety without fine-tuning.
Summary
The paper introduces HoloAegis, a minimally parametric framework for LLM safety that utilizes geometric reasoning over frozen representations, avoiding the pitfalls of fine-tuning and high inference costs.
Key contributions
- Introduction of a minimally parametric topological inference framework for LLM safety.
- Formalization of safety evaluation through geometric reasoning and Gibbs-Boltzmann Free Energy.
- Empirical validation across multiple benchmarks demonstrating state-of-the-art performance with minimal latency.
Notable insights
- The use of a Gibbs-Boltzmann Free Energy computation for safety evaluation is a novel approach that leverages topological properties.
- The Topological Boundary Stability Conjecture suggests a new way to stabilize decision boundaries against perturbations, which could influence future research in representation learning.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2608.08485v2 Announce Type: replace Abstract: Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs. We ask a complementary question: how far can safety be achieved through pure geometric reasoning over frozen representations, and where does it fail? We present HoloAegis, a minimally parametric topological inference framework that decouples representation from reasoning: an un-fine-tuned encoder maps text to the unit sphere S^{d-1}, and all decisions reduce to Gibbs-Boltzmann free-energy differences over pre-computed anchor centroids. We contribute a boundary-mapping study rather than a leaderboard claim. On a frozen three-benchmark protocol, HoloAegis (3.2 MB) statistically matches WildGuard-7B (14 GB) on toxicity (0.96 vs. 0.96), exceeds it on harmful behaviors (0.99 vs. 0.79), and cedes oversafety detection (0.62 vs. 0.98) -- while ShieldGemma-2B fails on indirect harms (0.34). These failure modes are complementary and mechanistically traceable: potential-difference scoring senses manifold clustering, whereas policy-conditioned LLM judging requires explicit taxonomy matching. We restate our Topological Boundary Stability conjecture in ratio form and validate it via reference-set bootstrap: anchor banks reduce score variance 4-15x and boundary displacement to approximately 0.44 + 0.23 sqrt(k/K) of the full-space estimator. Per-domain analysis further reveals that geometric separability tracks within-domain semantic homogeneity. Our results chart where geometric guardrails substitute for, and where they must defer to, LLM judges.