arXiv AI By Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee, Kin Chung Ho, Ping Shum, Michael K. Ng

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails

Read the original on arXiv AI →

arXiv:2608. 08485v1 Announce Type: new Abstract: Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 16

HoloAegis: Frozen Representation, Topological Inference --- Minimally Parametric Safety Manifolds and Their Capability Boundaries for LLM Guardrails

HoloAegis is a minimally parametric topological inference framework that uses frozen representations to map text onto the unit sphere and makes decisions via Gibbs‑Boltzmann free‑energy differences over pre‑computed anchor centroids. On a frozen three‑benchmark protocol, it matches WildGuard‑7B on toxicity, outperforms it on harmful behaviors, but underperforms on oversafety detection, while ShieldGemma‑2B fails on indirect harms. The study demonstrates that geometric guardrails can substitute for LLM judges in some cases and must defer to them in others, with anchor banks reducing score variance and boundary displacement.

By Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee, Kin Chung Ho, Ping Shum, Michael K. Ng
arXiv Computation and Language
Aug 24

Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds

The paper shows that large language models (LLMs) naturally organize their hidden state manifolds into small‑world networks, enabling efficient multi‑hop reasoning. By converting similarity matrices into unweighted graphs, the authors trace connectivity between distant semantic anchors and find a sharp topological phase transition: deep reasoning layers compress conceptual distances into paths bounded by six semantic hops, while early syntactic layers remain fragmented. The framework is applied to zero‑shot hallucination detection in Retrieval‑Augmented Generation, revealing that factual generations preserve a ~3‑hop structure, whereas hallucinations collapse the topology.

By Md. Faiyaz Abdullah Sayeedi