HoloAegis is a minimally parametric topological inference framework that uses frozen representations to map text onto the unit sphere and makes decisions via Gibbs‑Boltzmann free‑energy differences over pre‑computed anchor centroids. On a frozen three‑benchmark protocol, it matches WildGuard‑7B on toxicity, outperforms it on harmful behaviors, but underperforms on oversafety detection, while ShieldGemma‑2B fails on indirect harms. The study demonstrates that geometric guardrails can substitute for LLM judges in some cases and must defer to them in others, with anchor banks reducing score variance and boundary displacement.
By Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee, Kin Chung Ho, Ping Shum, Michael K. Ng
arXiv:2605. 22873v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning has become the default strategy for enhancing LLM capabilities, yet its application raises a fundamental question: when is explicit reasoning actually beneficial?
By Wei Xia, Haoqing Wang, Zhi-Hong Deng, Yehui Tang
The paper shows that large language models (LLMs) naturally organize their hidden state manifolds into small‑world networks, enabling efficient multi‑hop reasoning. By converting similarity matrices into unweighted graphs, the authors trace connectivity between distant semantic anchors and find a sharp topological phase transition: deep reasoning layers compress conceptual distances into paths bounded by six semantic hops, while early syntactic layers remain fragmented. The framework is applied to zero‑shot hallucination detection in Retrieval‑Augmented Generation, revealing that factual generations preserve a ~3‑hop structure, whereas hallucinations collapse the topology.
By Md. Faiyaz Abdullah Sayeedi
arXiv:2607. 03329v1 Announce Type: new Abstract: Conventional uniform convergence bounds and empirical risk minimization break down in massive over-parameterized models, such as large language transformers and biological sequence networks.
By Bing Cheng, Yi-Shuai Niu, Howell Tong, Shing-Tung Yau
arXiv:2603. 10384v3 Announce Type: replace Abstract: Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics of reasoning.
By Xinyan Jiang, Ninghao Liu, Di Wang, Lijie Hu
arXiv:2607. 17962v1 Announce Type: cross Abstract: TabPFN is a transformer-based foundation model for tabular prediction that performs inference without task-specific training by conditioning on a support set and query inputs.
By James Hu, Mahdi Ghelichi
arXiv:2606. 28589v1 Announce Type: new Abstract: Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and "Wait" prompts, primarily encourage models to think more, yet often fail to guide them toward Truth.
By Tianlong Wang, Yuhang Wang, Weibin Liao, Xin Gao, Xinyu Ma, Yang Lin, Yasha Wang, Liantao Ma
Multimodal large language models (MLLMs) have advanced image geolocalization mainly by improving how they reason about geographic cues. How that reasoning isdecoded into coordinates, however, has lagged behind.
MindTopo is a benchmark that tests foundation models on topological reasoning, covering five cognitive properties—continuity, separation, order, enclosure, and knots—across two cognitive levels: reasoning and planning. It contains 11,030 instances from 13 procedurally generated task types, and evaluates 14 multimodal large language models, including agent configurations with image and video generation. Results show that models perform better on reasoning than planning, and even the best model lags far behind human performance, with fine‑tuning and reinforcement learning improving reasoning more than planning.
By Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica, Jianwen Lyu, Zihan Wang, Reuben Tan, Jianfeng Gao, Ruohan Zhang, Yining Hong, Jiajun Wu, Manling Li
arXiv:2603. 04852v2 Announce Type: replace Abstract: Multi-step theorem prediction is a central challenge in geometry problem solving.
By Junbo Zhao, Ting Zhang, Can Li, Wei He, Jingdong Wang, Hua Huang
arXiv:2607. 01571v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve complex problems by generating intermediate reasoning steps.
By Aria Masoomi, Mahsa Bazzaz, Adel Javanmard, Vahab Mirrokni
arXiv:2608. 07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures.
By Shibo Chu, Yuze Liu, Tiehua Zhang, Zhishu Shen, Lianghua He, Haofen Wang, Zhijun Ding