GeoSteer introduces a geometry-aware, optimization-based approach to norm-preserving activation steering in large language models. By formulating steering as a Riemannian optimization problem, it updates activations through a sequence of small geodesic steps guided by a learned nonlinear objective, avoiding fixed steering directions. Experiments on TruthfulQA, RealToxicityPrompts, and UltraFeedback show that GeoSteer consistently outperforms existing activation steering baselines, offering smoother, more stable, and more consistent steering behavior.
By Xuan Cuong Ngo, Hao Vo, Ngan Le
The paper introduces a Nested Inductive Bias framework that uses a two‑stage diffeomorphic composition to pull back non‑Euclidean target geometries onto symmetric positive definite (SPD) manifolds. This approach allows the construction of curvature‑aligned Riemannian classifiers that respect both matrix constraints and the intrinsic relational geometry of data. Empirical results on kinematic, signal processing, and synthetic benchmarks show that class separability degrades when metric curvature does not match the data distribution, and the authors also propose the Rational Conformal Metric (RCM) for robust vectorized architectures.
By Tushar Das
arXiv:2603. 10718v3 Announce Type: replace Abstract: Flow Matching enables simulation-free training of generative models on Riemannian manifolds, yet sampling typically still relies on numerically integrating a probability-flow ODE.
By Zichen Zhong, Haoliang Sun, Yukun Zhao, Yongshun Gong, Yilong Yin
arXiv:2609.10305v1 Announce Type: new
Abstract: Language models under one million parameters matter for edge deployment, domain adaptation, and reproducible research, yet a two-layer LSTM or Transfor...
By Fang Li
arXiv:2607. 07047v1 Announce Type: cross Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety.
By Szczepan Konior, Alexandre Quemy, Przemys{\l}aw Klocek, Gr\'egoire Cattan, Bart{\l}omiej Sobieski
arXiv:2609.05575v1 Announce Type: new
Abstract: Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability,...
By Yiming Tang, Harshvardhan Saini, Samyak Jha, Huaming Chen, Xufeng Duan, Dianbo Liu
The paper introduces a Geometric-to-Semantic Spherical Transfer Learning framework for labeling cortical sulci on brain surfaces. It first pre‑trains a spherical encoder on ~30,000 unlabeled UK Biobank subjects using only curvature and depth, then injects sulcal fundi lines as a soft‑initialized Topological Prior Injector to bridge the geometric‑semantic gap. Experiments show the method surpasses fully supervised baselines, achieving a mean Dice score of 0.77 and delivering the largest gains on variable and tertiary sulci.
By Saeb Tounsi, Jo\"el Chavas, Pietro Gori, Vincent Frouin, Denis Rivi\`ere, Jean-Fran\c{c}ois Mangin
arXiv:2607. 03329v1 Announce Type: new Abstract: Conventional uniform convergence bounds and empirical risk minimization break down in massive over-parameterized models, such as large language transformers and biological sequence networks.
By Bing Cheng, Yi-Shuai Niu, Howell Tong, Shing-Tung Yau
arXiv:2606. 03270v1 Announce Type: cross Abstract: Foundation models have sparked a revolution via a pretraining-adaptation paradigm, with recent efforts extending this success to graphs.
By Li Sun, Zhenhao Huang, Yiding Wang, Qin Chen, Pietro Lio, Philip S. Yu
arXiv:2608. 06031v1 Announce Type: new Abstract: Dynamic graph prompting freezes a pre-trained temporal backbone and adapts it to label-scarce downstream tasks using lightweight prompts.
By Quanxin Wang, Xuanting Xie, Bingheng Li, Xingtong Yu, Shuo Wang, Ruiyi Fang, Zhao Kang
arXiv:2512. 12225v3 Announce Type: replace Abstract: Developing artificial agents that unify representation, memory, adaptation, and prediction remains a fundamental challenge in artificial intelligence.
By Laha Ale
arXiv:2606. 25347v1 Announce Type: new Abstract: Exemplar-free class-incremental learning (EFCIL) requires stable decision boundaries within a shifting feature space.
By Hongye Xu, Bartosz Krawczyk