arXiv Machine Learning

Mutual Information Constrained Chernoff Bottleneck

arXiv Machine Learning
Aug 4

Tail-Aware Information-Theoretic Bounds for LLM Alignment under Heavy-Tailed Rewards

arXiv:2604. 10727v2 Announce Type: replace-cross Abstract: Classical information-theoretic learning bounds typically rely on KL mutual information and moment-generating-function (MGF) arguments, which are well matched to bounded or sub-Gaussian losses but can be ineffective when losses or rewards are heavy-tailed.

By Huiming Zhang, Binghan Li, Wan Tian, Qiang Sun
arXiv Machine Learning
Aug 11

Optimal Learning Under Tsybakov Noise

arXiv:2608. 08416v1 Announce Type: new Abstract: Probably Approximately Correct (PAC) learning [Val84] is a fundamental learning model that has been extensively investigated.

By Steve Hanneke, Hongao Wang, Mingyue Xu
arXiv AI
Sep 12

A Mathematical Theory of Pragmatic Information

The paper introduces a pragmatic information theory that unifies communication, control, and decision-making through the isoteleia mapping, which formalizes equifinality by treating distinct semantic paths that lead to the same optimal action as pragmatically equivalent. It establishes a three-tier hierarchy of syntactic, semantic, and pragmatic information, defines pragmatic entropy, mutual information, channel capacity, and rate-distortion, and proves coding theorems that generalize Shannon’s results. The authors also present pragmatic value and cost of information, a Lagrangian dual framework for cross-layer optimization, and a pragmatic efficiency bound that quantifies the maximum net utility for resource-constrained intelligent systems, extending the theory to continuous messages and dynamic settings.

By Kai Niu, Ping Zhang
arXiv AI
Jun 2

Information-Theoretic Lower Bounds for Bit-Constrained Stochastic Optimization via a Reduction to Compressed Gaussian Mean Estimation

arXiv:2606. 00703v1 Announce Type: cross Abstract: Low-precision pretraining (FP8, MXFP4, NVFP4) is now standard for frontier language models, yet the literature is almost entirely achievability -- algorithms and empirical scaling laws -- with no matching characterization of what is information-theoretically possible.

By Munsik Kim
arXiv AI
Jun 10

Minimum Distortion Quantization with Specified Output Distribution

arXiv:2606. 10458v1 Announce Type: cross Abstract: We derive the optimal quantizer of a real-valued random variable $W$ with distribution $P_W$ such that 1) the distribution of the quantization output $X$ that can take $k$ values follows any specified distribution $P_X$ over $\{1,\ldots,k\}$, and 2) the minimum mean squared error (MMSE) of estimating $W$ from $X$ is minimized.

By Aolin Xu
arXiv Machine Learning
Jun 10

The hyper-scaled NLP bound for maximum-entropy remote sampling

arXiv:2601. 20970v3 Announce Type: replace-cross Abstract: The maximum-entropy remote sampling problem (MERSP) is to select a subset of $s$ random variables from a set of $n$ random variables, so as to maximize the information concerning a set of target random variables that are not directly observable.

By Gabriel Ponte, Marcia Fampa, Jon Lee