arXiv Machine Learning

Honest Physical-Support Inference after Latent Dictionary Learning: Collision Singularities and Minimax Resolution

arXiv:2607. 16813v1 Announce Type: new Abstract: Sparse-support uncertainty is usually quantified by treating the dictionary as known, an assumption that can produce overconfident, label-dependent conclusions when the dictionary is learned from latent sparse mixtures.

arXiv Machine Learning
Aug 21

Physical-Support Confidence Sets for Highly Coherent Dictionaries

arXiv:2608. 20295v1 Announce Type: new Abstract: Sparse pursuit after dictionary learning can yield a precise atom support even when its physical interpretation is not justified by the calibration data, especially for highly coherent dictionaries where alternative calibration-compatible dictionaries may assign different physical meanings to the same selected support.

By Guan-Ju Peng
arXiv Machine Learning
Sep 4

Resolution-Aware Experimental Design under Partial Identifiability

The paper introduces Resolution-Aware Experimental Design (RAED), a method that selects experiments by minimizing the expected size of the nonempty structural candidate set while controlling false-exclusion rates. RAED is shown to preserve expected ordering under a composite Blackwell comparison and is implemented via a learned score-based approach with finite-sample nuisance-average and positive-tail calibration. Experiments on subsurface-flow, fluvial, and methane-oxidation benchmarks demonstrate RAED’s ability to resolve structural ambiguities and provide finite-sample guarantees for tail-sensitive nuisance risk.

By Sofianos Panagiotis Fotias
arXiv Computation and Language
4d ago

The Cost of Compression: A Rate-Distortion Limit on Factual Hallucination

The paper presents a rate‑distortion framework for understanding factual hallucination in closed‑book question answering. It shows that even when a fact is observed, limited memory forces it to be stored approximately, leading to errors that can be bounded by a combination of compression distortion and missing coverage. The authors derive a theoretical lower bound on error and validate it with simulations and probes on modern language models.

By Xi Wang, Shijia Xu, Rongfeng Guo
arXiv Machine Learning
Aug 12

Optimistic Rates for Multiclass PAC Learning

arXiv:2608. 10869v1 Announce Type: new Abstract: Worst-case multiclass bounds do not become smaller when the best classifier is already nearly correct: what is missing is an optimistic rate, a guarantee whose fluctuation scales with the oracle risk itself.

By Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao
Hugging Face Trending Papers
Sep 3

Resolution-Aware Experimental Design under Partial Identifiability

The paper introduces Resolution-Aware Experimental Design (RAED), a method that selects experiments by minimizing the expected size of the nonempty structural candidate set while controlling false exclusions. RAED is shown to align with a composite Blackwell comparison and is implemented via a learned score-based approach with finite-sample calibration. Experiments on subsurface-flow, fluvial, and methane-oxidation benchmarks demonstrate that RAED can diverge from expected-information-gain selections, yielding clearer resolution and explicit ambiguity handling.