arXiv Machine Learning

Quality Control Algorithms for Pattern Counting

arXiv:2608. 03439v1 Announce Type: cross Abstract: In recent work, Marcussen, Rubinfeld, and Sudan introduced the notion of quality control problems, which aim to capture the task of determining if a given input is truly random.

arXiv Machine Learning
Jul 9

Is Randomness Necessary for Adaptive Data Analysis?

arXiv:2607. 07085v1 Announce Type: cross Abstract: The Adaptive Data Analysis (ADA) problem formalizes the challenge of preventing false discovery and overfitting when a dataset is repeatedly reused.

By Edith Cohen, Haim Kaplan, Yishay Mansour, Shay Sapir, Uri Stemmer
arXiv Machine Learning
Jul 28

Hallucination Rates in Language Generation

arXiv:2607. 23361v1 Announce Type: cross Abstract: Language generation in the limit is an elegant model introduced by Kleinberg and Mullainathan [KM24] to formally study language generation by an algorithm that learns solely based on example strings.

By Debmalya Panigrahi, Fan Wei, Ian Zhang
arXiv Machine Learning
4d ago

Local Search with Correlated Randomness

arXiv:2607.17469v2 Announce Type: replace-cross Abstract: How much does an algorithm's running-time distribution under independent randomness reveal about its behavior when independence is no longer...

By Yunbei Xu
arXiv Machine Learning
1d ago

Adaptive and oblivious statistical adversaries are equivalent

The paper resolves a key question in statistical learning under adversarial corruption by showing that sample‑adaptive and sample‑oblivious adversaries are equivalent up to polynomial factors in the sample size for all corruption types. It proves that any algorithm that succeeds against a sample‑oblivious adversary can be transformed into one that succeeds against the corresponding sample‑adaptive adversary by requesting a polynomially larger sample and running the original algorithm on a random subsample. The construction preserves computational efficiency and requires only a simple modification of the algorithm.

By Guy Blanc, Gregory Valiant
arXiv Machine Learning
Jul 28

Learning Distributions from Multiple Data Providers

arXiv:2607. 24732v1 Announce Type: cross Abstract: Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples.

By Jon Kleinberg, Amin Saberi, Xizhi Tan, Grigoris Velegkas