arXiv:2602. 17894v2 Announce Type: replace-cross Abstract: Data collection is a critical component of modern statistical and machine learning pipelines, particularly when data must be gathered from multiple heterogeneous sources to study a target population of interest.
By Michael O. Harding, Vikas Singh, Kirthevasan Kandasamy
arXiv:2606. 00563v1 Announce Type: cross Abstract: Selection bias is a common and often unavoidable aspect of real-world data that challenges the generalizability of machine learning models.
By Kara Liu, Maggie Wang, Russ B. Altman
arXiv:2510. 16882v4 Announce Type: replace-cross Abstract: Supervised fine-tuning (SFT) is a commonly used technique to adapt large language models (LLMs) to downstream tasks.
By Heming Zou, Yixiu Mao, Yun Qu, Qi Wang, Xiangyang Ji
arXiv:2609.16454v1 Announce Type: new
Abstract: Recent work by Doshi and Hauser (2024), Bisbee et al. (2024), and Xie et al. (2026) raises concerns that outputs from large language models (LLMs) tend...
By Kirill Skobelev, Eric Fithian, X. Y. Han
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert
arXiv:2604. 11305v3 Announce Type: replace Abstract: Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality requirement, while controlling the false discovery rate (FDR).
By Meiyi Zhu, Osvaldo Simeone
arXiv:2510. 16657v3 Announce Type: replace-cross Abstract: Synthetic data has been increasingly used to train frontier generative models.
By Bingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng Xu
The paper introduces SynthSentry, a model‑agnostic method for detecting synthetic data contamination in language‑model training corpora. It computes a distributional divergence score based on lexical diversity collapse, n‑gram tail truncation, and perplexity variance across reference models, requiring no access to the generating model or synthetic labels. Experiments on English corpora contaminated by small open‑weight generators and an instruction‑tuned model show that SynthSentry ranks contamination severity accurately, maintains low false‑positive rates after calibration, and does not degrade downstream fine‑tuning performance at the tested scale.
By Praveen Kumar Myakala, Ravichandra Namburi, Sowmya Keragodu Jayaramu, Sooraj George Thomas
arXiv:2511.22435v2 Announce Type: replace
Abstract: Invariant learning on graphs aims to build predictors that rely on causal substructures rather than on environment-specific shortcuts. Current meth...
By Ali Ghasemi, Farooq Ahmad Wani, Maria Sofia Bucarelli, Fabrizio Silvestri
arXiv:2607. 14157v1 Announce Type: cross Abstract: Retrieval over corpora that mix several domains often returns relevant but wrong-domain evidence that ranking metrics miss and that conformal risk control bounds only marginally, under-covering the worst domains.
By Jayakumar Manoharan
arXiv:2607. 20787v1 Announce Type: cross Abstract: For two decades, the standard remedy for class-imbalanced learning has been to fabricate synthetic minority examples, and the standard evidence of their validity has been a check that cannot fail: synthetic points are scored against the very data that generated them.
By Ahmad B. Hassanat, Ahmad S. Tarawneh, Ghada A. Altarawneh
arXiv:2603. 15158v2 Announce Type: replace Abstract: Addressing the domain adaptation problem becomes more challenging when distribution shifts across domains stem from latent confounders that affect both covariates and outcomes.
By Zahra Rahiminasab, Reza Soumi, Arto Klami, Samuel Kaski