arXiv:2606. 24589v1 Announce Type: new Abstract: Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm that resulting failures are real.
By Khanak Khandelwal (Indian Institute of Technology Jodhpur)
The study investigates whether provenance information can reliably identify the source of synthetic text and whether this identification improves the selection of training data. Using financial‑risk text, the authors achieve 98.7% accuracy in attributing original generated passages, but accuracy drops to 53.1% after paraphrasing and 29.0% after style rewriting. They compare two selection strategies—one based on source provenance and another on a reference model score—across three rounds of generation and retraining, finding that the two methods choose different examples but do not produce a consistent difference in model degradation. The results suggest that source attribution and useful data selection are distinct challenges, and neither provenance nor the tested proxy suffices to guarantee stable recursive training behavior.
By Joss Armstrong
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert
Large language model (LLM) agents require post-training methods that can improve long-horizon decision making from environment feedback. However, existing agentic post-training pipelines often treat data curation as a fixed preprocessing step, focusing mainly on data augmentation while neglecting filtering, refinement, and adaptation to downstream failures.
arXiv:2610.00767v1 Announce Type: cross
Abstract: Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later trai...
By Peter Nutter, Dani Roytburg, Cl\'ement Dumas, Jinghua Ou, Shi Feng
arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.
By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai