The study investigates whether provenance information can reliably identify the source of synthetic text and whether this identification improves the selection of training data. Using financial‑risk text, the authors achieve 98.7% accuracy in attributing original generated passages, but accuracy drops to 53.1% after paraphrasing and 29.0% after style rewriting. They compare two selection strategies—one based on source provenance and another on a reference model score—across three rounds of generation and retraining, finding that the two methods choose different examples but do not produce a consistent difference in model degradation. The results suggest that source attribution and useful data selection are distinct challenges, and neither provenance nor the tested proxy suffices to guarantee stable recursive training behavior.
By Joss Armstrong
arXiv:2606.26403v2 Announce Type: replace
Abstract: Foundation-model research increasingly needs data about people: user state, personal histories, relationships, contact-like fields, documents, and...
By Sriram Selvam, Anneswa Ghosh
arXiv:2606. 05403v1 Announce Type: new Abstract: Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions.
By Rohan N. Pradhan, Steve Goley
arXiv:2607. 23804v1 Announce Type: cross Abstract: Context attribution methods for large language models (LLMs) identify which input context contributes to the model response.
By Quoc-Huy Trinh, Lin Zhu, Sebastian Szyller
The paper introduces CAMS, a Claim‑Anchored Multi‑Document Summarization framework that decomposes source documents into atomic claims, resolves provenance deterministically from verbatim quotes to token spans, clusters equivalent claims across documents, and rewrites summaries so each sentence ends with claim identifiers linking back to source spans. CAMS separates provenance (an invariant for each emitted sentence) from faithfulness (an objective encouraged by selection, rewriting, and verification). Evaluations on MultiNews, DiverseSumm, and zero‑shot WCEP show that CAMS matches strong baselines in summary quality while improving faithfulness and citation precision, raising attribution accuracy from 38% to 64% and reducing human verification time per claim by 3.4×.
By Shuo Guan
arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.
By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai