arXiv Machine Learning By Joss Armstrong

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

Read the original on arXiv Machine Learning →

The study investigates whether provenance information can reliably identify the source of synthetic text and whether this identification improves the selection of training data. Using financial‑risk text, the authors achieve 98.7% accuracy in attributing original generated passages, but accuracy drops to 53.1% after paraphrasing and 29.0% after style rewriting. They compare two selection strategies—one based on source provenance and another on a reference model score—across three rounds of generation and retraining, finding that the two methods choose different examples but do not produce a consistent difference in model degradation. The results suggest that source attribution and useful data selection are distinct challenges, and neither provenance nor the tested proxy suffices to guarantee stable recursive training behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects

The paper proposes a generation‑provenance substrate for synthetic speech research objects, binding source specifications, generated content, waveform, target, fact requirements, quality signals, review lineage, and an immutable manifest identity. It audits this substrate in a private Japanese care‑handoff pipeline, documenting 113 assets and 1.552 hours of synthetic speech with linked audio, transcripts, notes, and fact checklists, while noting selective human evidence and source‑specific gaps. The authors argue that provenance is necessary but not sufficient for behavior attribution, requiring additional frozen training runs and intervention evidence, and they provide a compact provenance contract, audit protocol, and a bounded case study. "whyItMatters":"The study highlights the need for detailed provenance records to enable reliable auditing and attribution of synthetic data behavior, underscoring limitations in current practices and offering a structured framework for future research."

By Sidi Chang, Peiying Zhu
arXiv Computation and Language
Sep 14

SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data

The paper introduces SynthSentry, a model‑agnostic method for detecting synthetic data contamination in language‑model training corpora. It computes a distributional divergence score based on lexical diversity collapse, n‑gram tail truncation, and perplexity variance across reference models, requiring no access to the generating model or synthetic labels. Experiments on English corpora contaminated by small open‑weight generators and an instruction‑tuned model show that SynthSentry ranks contamination severity accurately, maintains low false‑positive rates after calibration, and does not degrade downstream fine‑tuning performance at the tested scale.

By Praveen Kumar Myakala, Ravichandra Namburi, Sowmya Keragodu Jayaramu, Sooraj George Thomas
arXiv AI
Jun 10

Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation

arXiv:2606. 11127v1 Announce Type: cross Abstract: Synthetic post-training pipelines commonly filter generated samples with reward models or holistic LLM judges, yet two practices remain rarely examined together: whether the filtering signal is grounded in the source evidence that induced each generation, and whether rejected samples can be systematically recovered rather than permanently discarded.

By Soham Bhattacharjee, Karun Sharma, Vinay Kumar Sankarapu, Pratinav Seth