arXiv:2606. 14965v1 Announce Type: new Abstract: Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically generate noise through imperfect annotators or classifier raters, leaving the source of ambiguity implicit.
By Shadman Islam, Agustinus Kristiadi, Mostafa Milani
arXiv:2606. 10229v1 Announce Type: cross Abstract: We study whether demonstration-curation metrics that detect defective training episodes also improve the downstream behavior-cloning policy that trains on the curated data.
By Aarav Bedi
arXiv:2604. 18245v3 Announce Type: replace Abstract: Large language models operate in protocols containing multiple calls, yet added calls are usually evaluated only by their net effect.
By Fernando Reitich
The paper introduces a training‑free, human‑in‑the‑loop anomaly detection framework that allows a domain expert to correct a PatchCore detector by editing its memory bank, without retraining or using gradients. Using only ten golden samples, operator corrections close a median 66% of the performance gap to a fully trained bank, improving 12 of 15 MVTec AD categories while harming none. The approach is evaluated with a rigorous held‑out protocol and shows that passive and active querying yield statistically indistinguishable gains, with a defect‑memory extension failing decisively.
By Ayusha Abbas, Saram Abbas, Kabita Adhikari
arXiv:2606. 11616v1 Announce Type: new Abstract: High-quality training data is essential for the success of machine learning models.
By Jiale Deng, Yanyan Shen, Xiaogang Shi, Chai Junjun
The paper investigates code-level autonomous research loops (ARLs) where a language model edits training pipelines to improve an in-loop metric. It identifies a failure mode called algorithmic mode collapse, where edits become semantically uniform despite surface diversity, leading to a growing gap between in-loop gains and independent evaluation. The authors propose Diversity‑Aware Proposal Sampling (DAPS), a lightweight method that reduces semantic decay by 69.1% and boosts faithfulness by over 80% while maintaining optimization speed.
By Bowei He, Weixu Zhang, Yili Jin, Xue Liu
arXiv:2607. 01280v1 Announce Type: new Abstract: Programming-by-example systems infer programs from a small set of input-output examples.
By Yuan Si, Jialu Zhang
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert
The paper introduces a 96‑item paired benchmark for evaluating large language models (LLMs) on backtest auditing, where each flawed backtest is matched with a clean control that keeps strategy, dates, code style, labels, and reporting scaffold constant while altering a single methodological detail. A deterministic scorer distinguishes flaw recall, clean‑control false positives, evidence localization, and fix relevance. Experiments on 1440 cached audits from four text endpoints show that the DeepSeek auditor achieves perfect closed and clean‑aware code recall, yet open prompts over‑flag 93.8% of clean controls, and clean‑aware specificity is 87.5% even when recall saturates. Introducing a clean‑aware warning eliminates 20.8% false positives to 0% without affecting recall, though the budget anchor still flags many clean controls. Reporting clean‑control rates provides a clearer differentiation among models than reporting recall alone.
By Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov
The paper investigates how two signals—input‑conditional uncertainty and prediction‑label loss—detect different types of data corruption in federated learning. Experiments on ResNet‑20 with CIFAR‑10 and SVHN show that prediction‑label loss excels at spotting persistent random label flips, while expected‑entropy uncertainty better identifies additive image noise. The authors argue that effective federated data‑quality assessment must match the chosen signal to the specific corruption type rather than rely solely on uncertainty measures.
By Bradley Scott, Zeqi Luo, Edmond S. L. Ho
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.
By Arastoo Zibaeirad, Marco Vieira