arXiv Machine Learning

Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning

arXiv:2608. 06511v1 Announce Type: new Abstract: Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions.

arXiv Machine Learning
Aug 19

Training-Free Human-in-the-Loop Anomaly Detection via Memory Bank Correction

The paper introduces a training‑free, human‑in‑the‑loop anomaly detection framework that allows a domain expert to correct a PatchCore detector by editing its memory bank, without retraining or using gradients. Using only ten golden samples, operator corrections close a median 66% of the performance gap to a fully trained bank, improving 12 of 15 MVTec AD categories while harming none. The approach is evaluated with a rigorous held‑out protocol and shows that passive and active querying yield statistically indistinguishable gains, with a defect‑memory extension failing decisively.

By Ayusha Abbas, Saram Abbas, Kabita Adhikari
arXiv Computation and Language
Sep 2

Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops

The paper investigates code-level autonomous research loops (ARLs) where a language model edits training pipelines to improve an in-loop metric. It identifies a failure mode called algorithmic mode collapse, where edits become semantically uniform despite surface diversity, leading to a growing gap between in-loop gains and independent evaluation. The authors propose Diversity‑Aware Proposal Sampling (DAPS), a lightweight method that reduces semantic decay by 69.1% and boosts faithfulness by over 80% while maintaining optimization speed.

By Bowei He, Weixu Zhang, Yili Jin, Xue Liu
arXiv AI
Sep 24

Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark

The paper introduces a 96‑item paired benchmark for evaluating large language models (LLMs) on backtest auditing, where each flawed backtest is matched with a clean control that keeps strategy, dates, code style, labels, and reporting scaffold constant while altering a single methodological detail. A deterministic scorer distinguishes flaw recall, clean‑control false positives, evidence localization, and fix relevance. Experiments on 1440 cached audits from four text endpoints show that the DeepSeek auditor achieves perfect closed and clean‑aware code recall, yet open prompts over‑flag 93.8% of clean controls, and clean‑aware specificity is 87.5% even when recall saturates. Introducing a clean‑aware warning eliminates 20.8% false positives to 0% without affecting recall, though the budget anchor still flags many clean controls. Reporting clean‑control rates provides a clearer differentiation among models than reporting recall alone.

By Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov
arXiv AI
6d ago

Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality

The paper investigates how two signals—input‑conditional uncertainty and prediction‑label loss—detect different types of data corruption in federated learning. Experiments on ResNet‑20 with CIFAR‑10 and SVHN show that prediction‑label loss excels at spotting persistent random label flips, while expected‑entropy uncertainty better identifies additive image noise. The authors argue that effective federated data‑quality assessment must match the chosen signal to the specific corruption type rather than rely solely on uncertainty measures.

By Bradley Scott, Zeqi Luo, Edmond S. L. Ho