arXiv AI

Data Quality Profiling at Scale with Progressive Sampling: A Benchmark for Data-Centric AI Pipelines

arXiv:2607. 25356v1 Announce Type: cross Abstract: Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is foundational for data-centric AI pipelines, yet exhaustive scans over millions of rows are prohibitively slow for near-real-time monitoring.

arXiv AI
Sep 10

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

The study investigates how the composition of data during the mid‑training phase of language models affects performance across multiple domains. Experiments with Qwen3‑8B‑Base on five distinct KOR‑Bench domains show that moderate coverage (10%‑40%) yields the best per‑domain results, and that alignment passes cannot fully close the performance gaps created by mid‑training data choices. Additionally, zero coverage in mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.

By Yunpeng Xu, Kun Zheng
arXiv AI
Sep 12

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

The study examines the composition of a random sample from the Model Context Protocol (MCP) registry, revealing that only 48.8% of the 400 sampled npm/stdio servers successfully complete an initialization handshake, compared to 66.7% for a hand‑curated frame. Among the servers that run, there are no fatal JSON Schema violations across 2,766 advertised tools, but optional safety annotations vary widely, with a 58.8% omission rate in the random draw versus 41.5% in the curated set. The authors also compare MCP tool descriptions to two benchmark corpora, finding minimal near‑duplication in real MCP tools (2.8%) and significant repetition in synthetic datasets (up to 85.6%).

By Haseeb Mohammed Afsar
arXiv AI
Sep 15

Sampling headroom is not selection gain: a compute-value audit of test-time scaling for video world models

The paper introduces the Compute-Value Audit (CVA), a sequential framework that evaluates whether extra sampling during test‑time scaling for video world models actually yields a net benefit after accounting for the compute cost of generation and verification. On 192 Physics‑IQ scenes, increasing the sample pool from 4 to 16 candidates improves oracle quality by +9.23 IQ, yet common metrics such as Flow, Cycle, and VideoReward fail to reliably recover this headroom, and adaptive‑depth policies recover only 42‑69% of the potential gain. Only a few specific interventions—anchor‑explorer in a sparse PRM800K setting, MMLU‑Pro exposing a predictive‑state gap, and a privileged paired‑future upper bound—successfully pass all CVA stages, indicating that sampling headroom is valuable only when it can be converted into a reliable decision that survives the full compute charge.

By Yuhua Jiang, Junjie Lu, Feifei Gao
arXiv AI
Aug 20

Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Candidate explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking methods outperform single‑action RL baselines, while KG‑path selectors achieve near‑perfect user satisfaction scores.

By Tanay Chowdhury, Saeideh Shahrokh Esfahani
arXiv Machine Learning
Jul 2

A Filtered Mixture-of-Generators for Fully Synthetic Survival Training

arXiv:2607. 00127v1 Announce Type: new Abstract: Survival analysis models time-to-event data, but in clinical settings training data are costly and scarce: events accrue over years of follow-up, cohorts are small, and privacy regulations restrict sharing across institutions.

By Niccol\`o Maria Rizzi, Eugenio Lomurno, Alberto Archetti, Matteo Matteucci
arXiv Machine Learning
Sep 4

Sharpening the Ensemble: An SSIM-Aligned Residual Refiner for Brain-MRI Inpainting Post-Processing

The paper proposes a lightweight residual refiner that post‑processes the outputs of a two‑model ensemble for brain‑MRI inpainting. By training the refiner with an λ‑weighted combination of α loss and SSIM, the authors achieve a modest but statistically significant SSIM improvement (from 0.8767 to 0.8780 on a held‑out set) without altering MSE. Ablations show that adding a third model or using classical unsharp masking does not yield similar gains, indicating the improvement comes from learned sharpening rather than generic post‑processing.

By Kubilay Ka\u{g}an K\"om\"urc\"u, \.Ilkay \"Oks\"uz
Hugging Face Trending Papers
Aug 19

Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking outperforms single‑action RL methods, while KG‑path selectors achieve near‑perfect unique‑output rates, and the overall end‑to‑end build cost is around $15 on commodity hardware.