arXiv Machine Learning

Testing Most Influential Sets

arXiv:2510. 20372v4 Announce Type: replace-cross Abstract: Small influential data subsets can dramatically impact model conclusions, with a few data points overturning key findings.

arXiv AI
Sep 10

From Citations to Contributions: LLM-Assisted Credit Scoring of Research Articles

The paper proposes a new method for credit scoring research articles that distinguishes between a paper’s original contribution and the prior work it builds upon. It introduces a hierarchical ‘contribution tree’ framework that conserves importance across a document’s structure and separates original from citation-derived credit. Large language models are employed as noisy comparative estimators to scale the analysis, and the approach is extended to collections of articles via weighted citation graphs to produce corpus-level contributions and normalized influence scores.

By Sana Ebrahimi, Suraj Shetiya, Abolfazl Asudeh
arXiv AI
Sep 10

Optimal Experiments for Partial Causal Effect Identification

The paper tackles selecting a cost‑constrained set of experiments that most effectively tighten bounds on a partially identifiable causal query. It formalizes this as the NP‑hard max‑potency problem, introduces efficient graphical pruning rules to reduce the search space, and demonstrates the approach on synthetic graphs and real NHANES data to estimate the effect of physical activity on diabetes.

By Tobias Maringgele, Jalal Etesami
arXiv Machine Learning
Sep 15

Data Attribution at Scale via Influence Matrix Estimation

Data Attribution at Scale via Influence Matrix Estimation proposes a scalable approach to quantify how individual training examples influence a model’s predictions. The authors introduce two algorithms, MAGE and SPELL, that reconstruct an influence matrix from a limited number of measurements without extra computational cost, improving over existing baselines across various training scales and budgets.

By Yuxi Chen, Hamza Golubovic, Han Tong, Arian Maleki, Andrew Ilyas
arXiv Machine Learning
Jul 31

What Is The Performance Ceiling of My Classifier? Utilizing Category-Wise Influence Functions for Pareto Frontier Analysis

arXiv:2510. 03950v2 Announce Type: replace Abstract: Data-centric learning seeks to improve model performance from the perspective of data quality, and has been drawing increasing attention in the machine learning community.

By Shahriar Kabir Nahin, Wenxiao Xiao, Joshua Liu, Anshuman Chhabra, Hongfu Liu
arXiv Machine Learning
Jul 23

Data-Poisoning Audits for Causal Effect Estimation

arXiv:2607. 19692v1 Announce Type: cross Abstract: Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect.

By Kwangho Kim