arXiv:2607. 09706v1 Announce Type: new Abstract: Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objective and point-valued coefficients, then solve once.
By Suyash Mishra
arXiv:2603. 01119v2 Announce Type: replace-cross Abstract: A fundamental challenge in causal inference with observational data is correct specification of a causal model.
By Rohit Bhattacharya, Ina Ocelli, Ted Westling
Counterfactual explanations are widely used to provide algorithmic recourse in high-stakes decision-making systems. Most existing methods seek the smallest change to an input that flips a model's decision.
arXiv:2606. 18832v1 Announce Type: cross Abstract: Counterfactual explanations are widely used to provide algorithmic recourse in high-stakes decision-making systems.
By K. Darshana Abeyrathna, Sara El Mekkaoui, Nils Enric Canut Taugb{\o}l, Anuja Vats
The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2607. 11653v1 Announce Type: new Abstract: Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management.
By Ivane Antonov, Sohom Mukherjee, Richard Pibernik, Yo Joong Choe
arXiv:2606. 30931v1 Announce Type: new Abstract: The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its statistical behavior remains poorly understood.
By Anish Acharya, Kris W Pan, Brian Verkhovsky
The paper introduces a counterfactual tool ranking framework that accounts for authority, historical support, and estimation nuances. Using eleven enterprise-inspired tools, synthetic and real-world experiments on the Berkeley Function Calling Leaderboard, the study compares direct regression and doubly robust (DR) methods, finding that DR performs better in shifted environments while direct regression excels in linear settings. The authors also evaluate Qwen2.5 models on held-out tasks, analyze policy differences under missing support, and present a falsifiable evaluation method with publicly available evidence.
By Jiapeng Li
JuryProbe is an empirical diagnostic tool designed to assess consensus risk in panels of reference‑free large language model judges used for factuality verification. It estimates risk by measuring false‑negative correlations and false‑consensus lift from a labeled calibration probe, and routes high‑risk majority decisions to judges with trusted references. The approach was validated on FEVER corruptions, showing that flagged decisions can be grounded without additional reference acquisition in most cases, while reducing false accepts by about 0.4% and avoiding 28% of reference acquisitions.
By Tianxin Zhou, Ruixi Lin
arXiv:2607. 19692v1 Announce Type: cross Abstract: Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect.
By Kwangho Kim
arXiv:2603. 05175v2 Announce Type: replace Abstract: The rapid proliferation of AI applications has intensified debate on effective regulation of these black-box services.
By Anurag Singh, Julian Rodemann, Rajeev Verma, Siu Lun Chau, Krikamol Muandet
The paper investigates how the definition of influence—specifically the behavior being attributed, the intervention on training data, and the counterfactual training process—affects rankings produced by influence estimators. It formalizes influence as a counterfactual estimand, distinguishes specification mismatch from approximation error, and categorizes existing estimators by their implied specifications. Experiments demonstrate that different specifications can lead to markedly different rankings, and that careful specification choice improves attribution quality in tasks such as noisy label detection and large‑language‑model attribution.
By Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun