The paper introduces a counterfactual tool ranking framework that accounts for authority, historical support, and estimation nuances. Using eleven enterprise-inspired tools, synthetic and real-world experiments on the Berkeley Function Calling Leaderboard, the study compares direct regression and doubly robust (DR) methods, finding that DR performs better in shifted environments while direct regression excels in linear settings. The authors also evaluate Qwen2.5 models on held-out tasks, analyze policy differences under missing support, and present a falsifiable evaluation method with publicly available evidence.
By Jiapeng Li
arXiv:2606. 10632v1 Announce Type: cross Abstract: Lipschitz-style individual fairness formalizes the idea that semantically similar examples should receive similar predictions, but its evaluation in multi-task learning (MTL) can be confounded by method-induced representation scales.
By Junbo Ding, Xin Zang, Chenchen Pan, Donghao Song, Jiaxin Zhu, Danhuai Guo
arXiv:2606. 27948v1 Announce Type: new Abstract: Counterfactual explanations (CFs) help understand machine learning models by identifying minimal input changes that would lead to alternative model outcomes.
By Xuan Zhao, Lena Krieger, Zhuo Cao, Arya Bangun, Hanno Scharr, Ira Assent
arXiv:2606. 08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observability) or whether it passed (evaluation), but not which step caused the failure.
By Jaineet Shah
The paper introduces a controlled contrast framework for estimating peer effects when interaction graphs evolve, indexing potential outcomes by own treatment, temporally aggregated peer exposure, and a post‑assignment evolution summary. It proposes the Dynamic Network Doubly Robust estimator (DynaNet‑DR), which uses a temporally factorized propensity and normalized augmentation to achieve consistency under standard causal assumptions. Semi‑synthetic benchmarks on real temporal graph sequences demonstrate that DynaNet‑DR achieves favorable estimation accuracy compared to other methods, and an observational study on MathOverflow illustrates its practical application.
By Xiaojing Du
arXiv:2608. 06469v1 Announce Type: cross Abstract: Collaborative machine learning among financial institutions must be both group-fair and robust against deliberate adversarial manipulation.
By Devharsh Trivedi, Nesrine Kaaniche, Nikos Triandopoulos, Maryline Laurent, Jackson Walters