arXiv:2607. 19692v1 Announce Type: cross Abstract: Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect.
By Kwangho Kim
arXiv:2603. 13356v2 Announce Type: replace Abstract: Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global budget.
By Majid Ghasemi, Mark Crowley
arXiv:2608. 05490v1 Announce Type: new Abstract: Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision.
By Ahmed Hassoon, Mark Dredze
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
By Kaihua Ding
arXiv:2607. 11607v1 Announce Type: new Abstract: Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring.
By Hari Prasad
arXiv:2608.22160v1 Announce Type: new
Abstract: Physical automation is scaling toward fleets of embodied machines commanded by an AI brain. Early deployments already run factories and warehouses at p...
By Zhixu Du, Yiran Chen
arXiv:2604. 07650v2 Announce Type: replace Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent?
By Chenchen Kuai, Jiwan Jiang, Zihao Zhu, Hao Wang, Keshu Wu, Zihao Li, Yunlong Zhang, Chenxi Liu, Zhengzhong Tu, Zhiwen Fan, Yang Zhou
arXiv:2606. 17005v1 Announce Type: new Abstract: Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series shaped by reporting rules, benchmark revisions, and missingness.
By Yanan Long
arXiv:2606. 15640v1 Announce Type: new Abstract: Audit risk assessment increasingly benefits from combining heterogeneous evidence sources, yet existing approaches typically produce point predictions without quantifying how well different evidence streams agree.
By Yuhan Wang, Manqing Wang, Yixuan Lu, Zhaoyue Peng, Shengda Lin
arXiv:2606. 30338v1 Announce Type: new Abstract: External evaluations are becoming increasingly central to the governance of AI systems.
By Ioannis Pitsiorlas, Martha V. Sourla, Marios Kountouris
arXiv:2605. 06340v2 Announce Type: replace-cross Abstract: Continuous post-deployment compliance audits, mandated by emerging regulations such as the EU AI Act and Digital Services Act, create a class of strategic gaming distinct from the one-shot input/output gaming studied in prior work.
By Florian A. D. Burnat, Brittany I. Davidson
arXiv:2606. 30931v1 Announce Type: new Abstract: The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its statistical behavior remains poorly understood.
By Anish Acharya, Kris W Pan, Brian Verkhovsky