arXiv AI

When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments

The paper introduces a causal type discipline for sequential experiments that use AI-generated covariates. It defines a framework—including a versioned representation map, causal role classifier, claim-status filter, and estimand lock—to ensure that generated features are correctly classified as treatments, mediators, outcomes, or other roles, thereby preserving the intended causal estimand. The authors apply this framework to analyze compression bias, mediator adjustment, leakage, and other issues, demonstrating through simulations that careful refinement of generated covariates can reduce bias while design erasure or improper selection can lead to bias or undercoverage.

arXiv AI
Sep 17

Information Set Emulation: Causal Certificates for AI Derived EHR Features

The paper introduces information set emulation, a method that attaches detailed causal certificates—such as source evidence, timing, and proposed causal roles—to AI‑derived features extracted from electronic health records (EHRs). These certificates provide auditable evidence for causal roles and guide whether a feature can be used for causal inference or should be routed to compatible reporting or separate analyses. The framework integrates with a joint EHR observation map and offers identification, estimation, and diagnostic tools under standard causal assumptions, illustrated through synthetic simulations and a finite‑world example.

By Takes Fujita (VRI), Nobutaka Hattori (Department of Neurology, Juntendo University School of Medicine)
arXiv AI
4d ago

Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation

The paper introduces the concept of causal retention in interactive agents, examining whether a frozen learned state can correctly answer a mechanism‑probe map that is fixed independently of training. It shows that for finite structural causal models the optimal probe error is a Bayes decision risk, vanishing only when each learning‑interface fiber lies within a single probe‑answer fiber, and provides theoretical results such as a posterior‑coverage theorem and an exact edit decomposition. Experiments on finite causal systems, continuous simulators, TD‑MPC2, and Qwen2.5‑7B‑Instruct demonstrate that causal retention can be achieved with high accuracy, outperforming task‑performance‑based approaches.

By Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang, Cheng Zeng
arXiv AI
Aug 11

From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents

arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.

By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai
arXiv AI
4d ago

Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

The paper investigates how the definition of influence—specifically the behavior being attributed, the intervention on training data, and the counterfactual training process—affects rankings produced by influence estimators. It formalizes influence as a counterfactual estimand, distinguishes specification mismatch from approximation error, and categorizes existing estimators by their implied specifications. Experiments demonstrate that different specifications can lead to markedly different rankings, and that careful specification choice improves attribution quality in tasks such as noisy label detection and large‑language‑model attribution.

By Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun
arXiv Machine Learning
Aug 19

TabCausal: Pretraining Across Causal Environments for Tabular Causal Discovery

TabCausal is a causal discovery foundation model that learns to map datasets directly to causal graphs by pretraining across diverse causal environments. It uses a dynamic task construction strategy to expose the model to varied graph priors, mechanisms, noise models, dimensions, sample sizes, and intervention regimes, improving transferability from observational and mixed‑interventional data. On large synthetic benchmarks and a new protocol‑guided semantic benchmark, TabCausal outperforms many classical baselines and shows robust structure recovery, especially when interventional evidence is available.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv AI
Sep 15

Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents

The paper introduces PAI‑Bench, a benchmark designed to evaluate persistent AI agents on how faithfully they adhere to a versioned identity contract. It separates several dimensions—recall, composition, behavioral enactment, resistance, persistence, lineage, and role‑conditioned updates—while keeping scoring oracles independent of the target process. Experiments on synthetic profiles show that explicit cues can significantly alter the presence of identity identifiers, revealing prompt‑dependent component selection and sensitivity to startup cues.

By Zhenyu Zhao, Roy Zhao