When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces a causal type discipline for sequential experiments that use AI-generated covariates. It defines a framework—including a versioned representation map, causal role classifier, claim-status filter, and estimand lock—to ensure that generated features are correctly classified as treatments, mediators, outcomes, or other roles, thereby preserving the intended causal estimand. The authors apply this framework to analyze compression bias, mediator adjustment, leakage, and other issues, demonstrating through simulations that careful refinement of generated covariates can reduce bias while design erasure or improper selection can lead to bias or undercoverage.
The paper introduces information set emulation, a method that attaches detailed causal certificates—such as source evidence, timing, and proposed causal roles—to AI‑derived features extracted from electronic health records (EHRs). These certificates provide auditable evidence for causal roles and guide whether a feature can be used for causal inference or should be routed to compatible reporting or separate analyses. The framework integrates with a joint EHR observation map and offers identification, estimation, and diagnostic tools under standard causal assumptions, illustrated through synthetic simulations and a finite‑world example.
arXiv:2604. 23904v3 Announce Type: replace-cross Abstract: Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference.
The paper introduces the concept of causal retention in interactive agents, examining whether a frozen learned state can correctly answer a mechanism‑probe map that is fixed independently of training. It shows that for finite structural causal models the optimal probe error is a Bayes decision risk, vanishing only when each learning‑interface fiber lies within a single probe‑answer fiber, and provides theoretical results such as a posterior‑coverage theorem and an exact edit decomposition. Experiments on finite causal systems, continuous simulators, TD‑MPC2, and Qwen2.5‑7B‑Instruct demonstrate that causal retention can be achieved with high accuracy, outperforming task‑performance‑based approaches.
arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.
arXiv:2609.06294v1 Announce Type: new Abstract: Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental s...