arXiv AI

Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation

arXiv:2606. 31184v1 Announce Type: cross Abstract: Adaptive experiments for average treatment effects (ATE) require randomized allocations balancing valid inference with statistical efficiency.

arXiv Machine Learning
Jun 25

Efficient Adaptive Data Acquisition via Pretrained Belief Representations

arXiv:2606. 25197v1 Announce Type: new Abstract: Learning effective policies for adaptive data acquisition remains challenging: posterior-based methods rely on surrogate models and posterior approximations that can be misspecified or biased, while direct policy-learning methods map from historical observations and fail to exploit available model representations, making learning harder.

By Daolang Huang, Zhuoyue Huang, Conor Hassan, Luigi Acerbi, Samuel Kaski, Tom Rainforth
arXiv Machine Learning
Sep 18

Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation

The paper introduces Bidirectional Behavior Prior Distillation (B2PD), a method that uses action‑value priors to train a conditional variational autoencoder for generating high‑value behavior support. These expert behavior priors are then distilled into the online reinforcement learning agent, reducing inefficient exploration and stabilizing policy updates. Experiments on state‑ and pixel‑based tasks show that B2PD improves sample efficiency while maintaining stable learning dynamics.

By Gong Gao, Xiao Lai, Jiaji Shen, Ning Jia, Xianhui Liu, Weidong Zhao
arXiv Machine Learning
Sep 17

Correcting Boundary Bias and Observation Independence in Bayesian Experimental Design

The paper tackles two shortcomings of Gaussian‑process based active learning: (1) the posterior variance is independent of observed values, reducing sensitivity to data structure, and (2) it over‑inflates variance near domain boundaries, causing excessive edge sampling. The authors propose a reconstruction‑driven design density that warps sampling toward regions where the posterior mean changes rapidly, and a geometric equalizer that corrects boundary bias. Experiments on sixteen synthetic and two real‑data benchmarks show that the equalizer consistently improves function reconstruction, while the warp further enhances performance by concentrating measurements where the target function varies most.

By Sanna Jarl, Jens Sj\"olund, Jonathan J. S. Scragg, Maria B{\aa}nkestad