arXiv Computation and Language By Hanbo Xie

Better Behavioral Prediction, More Faithful Model Ablations? Evidence from Sequential Choice

Read the original on arXiv Computation and Language →

The paper investigates whether input ablations on predictive models can reliably reveal the importance of information for explaining human sequential choice behavior. Using two synthetic bandit tasks with known generating policies, the authors compare GRUs, Transformers, a fine‑tuned LLaMA, and cognitive models under varied reward contributions. They find that while neural models can predict choices well, their responses to ablations often diverge from the true generating process, indicating that predictive accuracy alone does not guarantee faithful model ablations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
3d ago

What Pretraining and Midtraining Make Learnable from Rewards?

The paper investigates how pretraining and midtraining enable reward-based learning by providing necessary information and computation. It analyzes sequential state computation and contextual memory, showing that task‑independent source observations resolve ambiguities in reward adaptation. Experiments on pretrained Qwen2.5 checkpoints across eight worlds demonstrate that correct source and first‑operation supervision significantly improve success rates, and that memory replay and independent confirmation further enhance performance.

By Chiwun Yang, Xiaoyu Li
arXiv AI
Aug 26

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

CAFE (Coupled Agent–Feedback Evolution) is a framework that lets a shared‑parameter model alternate between acting as a search agent and as a critic that provides corrective feedback. By learning when to request feedback and how to use it, CAFE trains the agent to recover from its own failures and shapes rewards both online and offline. Experiments on seven search benchmarks show that CAFE outperforms other RL‑based agents, maintains gains on out‑of‑domain tests, and reduces hallucinations, indicating that co‑evolving feedback is essential for self‑improving search agents.

By Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou, Xinbing Liang, Shizheng Zhu, Yuhui Wang, Jingqi Tong, Zhiheng Xi, Jiazheng Zhang, Clive Bai, Clarenceai, Blaze Chen, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen