arXiv:2606. 30801v1 Announce Type: cross Abstract: Personalization algorithms determine what content users encounter on online platforms.
By Alessandro Morosini, Sarah H. Cen, Andrew Ilyas, Hedi Driss, Aleksander M\k{a}dry, Chara Podimata
arXiv:2603. 00829v2 Announce Type: replace-cross Abstract: Safe deployment of Large Language Model (LLM) agents in autonomous settings requires reliable oversight mechanisms.
By Simon Storf, Rich Barton-Cooper, James Peters-Gill, Marius Hobbhahn
The paper examines whether internal representations of agentic systems can better indicate task success than traditional confidence measures. It introduces two methods—Latent Trajectory Dynamics (LTD) and Action Representation Probe (ARP)—that analyze changes in residual-stream representations and action-level representations, respectively. Experiments on Bash, SQL, and Python benchmarks with Qwen and DeepSeek models show these methods outperform conventional surface-level and sequence-based calibration baselines, offering a zero‑overhead reliability monitor without prompt changes or multiple rollouts.
By Priyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla
arXiv:2606. 15306v1 Announce Type: cross Abstract: We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared across those tasks and use it to improve future decisions.
By Daksh Mittal, Tommaso Castellani, Thomson Yen, Naimeng Ye, Fangyu Wu, Minghui Chen, Tiffany Cai, Emmanouil Koukoumidis, William Zeng, Hongseok Namkoong
arXiv:2607. 11228v1 Announce Type: cross Abstract: While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases.
By Anqi Li, Jie Zhang, Zhongqi Wang, Songkai Xue, Jiahao Wang, Shiguang Shan, Xilin Chen
EvoTS-Agent is a self‑evolving large language model agent designed for autonomous change‑point detection in financial time series. It begins with curated exploratory data analysis to set up candidate models, then iteratively refines detection pipelines using three operators—Revision, Alternative Strategy, and Recombination—guided by validation feedback. Across four benchmark datasets, EvoTS-Agent consistently outperforms existing LLM‑based agents and achieves a 100% execution success rate with all tested backbone LLMs.