arXiv:2610.08669v1 Announce Type: cross
Abstract: On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficien...
By Mehmet Emre Akbulut, Johannes Geier, Ulf Schlichtmann
arXiv:2610.08670v1 Announce Type: cross
Abstract: Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does...
By Orion Reblitz-Richardson
arXiv:2507.07426v4 Announce Type: replace
Abstract: Recent advances in large language models have demonstrated considerable potential in scientific domains such as drug repositioning. However, their...
By Zerui Yang, Yuwei Wan, Siyu Yan, Yudai Matsuda, Tong Xie, Linqi Song
arXiv:2508.16821v2 Announce Type: replace
Abstract: We introduce PuzzleJAX, a GPU-accelerated puzzle game engine and description language designed to support rapid benchmarking of tree search, reinfo...
By Sam Earle, Graham Todd, Yuchen Li, Ahmed Khalifa, Muhammad Umair Nasir, Zehua Jiang, Andrzej Banburski-Fahey, Julian Togelius
arXiv:2511.20471v3 Announce Type: replace
Abstract: Recent advances in Large Language Model (LLM) reasoning have improved conventional problem solving, but creative reasoning remains comparatively un...
By Yuto Suzuki, Farnoush Banaei-Kashani
arXiv:2605.10791v2 Announce Type: replace
Abstract: Knowledge Graph Question Answering (KGQA) aims to answer user questions by reasoning over Knowledge Graphs (KGs). Recent methods use supervision de...
By Shengxiang Gao, Chao Lei, Jey Han Lau, Linhao Luo, Jianzhong Qi
arXiv:2605.12922v2 Announce Type: replace
Abstract: Large language models can follow complex instructions in a single turn, yet over long multi-turn interactions they often lose the thread of instruc...
By Vardhan Dongre, Joseph Hsieh, Viet Dac Lai, Seunghyun Yoon, Trung Bui, Dilek Hakkani-T\"ur
arXiv:2605.24154v2 Announce Type: replace
Abstract: Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users an...
By Qitao Tan, Xiaoying Song, Arman Akbari, Arash Akbari, Yanzhi Wang, Xiaoming Zhai, Lingzi Hong, Zhen Xiang, Jin Lu, Geng Yuan
arXiv:2605.29695v2 Announce Type: replace
Abstract: Approximately 10% of newborns require assistance to initiate breathing at birth, and around 5% need ventilation support. Fetal heart rate (FHR) mon...
By Kjersti Engan, Neel Kanwal, Anita Yeconia, Ladislaus Blacy, Yuda Munyaw, Estomih Mduma, Hege Ersdal
arXiv:2606.21399v2 Announce Type: replace
Abstract: Runtime oversight often intervenes when an LLM agent's calibrated failure score crosses a threshold. Yet states with the same failure risk can diff...
By Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Qi Wen, Pengfei Zhou, Wangbo Zhao, Ivor Tsang
arXiv:2610.04188v2 Announce Type: replace
Abstract: Recent studies show that artificial intelligence (AI) with language and vision capabilities still experiences limitations in spatial reasoning. In...
By Uttamasha Monjoree, Wei Yan
arXiv:2610.04206v2 Announce Type: replace
Abstract: Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Archit...
By Uttamasha Monjoree, Wei Yan
arXiv:2610.05370v2 Announce Type: replace
Abstract: Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside...
By Xuancheng Li, Beining Wang, Haitao Li, Heng Wang, Yujia Zhou, Qingyi Pan, Blaze Chen, Yiqun Liu, Min Zhang, Qingyao Ai
arXiv:2610.05828v2 Announce Type: replace
Abstract: Large language models (LLMs) offer new opportunities for public opinion research by enabling early prediction of survey responses, potentially redu...
By Dongryeol Lee, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour
arXiv:2503.22764v3 Announce Type: replace-cross
Abstract: The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether m...
By Mingyuan Zhang, Yue Bai, Huan Wang, Yizhou Wang, Qihua Dong, Yitian Zhang, Yun Fu
arXiv:2505.14777v2 Announce Type: replace-cross
Abstract: The design of effective optimization algorithms for neural networks remains a fundamental challenge, and most existing methods rely on heuris...
By Mingquan Feng, Yixin Huang, Yifan Fu, Shaobo Wang, Junchi Yan
arXiv:2508.05880v3 Announce Type: replace-cross
Abstract: Understanding human emotions is central to user-facing AI applications, safety alignment, and the simulation of human behavior. As emotional...
By Sree Bhattacharyya, Evgenii Kuriabov, Lucas Craig, Tharun Dilliraj, Reginald B. Adams, Jr., Jia Li, James Z. Wang
arXiv:2508.10020v2 Announce Type: replace-cross
Abstract: Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especiall...
By Chuan Li, Qianyi Zhao, Fengran Mo, Cen Chen
arXiv:2510.11593v3 Announce Type: replace-cross
Abstract: For reliable large-scale quantum computation, quantum error correction (QEC) is essential to protect logical information distributed across m...
By Seong-Joon Park, Hee-Youl Kwak, Yongjune Kim
The paper introduces SORL, a framework that stabilizes off‑policy reinforcement learning for long‑horizon large language model agents. It identifies two key instability sources—token‑level policy granularity mismatches and high‑variance off‑policy updates—and proposes turn‑level importance sampling and clipping‑triggered normalization to align optimization with multi‑turn interactions. Two instantiations, SO‑PPO and SO‑GRPO, are evaluated on open‑domain, multi‑hop, and medical QA benchmarks, as well as on asynchronous RL for mathematical reasoning, showing improved robustness without the need for early stopping or heuristic tuning.
By Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong