arXiv:2605.29695v2 Announce Type: replace
Abstract: Approximately 10% of newborns require assistance to initiate breathing at birth, and around 5% need ventilation support. Fetal heart rate (FHR) mon...
By Kjersti Engan, Neel Kanwal, Anita Yeconia, Ladislaus Blacy, Yuda Munyaw, Estomih Mduma, Hege Ersdal
arXiv:2606.21399v2 Announce Type: replace
Abstract: Runtime oversight often intervenes when an LLM agent's calibrated failure score crosses a threshold. Yet states with the same failure risk can diff...
By Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Qi Wen, Pengfei Zhou, Wangbo Zhao, Ivor Tsang
arXiv:2610.04188v2 Announce Type: replace
Abstract: Recent studies show that artificial intelligence (AI) with language and vision capabilities still experiences limitations in spatial reasoning. In...
By Uttamasha Monjoree, Wei Yan
arXiv:2610.04206v2 Announce Type: replace
Abstract: Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Archit...
By Uttamasha Monjoree, Wei Yan
arXiv:2610.05370v2 Announce Type: replace
Abstract: Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside...
By Xuancheng Li, Beining Wang, Haitao Li, Heng Wang, Yujia Zhou, Qingyi Pan, Blaze Chen, Yiqun Liu, Min Zhang, Qingyao Ai
arXiv:2610.05828v2 Announce Type: replace
Abstract: Large language models (LLMs) offer new opportunities for public opinion research by enabling early prediction of survey responses, potentially redu...
By Dongryeol Lee, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour
arXiv:2503.22764v3 Announce Type: replace-cross
Abstract: The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether m...
By Mingyuan Zhang, Yue Bai, Huan Wang, Yizhou Wang, Qihua Dong, Yitian Zhang, Yun Fu
arXiv:2505.14777v2 Announce Type: replace-cross
Abstract: The design of effective optimization algorithms for neural networks remains a fundamental challenge, and most existing methods rely on heuris...
By Mingquan Feng, Yixin Huang, Yifan Fu, Shaobo Wang, Junchi Yan
arXiv:2508.05880v3 Announce Type: replace-cross
Abstract: Understanding human emotions is central to user-facing AI applications, safety alignment, and the simulation of human behavior. As emotional...
By Sree Bhattacharyya, Evgenii Kuriabov, Lucas Craig, Tharun Dilliraj, Reginald B. Adams, Jr., Jia Li, James Z. Wang
arXiv:2508.10020v2 Announce Type: replace-cross
Abstract: Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especiall...
By Chuan Li, Qianyi Zhao, Fengran Mo, Cen Chen
arXiv:2510.11593v3 Announce Type: replace-cross
Abstract: For reliable large-scale quantum computation, quantum error correction (QEC) is essential to protect logical information distributed across m...
By Seong-Joon Park, Hee-Youl Kwak, Yongjune Kim
The paper introduces SORL, a framework that stabilizes off‑policy reinforcement learning for long‑horizon large language model agents. It identifies two key instability sources—token‑level policy granularity mismatches and high‑variance off‑policy updates—and proposes turn‑level importance sampling and clipping‑triggered normalization to align optimization with multi‑turn interactions. Two instantiations, SO‑PPO and SO‑GRPO, are evaluated on open‑domain, multi‑hop, and medical QA benchmarks, as well as on asynchronous RL for mathematical reasoning, showing improved robustness without the need for early stopping or heuristic tuning.
By Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong
arXiv:2512.11147v2 Announce Type: replace-cross
Abstract: AI agents are increasingly granted autonomous access to sensitive user data and third-party services, making effective permission management...
By Jinhao Zhu, Xiao Huang, Kevin Tseng, Gil Vernik, Shishir G. Patil, Vivian Fang, Raluca Ada Popa
arXiv:2603.01295v2 Announce Type: replace-cross
Abstract: Joint lesion segmentation and tissue classification in breast ultrasound are usually trained with a shared encoder, so the two branches stop...
By Abdullah Al Shafi, Md Kawsar Mahmud Khan Zunayed, Safin Ahmmed, Sk Imran Hossain, Engelbert Mephu Nguifo
arXiv:2603.04317v2 Announce Type: replace-cross
Abstract: A growing literature shows that variables can be linearly decoded from the activations of large language models (LLMs). These range from prop...
By Elan Barenholtz
arXiv:2603.16017v2 Announce Type: replace-cross
Abstract: Large language models (LLMs) increasingly participate in morally sensitive decision-making, yet how they organize ethical frameworks across r...
By Fan Huang, Haewoon Kwak, Jisun An
arXiv:2603.23184v2 Announce Type: replace-cross
Abstract: Despite the success of reinforcement learning from human feedback (RLHF), existing reward modeling methods largely rely on explicit feedback,...
By Hao Wang, Haocheng Yang, Licheng Pan, Lei Shen, Xiaoxi Li, Yinuo Wang, Zhichao Chen, Yuan Lu, Haoxuan Li, Zhouchen Lin
arXiv:2604.07925v2 Announce Type: replace-cross
Abstract: The self-attention mechanism is central to the success of Transformer architectures. However, standard row-stochastic attention has been show...
By Michela Lapenna, Rita Fioresi, Bahman Gharesifard
arXiv:2605.18850v2 Announce Type: replace-cross
Abstract: We introduce KadiAssistant, a privacy-by-design AI assistant integrated into the Kadi research data ecosystem, enabling researchers to effici...
By Adrian Cierpka, Mohammad Shafiqul Islam, Johannes Steinh\"ulb, Eric Dietriche Sesso Domtchoueng, Michael Selzer, Arnd Koeppe
arXiv:2605.18882v2 Announce Type: replace-cross
Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, si...
By Wei Shi, Ziheng Peng, Sihang Li, Xiting Wang, Xiang Wang, Mengnan Du, Na Zou