arXiv:2508.16821v2 Announce Type: replace
Abstract: We introduce PuzzleJAX, a GPU-accelerated puzzle game engine and description language designed to support rapid benchmarking of tree search, reinfo...
By Sam Earle, Graham Todd, Yuchen Li, Ahmed Khalifa, Muhammad Umair Nasir, Zehua Jiang, Andrzej Banburski-Fahey, Julian Togelius
arXiv:2511.20471v3 Announce Type: replace
Abstract: Recent advances in Large Language Model (LLM) reasoning have improved conventional problem solving, but creative reasoning remains comparatively un...
By Yuto Suzuki, Farnoush Banaei-Kashani
arXiv:2605.10791v2 Announce Type: replace
Abstract: Knowledge Graph Question Answering (KGQA) aims to answer user questions by reasoning over Knowledge Graphs (KGs). Recent methods use supervision de...
By Shengxiang Gao, Chao Lei, Jey Han Lau, Linhao Luo, Jianzhong Qi
arXiv:2605.27593v2 Announce Type: replace
Abstract: Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret c...
By Xijie Zeng, Frank Rudzicz
arXiv:2605.29695v2 Announce Type: replace
Abstract: Approximately 10% of newborns require assistance to initiate breathing at birth, and around 5% need ventilation support. Fetal heart rate (FHR) mon...
By Kjersti Engan, Neel Kanwal, Anita Yeconia, Ladislaus Blacy, Yuda Munyaw, Estomih Mduma, Hege Ersdal
arXiv:2610.04188v2 Announce Type: replace
Abstract: Recent studies show that artificial intelligence (AI) with language and vision capabilities still experiences limitations in spatial reasoning. In...
By Uttamasha Monjoree, Wei Yan
arXiv:2610.04875v2 Announce Type: replace
Abstract: Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this pro...
By Chung-En Ho, Weiyu Sun, Cheng-Jhih Shih, He Li, Yong Liu, Yingyan Celine Lin
arXiv:2610.05370v2 Announce Type: replace
Abstract: Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside...
By Xuancheng Li, Beining Wang, Haitao Li, Heng Wang, Yujia Zhou, Qingyi Pan, Blaze Chen, Yiqun Liu, Min Zhang, Qingyao Ai
arXiv:2610.05828v2 Announce Type: replace
Abstract: Large language models (LLMs) offer new opportunities for public opinion research by enabling early prediction of survey responses, potentially redu...
By Dongryeol Lee, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour
arXiv:2503.22764v3 Announce Type: replace-cross
Abstract: The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether m...
By Mingyuan Zhang, Yue Bai, Huan Wang, Yizhou Wang, Qihua Dong, Yitian Zhang, Yun Fu
arXiv:2508.14748v2 Announce Type: replace-cross
Abstract: The increasing variety of molecular data creates a need for generative models that can flexibly incorporate heterogeneous constraints across...
By Yunzhe Zhang, Yifei Wang, Khanh Vinh Nguyen, Pengyu Hong
arXiv:2510.11593v3 Announce Type: replace-cross
Abstract: For reliable large-scale quantum computation, quantum error correction (QEC) is essential to protect logical information distributed across m...
By Seong-Joon Park, Hee-Youl Kwak, Yongjune Kim
arXiv:2511.21075v4 Announce Type: replace-cross
Abstract: Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text...
By Zhenchao Tang, Fang Wang, Haohuai He, Jiale Zhou, Tianxu Lv, Jun Zhu, Shouzhi Chen, Minghao Yang, Yu Wang, Jiayang Wu, Yidong Song, Yaokun Li, Jiehui Huang, Jun Zhou, Bing He, Jianhua Yao
arXiv:2610.01936v2 Announce Type: replace
Abstract: Large Language Models (LLMs) have demonstrated remarkable fluency and versatility across natural language tasks but remain fundamentally limited by...
By Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR
arXiv:2603.01295v2 Announce Type: replace-cross
Abstract: Joint lesion segmentation and tissue classification in breast ultrasound are usually trained with a shared encoder, so the two branches stop...
By Abdullah Al Shafi, Md Kawsar Mahmud Khan Zunayed, Safin Ahmmed, Sk Imran Hossain, Engelbert Mephu Nguifo
arXiv:2603.02655v2 Announce Type: replace-cross
Abstract: Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in...
By Anum Afzal, Yuki Saito, Hiroya Takamura, Katsuhito Sudoh, Shinnosuke Takamichi, Graham Neubig, Florian Matthes, Tatsuya Ishigaki
arXiv:2603.21276v2 Announce Type: replace-cross
Abstract: The growing demand for on-device large language model (LLM) services on mobile edge devices has driven the adoption of Mixture-of-Experts (Mo...
By Zihan Fang, Qianru Wang, Haonan An, Zheng Lin, Yiqin Deng, Symeon Chatzinotas, Yuguang Fang
arXiv:2603.23184v2 Announce Type: replace-cross
Abstract: Despite the success of reinforcement learning from human feedback (RLHF), existing reward modeling methods largely rely on explicit feedback,...
By Hao Wang, Haocheng Yang, Licheng Pan, Lei Shen, Xiaoxi Li, Yinuo Wang, Zhichao Chen, Yuan Lu, Haoxuan Li, Zhouchen Lin
arXiv:2604.07925v2 Announce Type: replace-cross
Abstract: The self-attention mechanism is central to the success of Transformer architectures. However, standard row-stochastic attention has been show...
By Michela Lapenna, Rita Fioresi, Bahman Gharesifard
arXiv:2605.18882v2 Announce Type: replace-cross
Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, si...
By Wei Shi, Ziheng Peng, Sihang Li, Xiting Wang, Xiang Wang, Mengnan Du, Na Zou