LiveMathematicianBench is a dynamic multiple‑choice benchmark for research‑level mathematical reasoning, built from recent arXiv papers published after model training cutoffs. It introduces a thirteen‑category logical taxonomy of theorem types and uses a proof‑sketch‑guided distractor pipeline to create plausible but invalid answer choices, enhancing sensitivity to genuine reasoning. Evaluation shows current large language models perform poorly, with the best model scoring 43.5% overall and only 17.6% under substitution‑resistant conditions, indicating the benchmark’s difficulty and realism.
By Linyang He, Qiyao Yu, Hanze Dong, Baohao Liao, Xinxing Xu, Micah Goldblum, Jiang Bian, Nima Mesgarani
arXiv:2605. 28742v2 Announce Type: replace Abstract: Language models can use verifiable rewards to improve at a wide variety of reasoning tasks.
By Linas Nasvytis, Simon Jerome Han, Ben Prystawski, Satchel Grant, Noah D. Goodman, Judith E. Fan
arXiv:2608. 08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge.
By Donghong Jiang, Endian Lin, Luoping Cui, Hanqing Liu, Mingjie Liu, Fan Yang, Hong Wang, Zhao Yang, Chuang Zhu
arXiv:2508. 06165v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have shown strong capabilities through two complementary paradigms: Retrieval-Augmented Generation (RAG) for knowledge grounding and Reinforcement Learning from Verifiable Rewards (RLVR) for complex reasoning.
By Weitao Li, Boran Xiang, Xiaolong Wang, Zhinan Gou, Weizhi Ma, Yang Liu
arXiv:2607. 25182v1 Announce Type: cross Abstract: The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval.
By Adarsh Singh, Kushal Raj Bhandari, Jianxi Gao, Soham Dan, Vivek Gupta
Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit.