arXiv:2606. 29493v1 Announce Type: new Abstract: Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-checked proof.
By Pawan Sasanka Ammanamanchi, Siddharth Bhat, Stella Biderman
arXiv:2606. 27944v1 Announce Type: cross Abstract: Phone-use Agents can execute complex tasks end to end across real mobile applications.
By Yiming Sun, Chen Chen, Zifan Zhou, Mi Zhang
arXiv:2606. 30116v1 Announce Type: new Abstract: Pairwise preference data is widely used for training and evaluating language models (e.
By Eleanor Clifford, Michael Amir, Arduin Findeis, Aaron Zhao, Robert Mullins
arXiv:2605. 24844v2 Announce Type: replace Abstract: While general-purpose Large Language Models (LLMs) applied to Geology often hallucinate when reasoning about subsurface structures and deep-time evolution, current AI in Earth sciences predominantly targets surface remote sensing and GIS.
By Chenyou Guo, Zongqi Liu, Yizhou Zhang, Zhaorui Jiang, Ze Liu
arXiv:2606. 28360v1 Announce Type: cross Abstract: University students often struggle to navigate complex academic policies, leading to advising bottlenecks and delayed access to critical information.
By Ben Torsion, Jun Zhou
arXiv:2606. 28370v1 Announce Type: cross Abstract: Enterprise business intelligence queries span structured warehouses and unstructured document repositories -- modalities with fundamentally different access methods, cost profiles, and correctness semantics.
By Darshita Rathore, Vineet Kumar, Vaibhav Singal, Ankur Vivek Singh, Anindya Moitra
arXiv:2606. 28376v1 Announce Type: cross Abstract: Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and existing memory methods either truncate aggressively and lose non-local evidence or retain boilerplate that degrades decision quality.
By Mellow Baixuan Chen, Xiangguo Sun
arXiv:2606. 28358v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) aims to enhance the trustworthiness of Large Language Models (LLMs) by grounding their outputs in external documents, often using inline citations for verifiability.
By Ian van Dort (University of Amsterdam), Maria Heuss (University of Amsterdam)
arXiv:2605. 09038v3 Announce Type: replace Abstract: Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries.
By Jinchao Hu, Meizhi Zhong, Kehai Chen, Min Zhang
arXiv:2602. 18452v3 Announce Type: replace-cross Abstract: As conversational multimodal AI tools are increasingly adopted to process patient data for health assessment, robust benchmarks are needed to measure progress and expose failure modes under realistic conditions.
By Gaia A. Bertolino, Yuwei Zhang, Tong Xia, Domenico Talia, Cecilia Mascolo
arXiv:2606. 30113v1 Announce Type: cross Abstract: Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actions from discrete codes remains challenging.
By Tengyue Jiang, Chunpu Xu, Jiayue Kang, Yao Mu
arXiv:2606. 28925v1 Announce Type: cross Abstract: Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agents, while over-selection increases execution cost.
By Ananto Nayan Bala, Faisal Muhammad Shah
arXiv:2606. 29090v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become the standard way to ground large language models in external knowledge, yet most systems retrieve a fixed number of passages for every question regardless of its difficulty.
By Ansh Kamthan
arXiv:2606. 28419v1 Announce Type: cross Abstract: Limited data availability, class imbalance, and domain variability remain major barriers to reliable medical image classification.
By Teerath Kumar, Raja Vavekanand, Muhammad Turab
arXiv:2606. 28409v1 Announce Type: cross Abstract: Software-compilable C programs routinely fail to complete the four-stage pipeline of a high-level synthesis (HLS) toolchain -- compilation, C simulation (CSim), synthesis, and C/RTL co-simulation (CoSim) -- because HLS accepts only a synthesizable subset of C (HLS-C).
By Zhe Zhao, Hongbing Lang, Zhihan Xiao, Luke Ztz Hu, John Imoleayo Adebisi, Songping Mai
arXiv:2606. 29117v1 Announce Type: cross Abstract: Post-hurricane damage assessment and repair scheduling can require computationally intensive simulation and optimization.
By Hooman Torkaman, Ellis Oti Boateng, Jignesh Solanki, Anurag Srivastava
arXiv:2606. 29503v1 Announce Type: cross Abstract: The verbose context problem occurs when structured concepts have token-inefficient textual representations.
By Shiva Kaul, Min-Gyu Kim, Anjum Khurshid, Sriram Vishwanath
arXiv:2606. 29776v1 Announce Type: cross Abstract: Nuclear Magnetic Resonance (NMR) spectroscopy is the gold standard for molecular structure elucidation, yet interpreting complex spectra for unknown molecules remains a bottleneck reliant on human expertise.
By Zheng Fang, Chen Yang, Yusen Tan, Yunpeng Zhao, Fanjie Xu, Hongxin Xiang, Hanyu Sun, Hanyu Gao, Xiaojian Wang, Wenjie Du, Yuqiang Li, Jun Xia
arXiv:2606. 29824v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel as static solvers, transforming them into autonomous agents remains challenging.
By Chengfeng Zhao, Yuqiao Tan, Shizhu He, Yequan Wang, Jun Zhao, Kang Liu
arXiv:2407. 19633v4 Announce Type: replace Abstract: Optimization problems are pervasive in sectors from manufacturing and distribution to healthcare.
By Ali AhmadiTeshnizi, Wenzhi Gao, Herman Brunborg, Shayan Talaei, Connor Lawless, Madeleine Udell