arXiv:2608. 13221v1 Announce Type: new Abstract: The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search.
By Shunwen Bai, Ziping Ma, Chaoyang Zhang, Yarong Wang, Jiale Liu, Zhen Qin, Qingpei Guo
arXiv:2608. 12477v1 Announce Type: new Abstract: Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient.
By Xiaobin Shen, Chloe Y. H. Huang, Jonathan Elmer, George H. Chen
arXiv:2604. 17299v3 Announce Type: replace-cross Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusing harmful ones.
By Tiankai Yang, Yi Nian, Xinyuan Li, Ruiyao Xu, Henry Peng Zou, Kaize Ding, Xiyang Hu, Yan Liu, Yue Zhao
arXiv:2608. 13345v1 Announce Type: new Abstract: Artificial Intelligence (AI) safety systems combine character shaping (e.
By Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu, Ryuji Hamamoto
arXiv:2412. 17228v4 Announce Type: replace Abstract: Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic trials.
By Jennifer Altreuter, Pavel Trukhanov, Morgan A. Paul, Michael J. Hassett, Irbaz B. Riaz, Muhammad Umar Afzal, Arshad A. Mohammed, Ayub Umair, Huan He, Chueh Husan Hsu, Sarah Sammons, James Lindsay, Emily Mallaber, Harry R. Klein, Gufran Gungor, Matthew Galvin, Michael Deletto, Sabrina Y. Camp, Stephen C. Van Nostrand, James Provencher, Joyce Yu, Naeem Tahir, Jonathan Wischhusen, Olga Kozyreva, Taylor Ortiz, Hande Tuncer, Jad El Masri, Alys Malcolm, Tali Mazor, Ethan Cerami, Kenneth L. Kehl
arXiv:2508. 14390v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often express verbal confidence that is poorly aligned with actual correctness, limiting their reliability in safety-critical applications.
By Ke Fang, Tianyi Zhao, Qianwen Wang, Lu Cheng
arXiv:2608. 13463v1 Announce Type: cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels.
By Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, Linda Ungerboeck
arXiv:2608. 12874v1 Announce Type: new Abstract: Plasticity loss has emerged as a critical challenge in continual learning that significantly hinders the acquisition of sequential tasks.
By Zeyang Zhang, Tieliang Gong, Junyan Lu, Weizhan Zhang
arXiv:2608. 12355v1 Announce Type: cross Abstract: Recent progress in AI coding agent research has led to rapid improvements in agents' ability to autonomously perform complex software engineering tasks, from editing large codebases to executing long-horizon development workflows.
By Zora Z. Wang, John Yang, Kilian Lieret, Alexa Tartaglini, Valerie Chen, Yuxiang Wei, Zijian Wang, Lingming Zhang, Karthik Narasimhan, Ludwig Schmidt, Graham Neubig, Daniel Fried, Diyi Yang
arXiv:2608. 13100v1 Announce Type: new Abstract: Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character recognition, and automated scraping.
By Gupta Lovi Raj, Kaur Kamalpreet, Dama Sri Ram, Parani Prajithaa
arXiv:2608. 13484v1 Announce Type: cross Abstract: When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims.
By Dananjay Srinivas, Saksham Khatwani, Maria Pacheco
arXiv:2608. 12779v1 Announce Type: cross Abstract: Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment.
By Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang
arXiv:2510. 05678v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have achieved notable progress in multilingual settings, their performance remains uneven across languages as LLMs often rely on English-centric latent representations.
By Haneul Yoo, Jiho Jin, Kyunghyun Cho, Alice Oh
arXiv:2608. 13453v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and performing a wide range of manipulation tasks.
By Yukun Dai, Mingzhe Dai, Tianshi Wang, Fengling Li, Jingjing Li, Lei Zhu
arXiv:2602. 18733v2 Announce Type: replace Abstract: Training data leakage from Large Language Models (LLMs) raises serious concerns related to privacy, security, and copyright compliance.
By Trishita Tiwari, Ari Trachtenberg, G. Edward Suh
arXiv:2608. 13495v1 Announce Type: cross Abstract: Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis.
By Yi-Chung Chen, Philip Jacobson, Tom Lampo, Yiren Lu, Jin Yao, David I. Inouye, Jing Gao, Danhua Guo, Burhan Yaman
arXiv:2603. 27803v2 Announce Type: replace Abstract: We provide a distributed online algorithm for multi-agent submodular maximization under communication delays.
By Zirui Xu, Vasileios Tzoumas
arXiv:2608. 13118v1 Announce Type: new Abstract: Verification of neural networks against relational specifications, such as global robustness, is crucial for safety-critical applications of cyber-physical systems (CPS), given their increasing adoption of AI components.
By Kota Fukuda, Zhenya Zhang, Guanqin Zhang, Jianjun Zhao
arXiv:2605. 20088v2 Announce Type: replace-cross Abstract: Discovering shapelets -- i.
By Seongjun Lee, Seokhyun Lee, Changhee Lee
arXiv:2608. 13328v1 Announce Type: cross Abstract: Professional communication is increasingly mediated by LLMs - but do these models serve all users equally?
By Katherine Van Koevering, Anjalie Field