arXiv:2509. 16749v1 Announce Type: cross Abstract: LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners.
By Anna Bertiger, Bobby Filar, Aryan Luthra, Stefano Meschiari, Aiden Mitchell, Sam Scholten, Vivek Sharath
arXiv:2608. 11847v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning.
By Beomsik Cho, Jinhyeong Kim, Dongseok Lee, Jaehyung Kim
arXiv:2608. 11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems.
By Touseef Hasan, Mounika Ghanta, Souvika Sarkar, Ujjwal Guin
arXiv:2608. 11947v1 Announce Type: cross Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge.
By Karl Hanna, Chen Feng
arXiv:2608. 11749v1 Announce Type: cross Abstract: Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation.
By Shiji Zhou, Kunlin Lyu, Lei Zhang, Ruodong Wang, Yifan Sun
arXiv:2608. 10010v2 Announce Type: replace Abstract: Low-precision formats usually optimize scalar fidelity while inheriting conventional product arithmetic.
By Ye Qiao
arXiv:2608. 12086v1 Announce Type: cross Abstract: Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence.
By Nikolette Pedersen, Regitze Sydendal, Veronika Cheplygina, Th\'eo Sourget
arXiv:2507. 15100v3 Announce Type: replace-cross Abstract: Natural Language Inference (NLI) determines whether a premise entails, contradicts, or is neutral with respect to a hypothesis.
By Chathuri Jayaweera, Brianna Yanqui, Bonnie J. Dorr
arXiv:2608. 12025v1 Announce Type: cross Abstract: Medical devices are becoming more software-intensive, connected, and AI-enabled.
By Tuhinangshu Gangopadhyay, Rasmus Adler, Peter Liggesmeyer, Jan Reich
arXiv:2608. 11705v1 Announce Type: new Abstract: Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones.
By Lang Cao
arXiv:2608. 11510v1 Announce Type: cross Abstract: Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology and neuroscience, but their mechanistic basis is not fully understood.
By Xiaoyang Hu, Mike Angstadt, Shane Storks, Zan Huang, Aman Taxali, Alex Weigard, Richard L. Lewis, Chandra Sripada
arXiv:2608. 11604v1 Announce Type: new Abstract: Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents.
By Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou
arXiv:2602. 13452v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication.
By Belu Ticona, Antonis Anastasopoulos
arXiv:2512. 06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setting.
By Xiaocan Li, Shiliang Wu, Zheng Shen
arXiv:2608. 12218v1 Announce Type: cross Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories.
By Arda Uzunoglu, Benjamin van Durme, Daniel Khashabi
arXiv:2608. 11788v1 Announce Type: cross Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models.
By Minjun Kim, Inho Won, Hyeonseok Lim, MinKyu Kim, Junghun Yuk, Wooyoung Go, Jongyoul Park, Jungyeul Park, KyungTae Lim
arXiv:2608. 11250v1 Announce Type: new Abstract: Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced.
By Weicheng Ye, Youran Sun, Xingyu Ren, Shunyao Yu, Chugang Yi, Haizhao Yang
arXiv:2606. 00671v3 Announce Type: replace Abstract: We present Moxia (formerly AXIOM), a trust-first neuro-symbolic architecture for self-explaining mathematical reasoning over natural-language input.
By Alessio Bruno
arXiv:2608. 12125v1 Announce Type: cross Abstract: As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes.
By Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer
arXiv:2608. 12249v1 Announce Type: new Abstract: Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of computational science the work simply goes undone.
By Yuzhong Shen, Masha Sosonkina, Peng Xu, Mark S. Gordon