arXiv:2605.30557v2 Announce Type: replace-cross
Abstract: Spatial reasoning benchmarks typically evaluate whether vision-language models can derive the correct answer from a visual observation. Yet i...
By Yue Zhang, Zun Wang, Han Lin, Yonatan Bitton, Idan Szpektor, Mohit Bansal
arXiv:2606.16193v2 Announce Type: replace-cross
Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representat...
By Yusong Zhao, Hengyi Wang, Tanuja Ganu, Akshay Nambi, Hao Wang
arXiv:2609.15039v3 Announce Type: replace-cross
Abstract: User prompts provided to large language models (LLMs) may contain private information. One way to protect them is to execute the LLM inside a...
By Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar
arXiv:2610.04753v2 Announce Type: replace-cross
Abstract: Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Spa...
By Noam Elata, Itay Lamprecht, Mikey Shechter, Daniel Ohayon, Itay Hubara, Daniel Soudry
arXiv:2601.16390v2 Announce Type: replace-cross
Abstract: Large language models exhibit strong multilingual capabilities, yet significant performance gaps persist between dominant and non-dominant la...
By Rhitabrat Pokharel, Ameeta Agrawal, Tanay Nagar
arXiv:2610.07168v1 Announce Type: new
Abstract: Representations of translated sentences are similar in the inner layers of multilingual language models -- an observation connected to the platonic rep...
By Darshil Doshi, Wenjie Zhou, Corinna Elena Wegner, Daniel J. Korchinski, Santiago Acevedo, Matthieu Wyart
arXiv:2610.07121v1 Announce Type: new
Abstract: Diffusion large language models dLLMs) have emerged as a promising alternative to autoregressive language models through bidirectional diffusion-based...
By Donghyun Lee, Arkapravo Ghosh, Varun Manjunath, Bumjoon Kyle Rhee, Hyunho Kook, Shiting Xiao, Youngeun Kim, Priyadarshini Panda
arXiv:2610.07212v1 Announce Type: new
Abstract: Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model ca...
By Jiachen Zhao, Antonia Januszewicz, Taeho Jung
arXiv:2610.07332v1 Announce Type: new
Abstract: Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structur...
By Bolian Li, Ting-Yao Hu, Cheng-Yu Hsieh, Sanjoy Chowdhury, Oncel Tuzel, Raviteja Vemulapalli
arXiv:2610.07334v1 Announce Type: new
Abstract: Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring kn...
By Krishna Kabra, Constantin Venhoff, Christian Schroeder de Witt
arXiv:2610.07754v1 Announce Type: new
Abstract: Adversarial training is one of the most reliable defenses against adversarial attacks, but its high computational cost must generally be paid anew for...
By Soichiro Kumano
arXiv:2610.07804v1 Announce Type: new
Abstract: Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks. A natural explanation...
By Martin Eppert, Krishna Balasubramanian, Subhro Ghosh, Jason Klusowski, Yan Shuo Tan
arXiv:2610.07810v1 Announce Type: new
Abstract: Search filters help guests navigate vast catalogs in two-sided marketplaces like Airbnb, and recommending the right filters can meaningfully lift booki...
By Shashank Dabriwal, Tanya Piplani, Hao Li, Yiwei Wang, Ashish Jain, Kedar Bellare, Stephanie Moyerman
arXiv:2610.07819v1 Announce Type: new
Abstract: Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, find...
By Shih-Cheng Huang, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee, Shao-Hua Sun
arXiv:2610.08108v1 Announce Type: new
Abstract: Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders...
By Yiming Qin, Ke Wang, Amel Abdelraheem, Adam Hazimeh, Pascal Frossard
arXiv:2610.08129v1 Announce Type: new
Abstract: Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlle...
By Yuhwan Jeong, Jinnyeong Yang, Kuk-Jin Yoon
arXiv:2610.08164v1 Announce Type: new
Abstract: Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ ada...
By Seobin Song, Geonho Lee, Janghwan Lee, Jungwook Choi
arXiv:2610.08403v1 Announce Type: new
Abstract: Large Language Models (LLMs) require substantial computational resources, limiting their deployment on resource-constrained hardware. Ternary LLMs miti...
By Adeline Pittet, Shien Zhu, Val\'erie Verdan, Gustavo Alonso
arXiv:2610.08534v1 Announce Type: new
Abstract: Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a prec...
By Bing Liu, Wenjie Zhou, Chengcheng Zhao, Hongtao Zhang, Boao Kong, Felix Dangel, Wu Lin
arXiv:2610.08578v1 Announce Type: new
Abstract: Transformers provide a state-of-the-art modeling framework, yet poor calibration limits their reliability in safety-critical applications. A promising...
By Amir Mohammad Mahfoozi, Zi Yang, Ying Li, Michael Minyi Zhang