arXiv:2610.08716v1 Announce Type: cross
Abstract: Generative retrieval trains a language model to generate the identifier of a relevant document. Recent work replaces the autoregressive decoder with...
By Hicham Randrianarivo, Logan Renaud, Alexia Allal
arXiv:2610.07522v1 Announce Type: new
Abstract: Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but qu...
By Yan Scholten, Rachel Lawrence, James Hensman, Stephan G\"unnemann, Alicia Curth, Riccardo Grazzi
arXiv:2610.07208v1 Announce Type: new
Abstract: Predicting migration flows remains a significant challenge for traditional gravity-based forecasting models, which primarily rely on structured socio-e...
By Nathaniel T. Hindman, Fabricio Murai
arXiv:2610.07559v1 Announce Type: new
Abstract: Recent progress in tabular foundation models suggests that training on synthetic tasks can substantially improve in-context learning capabilities, with...
By Zijian Li, Xiangchen Song, Gongxu Luo, Jie Qiao, Ruichu Cai, Zhenhao Chen, Xinshuai Dong, Fan Feng, Guangyi Chen, Kun Zhang
arXiv:2610.07767v1 Announce Type: new
Abstract: Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation...
By Xin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao, Zheng Li, Junda Feng, Yuyan Luo, Yi Zhang, Yizhong Cao, Mi Zhang, Dayiheng Liu, Jianwei Zhang
arXiv:2610.07904v1 Announce Type: new
Abstract: We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top o...
By Aksel Fristrup, Sumit Pandey, Ankit Kariryaa
arXiv:2610.07967v1 Announce Type: new
Abstract: As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their r...
By Yiming Xu, Hongyue Yu, Beihua Yang, Zihan Chen, Yixin Liu, Zhen Peng, Bin Shi, Bo Dong, Chao Shen, Irwin King, Qinghua Zheng
arXiv:2610.08075v1 Announce Type: new
Abstract: Conditional neural fields represent signals continuously, but their effectiveness depends on how the conditional latent representations are inferred fr...
By Rudolf L. M. van Herten, Soufiane Ben Haddou, Rachit Saluja, Johannes C. Paetzold
arXiv:2610.08402v1 Announce Type: new
Abstract: Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two r...
By Jiaju Chen, Min Yang, Jinghua Piao, Xiaochong Lan, Xu Xia, Xiangnan He, Yong Li
arXiv:2610.08694v1 Announce Type: new
Abstract: Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature corre...
By Michael Bohl, Alexander Theus, David Wissel, Valentina Boeva
arXiv:2610.07269v1 Announce Type: cross
Abstract: Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles thr...
By Ayesh Abu Lehyeh, Jay Hwasung Jung, Safwan Wshah
arXiv:2610.07572v1 Announce Type: cross
Abstract: In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every q...
By Xi Ding, Naichen Shi, Jiawei Zhang
arXiv:2610.07585v1 Announce Type: cross
Abstract: We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical fe...
By Sheir A. Zaheer, Jihwan Moon, Chan Y. Park
arXiv:2610.07720v1 Announce Type: cross
Abstract: Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existin...
By Wanning He, Yuyao Zhang, Yu-Wing Tai
arXiv:2610.07723v1 Announce Type: cross
Abstract: Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activat...
By Yibo Zhang, Tianrong Guan, Liang Lin, Puze Wang, Jin Wang, Qingsong Wen
arXiv:2610.07825v1 Announce Type: cross
Abstract: Time series forecasting models are typically compared on pointwise error, which scores a prediction in isolation from the decision it is produced for...
By Jonathan Chang, Zimeng Lyu
arXiv:2610.07894v1 Announce Type: cross
Abstract: Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is in...
By Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Borui Jiang, Bo Han
arXiv:2610.08463v1 Announce Type: cross
Abstract: Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an...
By Edan Kinderman, Elad Hoffer, Yochai Blau, Brian Chmiel, Ron Banner, Daniel Soudry, Boris Ginsburg
arXiv:2509.18169v4 Announce Type: replace
Abstract: Tasks on complex systems require high-precision numerical computation to support decisions. However, current large language models (LLMs), even wit...
By Jingyuan Fan, Purui Liu, Hengbo Xiao, Yuxuan Zheng, Jingzhao Zhang, Chao Lu, Guannan He
arXiv:2602.02427v4 Announce Type: replace
Abstract: Large Language Models (LLMs) have achieved significant breakthroughs across various domains, but they can still produce unreliable or misleading ou...
By Qihao Wen, Jiahao Wang, Yang Nan, Pengfei He, Ravi Tandon, Han Xu