arXiv:2505.14777v2 Announce Type: replace-cross
Abstract: The design of effective optimization algorithms for neural networks remains a fundamental challenge, and most existing methods rely on heuris...
By Mingquan Feng, Yixin Huang, Yifan Fu, Shaobo Wang, Junchi Yan
arXiv:2508.05880v3 Announce Type: replace-cross
Abstract: Understanding human emotions is central to user-facing AI applications, safety alignment, and the simulation of human behavior. As emotional...
By Sree Bhattacharyya, Evgenii Kuriabov, Lucas Craig, Tharun Dilliraj, Reginald B. Adams, Jr., Jia Li, James Z. Wang
arXiv:2508.10020v2 Announce Type: replace-cross
Abstract: Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especiall...
By Chuan Li, Qianyi Zhao, Fengran Mo, Cen Chen
arXiv:2512.11147v2 Announce Type: replace-cross
Abstract: AI agents are increasingly granted autonomous access to sensitive user data and third-party services, making effective permission management...
By Jinhao Zhu, Xiao Huang, Kevin Tseng, Gil Vernik, Shishir G. Patil, Vivian Fang, Raluca Ada Popa
arXiv:2603.04317v2 Announce Type: replace-cross
Abstract: A growing literature shows that variables can be linearly decoded from the activations of large language models (LLMs). These range from prop...
By Elan Barenholtz
arXiv:2603.16017v2 Announce Type: replace-cross
Abstract: Large language models (LLMs) increasingly participate in morally sensitive decision-making, yet how they organize ethical frameworks across r...
By Fan Huang, Haewoon Kwak, Jisun An
arXiv:2605.18850v2 Announce Type: replace-cross
Abstract: We introduce KadiAssistant, a privacy-by-design AI assistant integrated into the Kadi research data ecosystem, enabling researchers to effici...
By Adrian Cierpka, Mohammad Shafiqul Islam, Johannes Steinh\"ulb, Eric Dietriche Sesso Domtchoueng, Michael Selzer, Arnd Koeppe
arXiv:2605.20740v2 Announce Type: replace-cross
Abstract: Large language models (LLMs) have emerged as flexible regressors capable of predicting real-valued quantities from heterogeneous inputs. Yet...
By Jungsoo Park, Hyungjoo Chae, Ethan Mendes, Jay DeYoung, Varsha Kishore, Wei Xu, Alan Ritter
arXiv:2610.07002v1 Announce Type: new
Abstract: Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides fea...
By Yiping Ji, James Martens, Simon Lucey
arXiv:2609.33155v2 Announce Type: replace
Abstract: Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stan...
By Yuyang Zhao, Xuan Liu, HaoYang Shang, Haojian Jin
arXiv:2604.18226v2 Announce Type: replace
Abstract: Automated analysis of customer feedback on social media is hindered by three challenges: the high cost of annotated training data, the scarcity of...
By Pierre-Carl Langlais, Pavel Chizhov, Yannick Detrois, Carlos Rosas-Hinostroza, Ivan P. Yamshchikov, Bastien Perroy
arXiv:2512.23065v4 Announce Type: replace
Abstract: The introduction of BERT established encoder-only transformer models as a foundational paradigm in natural language processing. Encoder-only models...
By Melik\c{s}ah T\"urker, A. Ebrar K{\i}z{\i}lo\u{g}lu, Onur G\"ung\"or, Susan \"Usk\"udarl{\i}
arXiv:2504.18225v2 Announce Type: replace
Abstract: We introduce a new generation of small reasoning models for RAG, search, and source summarization. Pleias-RAG-350m and Pleias-RAG-1B are mid-traine...
By Pierre-Carl Langlais, Pavel Chizhov, Mattia Nee, Carlos Rosas-Hinostroza, Matthieu Delsart, Ir\`ene Girard, Othman Hicheur, Anastasia Stasenko, Ivan P. Yamshchikov
arXiv:2610.08716v1 Announce Type: cross
Abstract: Generative retrieval trains a language model to generate the identifier of a relevant document. Recent work replaces the autoregressive decoder with...
By Hicham Randrianarivo, Logan Renaud, Alexia Allal
arXiv:2610.07522v1 Announce Type: new
Abstract: Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but qu...
By Yan Scholten, Rachel Lawrence, James Hensman, Stephan G\"unnemann, Alicia Curth, Riccardo Grazzi
arXiv:2610.07208v1 Announce Type: new
Abstract: Predicting migration flows remains a significant challenge for traditional gravity-based forecasting models, which primarily rely on structured socio-e...
By Nathaniel T. Hindman, Fabricio Murai
arXiv:2610.07559v1 Announce Type: new
Abstract: Recent progress in tabular foundation models suggests that training on synthetic tasks can substantially improve in-context learning capabilities, with...
By Zijian Li, Xiangchen Song, Gongxu Luo, Jie Qiao, Ruichu Cai, Zhenhao Chen, Xinshuai Dong, Fan Feng, Guangyi Chen, Kun Zhang
arXiv:2610.07767v1 Announce Type: new
Abstract: Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation...
By Xin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao, Zheng Li, Junda Feng, Yuyan Luo, Yi Zhang, Yizhong Cao, Mi Zhang, Dayiheng Liu, Jianwei Zhang
arXiv:2610.07904v1 Announce Type: new
Abstract: We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top o...
By Aksel Fristrup, Sumit Pandey, Ankit Kariryaa
arXiv:2610.07967v1 Announce Type: new
Abstract: As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their r...
By Yiming Xu, Hongyue Yu, Beihua Yang, Zihan Chen, Yixin Liu, Zhen Peng, Bin Shi, Bo Dong, Chao Shen, Irwin King, Qinghua Zheng