arXiv:2506.01732v4 Announce Type: replace
Abstract: Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of...
By Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas-Hinostroza, Mattia Nee, Eliot Krzystof Jones, Ir\`ene Girard, David Mach, Anastasia Stasenko, Ivan P. Yamshchikov
arXiv:2604.20658v2 Announce Type: replace
Abstract: Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-s...
By Shivani Kumar, Adarsh Bharathwaj, David Jurgens
arXiv:2608.27128v4 Announce Type: replace
Abstract: Compressing a key-value cache before its next question is known requires choosing what to retain without knowing which evidence will matter. Value...
By Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Dongnan Zheng, Junyan Zhang, Yubo Gao, Xuming Hu
arXiv:2610.05275v2 Announce Type: replace
Abstract: A large part of political discourse on social media is built and expressed at a level of claims: i.e. declarative, typically single-clause statemen...
By Sandrine Chausson, Bj\"orn Ross
arXiv:2610.06650v2 Announce Type: replace
Abstract: Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right en...
By Mohamed Chenene, Carlos Rosas-Hinostroza, Anastasia Stasenko, Shani Evenstein Sigalov, Pierre-Carl Langlais
arXiv:2607.10162v2 Announce Type: replace-cross
Abstract: Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the ac...
By Yun-Shao Tsai, Chun-Wei Chen, Chee-En Yu, Yi-Cheng Lin, Hung-yi Lee
arXiv:2610.06932v1 Announce Type: new
Abstract: Cache-based test-time adaptation (TTA) for vision-language models is often hindered by background bias in global representations and unreliable entropy...
By Siyu Huang, Yueyong Chen, Xuejiao Li, Jun Zhou
arXiv:2610.06972v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual understanding and multimodal reasoning, yet they remain fu...
By Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert
arXiv:2610.07566v1 Announce Type: new
Abstract: Feed-forward novel view synthesis methods achieve strong generalization from posed multi-view inputs, but scaling them to large input view sets remains...
By JooHyun Park, HanYoung Jang, HyeongYeop Kang
arXiv:2610.07729v1 Announce Type: new
Abstract: Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression b...
By Donghyun Han, Jangho Park, Yuseok Bae
arXiv:2610.07982v1 Announce Type: new
Abstract: Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), f...
By Jinsong Zhang, Kejun Wu, Ming Zhu, Renjie Qiao, Chengtao Cai, Zhengguo Li
arXiv:2610.08068v1 Announce Type: new
Abstract: Realistic physical interaction is a cornerstone of embodied intelligence, yet collecting paired visual--tactile data remains costly. Visual-to-tactile...
By Guo Tang, Yongtao Wang
arXiv:2610.08365v1 Announce Type: new
Abstract: We adapt Stevens's power law to measure the innate ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of...
By Kaichun Yang, Jian Chen
arXiv:2610.08414v1 Announce Type: new
Abstract: Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded...
By Zhen Yu, Wenyang Liu, Kejun Wu, Chengwang Xiao, Renjie Qiao, Chengtao Cai
arXiv:2610.08433v1 Announce Type: new
Abstract: Modern action recognition models rely on video transformers pretrained on massive collections of web-crawled videos, such as Kinetics-700. However, the...
By Ricardo Pizarro, Roberto Valle, Jos\'e M. Buenaposada, Luis M. Bergasa, Luis Baumela
arXiv:2610.08649v1 Announce Type: new
Abstract: Video-language models are ranked by multiple-choice accuracy on frames from a uniform grid. The grid has two parameters, a rate and a phase, and benchm...
By Lichen Zhu, Yiheng Wang, Yueqian Lin, Hai "Helen" Li, Yiran Chen
arXiv:2610.08674v1 Announce Type: new
Abstract: Current large video-language models (LVLMs) still face challenges when dealing with long videos, mainly because frames are often processed independentl...
By Yuhao Qin, Junbo Wang, Yuke Li, Yining Zhu
arXiv:2610.08713v1 Announce Type: new
Abstract: Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene. VL...
By Hairong Yin, Huangying Zhan, Shin-Fang Chng, Yi Xu, Raymond A. Yeh
arXiv:2610.06955v1 Announce Type: cross
Abstract: Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we nat...
By Ruoxuan Feng, Yutong Chen, Ruihua Song, Huan Yang, Zhongyuan Wang, Guocai Yao, Di Hu
arXiv:2510.16320v2 Announce Type: replace
Abstract: This paper presents a systematic study of scaling laws for the deepfake detection task. Specifically, we analyze the model performance against the...
By Wenhao Wang, Jusheng Zhang, Longqi Cai, Taihong Xiao, Yuxiao Wang, Ming-Hsuan Yang