arXiv:2610.04687v2 Announce Type: replace
Abstract: Answering questions over imperfect tables requires handling errors that can affect the answer. We investigate two challenges for large language mod...
By Baowen Zhang, Wei Fan, Ruman Wang, Hangting Ye
arXiv:2610.06625v2 Announce Type: replace
Abstract: Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe marke...
By Steven Denney, Matthew DiGiuseppe
arXiv:2610.06729v2 Announce Type: replace
Abstract: Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative...
By Zahra Solati Dehkordi, Vasileios Lampos
arXiv:2610.06896v1 Announce Type: new
Abstract: Frontier multimodal large language models (MLLMs) are increasingly positioned as general purpose visual reasoners as part of the quest for artificial g...
By Ross Callaghan, Niannu Gao, Hojjat Azadbakht, Hui Zhang
arXiv:2610.06945v1 Announce Type: new
Abstract: Mechanistic interpretability often relies on the Linear Representation Hypothesis (LRH), which assumes that high-level concepts are encoded as linear d...
By Muhammad Atif Butt, Pawe{\l} Skier\'s, Joost Van De Weijer, Kamil Deja
arXiv:2610.07031v1 Announce Type: new
Abstract: Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics a...
By Sitian Shen, Jiuming Liu, Mengmeng Liu, Yian Wang, Michael Ying Yang, Francesco Nex, Hao Cheng, Daniele De Martini, Ayush Tewari, Per Ola Kristensson
arXiv:2610.07381v1 Announce Type: new
Abstract: Modeling 3D scene geometry and its evolution over time is essential for autonomous driving and robotics. A common paradigm is to use world models to pr...
By Mehrdad Noori, Guile Wu, Sam Hosseini, Dongfeng Bai
arXiv:2610.07689v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense p...
By Juntong Li, Lingwei Dang, Haomin Wu, Ziyan Qiu, Qingxin Xiao, Qingyao Wu
arXiv:2610.07868v1 Announce Type: new
Abstract: Large language model agents have been used to search over symbolic structures such as programs and equations. We propose CueRator, an agentic framework...
By Sunchan Park, Beomkwon Cho, Kyeongbo Kong
arXiv:2610.07903v1 Announce Type: new
Abstract: Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing...
By Keliang Chen, Yaxin Hou, Hui Liu, Yuheng Jia
arXiv:2610.07925v1 Announce Type: new
Abstract: Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress,...
By Giyeol Kim, Chanho Eom
arXiv:2610.08109v1 Announce Type: new
Abstract: Cardiac magnetic resonance imaging reconstruction aims to recover high-quality images from undersampled acquisitions, enabling faster scans while prese...
By Anam Hashmi, Mayug Maniparambil, Julia Dietlmeier, Kathleen M. Curran, Noel E. O'Connor
arXiv:2610.08539v1 Announce Type: new
Abstract: Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis. Existing approaches mainly follow three paradig...
By Dongchen Si, Di Wang, Mingzhen Xu, Jing Zhang, Bo Du, Liangpei Zhang
arXiv:2610.08777v1 Announce Type: new
Abstract: Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise...
By Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh, Xinyu Zhang
arXiv:2610.08791v1 Announce Type: new
Abstract: Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and p...
By Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin, Ziming Qin, Zheng Jiang, Wenyi Li, Calvin Xiao, Youjie Zheng, Kaisen Yang, Qinhuai Na
arXiv:2610.07922v1 Announce Type: cross
Abstract: World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, su...
By Heng Yu, David D. Yuan, Juze Zhang, Changan Chen, Yao Feng, Michelle Baldonado, Steve Cousins, Li Fei-Fei, Jiajun Wu, Ehsan Adeli
arXiv:2602.14512v3 Announce Type: replace
Abstract: Autoregressive pretraining has been key to the scalability of large language models, yet medical generative foundation models remain predominantly...
By Zhicheng He, Yunpeng Zhao, Junde Wu, Ziwei Niu, Ziyue Wang, Bohan Li, Zijun Li, Lanfen Lin, Nan Liu, Yueming Jin
arXiv:2606.15158v3 Announce Type: replace
Abstract: Reference-guided generation (e.g., object compositing, customization) has progressed rapidly, yet current pipelines share a fundamental limitation:...
By Jeahun Sung, Dahyeon Kye, Soo Ye Kim, Jihyong Oh
arXiv:2606.03715v3 Announce Type: replace
Abstract: Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that co...
By Nurit Spingarn, Noa Cohen, Tamar Rott Shaham, Tomer Michaeli
arXiv:2610.07231v1 Announce Type: cross
Abstract: This work presents a novel learning-based pipeline for pose estimation of unknown spacecraft using only monocular images from a single servicer. The...
By Pol Francesch Huc, Simone D'Amico