arXiv:2603. 13683v4 Announce Type: replace-cross Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts.
By Hanwen Shen, Ting Ying, Jiajie Lu, Shanshan Wang
arXiv:2605. 05895v2 Announce Type: replace-cross Abstract: Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detection.
By Minsuk Jang, Yujin Yang, Hee-Seon Kim, Minseok Son, Younghun Kim, Changick Kim
arXiv:2607. 29218v1 Announce Type: new Abstract: With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft?
By Jianxin Gao, Beini Hu, Runze Li, Wanli Peng, Ruohan Lei, Jinyuan Zhang, Linna Deng, Tianyi Yu, Zining Wang
arXiv:2607. 28896v1 Announce Type: cross Abstract: Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly.
By Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji, Dinesh Manocha
arXiv:2607. 29601v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA).
By Jiajia Tang, Sizhe Yuen, Francisco Gomez Medina, Yali Du, Adam Sobey
arXiv:2607. 29422v1 Announce Type: cross Abstract: Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report.
By Michael Fu, Qiyue Mei, Patanamon Thongtanunam, Kla Tantithamthavorn
arXiv:2604. 08342v2 Announce Type: replace Abstract: Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains.
By Qiance Tang, Ziqi Wang, Jieyu Lin, Ziyun Li, Barbara De Salvo, Sai Qian Zhang
arXiv:2607. 29621v1 Announce Type: cross Abstract: Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires understanding the temporal and spectral patterns that drive their predictions.
By Antonia Holzapfel, Andres Felipe Posada Moreno, Sebastian Trimpe
arXiv:2509. 25146v2 Announce Type: replace-cross Abstract: This paper develops a mathematical argument and algorithms for building representations of data from event-based cameras, that we call Fast Feature Field ($\text{F}^3$).
By Richeek Das, Kostas Daniilidis, Pratik Chaudhari
arXiv:2607. 29240v1 Announce Type: cross Abstract: In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypical state.
By Kesheng Chen, Yamin Hu, Wenjian Luo
arXiv:2607. 28974v1 Announce Type: cross Abstract: The rapid advancement of image generation models has made it increasingly difficult for people to distinguish AI-generated images from real ones.
By Renxi Cheng, Jie Gui, Hongsong Wang
arXiv:2607. 29177v1 Announce Type: cross Abstract: Utility data (e.
By Rongchao Xu, Lin Jiang, Dahai Yu, Ximiao Li, Guang Wang
arXiv:2607. 29602v1 Announce Type: cross Abstract: Reading a social situation often depends on behavior, not words alone.
By Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D'Avirro, Benjamin Peloquin
Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set of context views and process them jointly. This limits their applicability to online scenarios where calibrated views arrive sequentially and the scene must be updated causally.
Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring. The process is tedious, time-consuming, and error-prone.
In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the spatio-temporal dimension, yet a high compression ratio is often accompanied by the loss of critical details.
Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excessive cost of visual token processing.
Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference. We introduce F-ICL, an in-context-learning benchmark that supplies one exactly.
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by its statistical correlation with human ratings.
Two prompts can request the same code change and produce the same correct patch, yet cause a coding agent to perform radically different kinds and amounts of work. We study this effect in a preregistered benchmark spanning 4,644 valid runs, 24 deterministic coding tasks, seven reasoning models, and two real agent harnesses.