arXiv:2610.07689v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense p...
By Juntong Li, Lingwei Dang, Haomin Wu, Ziyan Qiu, Qingxin Xiao, Qingyao Wu
arXiv:2610.07868v1 Announce Type: new
Abstract: Large language model agents have been used to search over symbolic structures such as programs and equations. We propose CueRator, an agentic framework...
By Sunchan Park, Beomkwon Cho, Kyeongbo Kong
arXiv:2610.07903v1 Announce Type: new
Abstract: Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing...
By Keliang Chen, Yaxin Hou, Hui Liu, Yuheng Jia
arXiv:2610.07925v1 Announce Type: new
Abstract: Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress,...
By Giyeol Kim, Chanho Eom
arXiv:2610.08109v1 Announce Type: new
Abstract: Cardiac magnetic resonance imaging reconstruction aims to recover high-quality images from undersampled acquisitions, enabling faster scans while prese...
By Anam Hashmi, Mayug Maniparambil, Julia Dietlmeier, Kathleen M. Curran, Noel E. O'Connor
arXiv:2610.08539v1 Announce Type: new
Abstract: Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis. Existing approaches mainly follow three paradig...
By Dongchen Si, Di Wang, Mingzhen Xu, Jing Zhang, Bo Du, Liangpei Zhang
arXiv:2610.08777v1 Announce Type: new
Abstract: Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise...
By Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh, Xinyu Zhang
arXiv:2610.08791v1 Announce Type: new
Abstract: Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and p...
By Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin, Ziming Qin, Zheng Jiang, Wenyi Li, Calvin Xiao, Youjie Zheng, Kaisen Yang, Qinhuai Na
arXiv:2610.07922v1 Announce Type: cross
Abstract: World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, su...
By Heng Yu, David D. Yuan, Juze Zhang, Changan Chen, Yao Feng, Michelle Baldonado, Steve Cousins, Li Fei-Fei, Jiajun Wu, Ehsan Adeli
arXiv:2602.14512v3 Announce Type: replace
Abstract: Autoregressive pretraining has been key to the scalability of large language models, yet medical generative foundation models remain predominantly...
By Zhicheng He, Yunpeng Zhao, Junde Wu, Ziwei Niu, Ziyue Wang, Bohan Li, Zijun Li, Lanfen Lin, Nan Liu, Yueming Jin
arXiv:2606.15158v3 Announce Type: replace
Abstract: Reference-guided generation (e.g., object compositing, customization) has progressed rapidly, yet current pipelines share a fundamental limitation:...
By Jeahun Sung, Dahyeon Kye, Soo Ye Kim, Jihyong Oh
arXiv:2606.03715v3 Announce Type: replace
Abstract: Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that co...
By Nurit Spingarn, Noa Cohen, Tamar Rott Shaham, Tomer Michaeli
arXiv:2610.07231v1 Announce Type: cross
Abstract: This work presents a novel learning-based pipeline for pose estimation of unknown spacecraft using only monocular images from a single servicer. The...
By Pol Francesch Huc, Simone D'Amico
The paper investigates how to extend music annotation schemas when new attributes need to be added to commercial music catalogs. It compares zero‑shot prediction using audio‑language models, learning from pretrained representations, and supervised adaptation on existing annotations, using a benchmark built on the MGPHot dataset. The findings show that supervised adaptation outperforms zero‑shot prediction even with limited annotation budgets, while reusing frozen representations remains the best option for very modest budgets.
By Christos Plachouras, Emmanouil Benetos, Johan Pauwels
SkillFormer is a method for audio language models that decomposes audio understanding into skill‑specific low‑rank adapters and uses a learned router to activate the appropriate adapters at inference time. The router selects which adapters to engage based on the question, allowing different parameters to be used for tasks such as pitch comparison versus genre classification. An alternating training schedule updates each adapter on its own skill cluster before jointly calibrating the router, reducing gradient conflicts and adding fewer than 4% of the base model’s parameters.
The approach improves average accuracy by 2.5 to 4.1 points across three distinct models on MMSU, MMAU‑Pro, and MMAR, achieving balanced gains across perception, reasoning, and semantic subcategories.
By Lee Seung-woo, Bowen Qi, Kim Min-jun, Jang Won-young
The paper introduces SAGA, a framework that enables large language model agents to evolve by abstracting experiences into reusable principles, procedures, and episodic descriptions. SAGA transforms interaction trajectories into hierarchical knowledge with explicit applicability conditions, linking them back to execution evidence. Experiments on ScienceWorld and ALFWorld show that this execution–abstraction feedback loop improves task performance, and ablation studies confirm the importance of contextual instantiation and action regulation.
By Bowen Ye, Yongchao Xu, Junkai Ma, Xiang Yin, Wenzhao Li
The paper introduces a cost‑aware hybrid approach that uses small, open‑source language models to classify smart data models (SDMs) in Internet of Things (IoT) environments. It benchmarks general‑purpose, reasoning‑specialized, and code‑specialized models across domain‑specific datasets, highlighting their suitability for edge devices with limited resources. The study also compares these lightweight models to large language models and near‑zero‑cost baselines such as TF‑IDF and a lightweight sentence encoder to demonstrate practical performance gains.
By Cristian Martella, Angelo Martella, Antonella Longo, Motaz Saad
The paper presents a vision‑language navigation system that transfers from simulation to a real Ackermann‑steered mobile robot without relying on navigation graphs or panoramic views. It uses a Cross‑Modal Attention architecture trained on simulated data and fine‑tuned with limited real‑world episodes, leveraging linear photometric adjustments and a camera‑LiDAR sensor suite. Evaluation with SPL and nDTW metrics shows robust, adaptable navigation in continuous environments.
By Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara
MemMux is a local runtime designed to provide runtime verification and honest resource attribution for fleets of parallel coding agents. It emits observable signals that track per‑agent memory usage, ensure complete reclamation of terminated agents, detect escaped child processes, and keep the system from exceeding a bounded memory footprint. In benchmarks against tmux and a raw‑process baseline, MemMux keeps a fleet under a 7.5 GiB budget with zero swap, while ungoverned tools exceed the budget and spill into swap, and it achieves 100 % attribution with low overhead.
By Sumanyu Muku
The paper investigates how large language model agents can over-rely on agentic memory, a phenomenon where retrieved memories distort inference even when they are correctly stored and retrieved. It shows that memory is helpful when past experience fully transfers to the current task but becomes misleading under partial query-memory overlap, a pattern confirmed by controlled experiments. To address this, the authors propose MEMTRIM, a plug‑and‑play framework that indexes memory evidence at write time and limits its reuse at read time, thereby reducing over-reliance without retraining and preserving useful memory benefits across models and memory architectures.
By Luoxi Tang, Yuqiao Meng, Nilesh Auradkar, Muchao Ye, Dazheng Zhang, Zhaohan Xi