Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,449 stories · RSS feed

arXiv Computer Vision
3d ago

World Models' Last Exam in Physics

arXiv:2610.08791v1 Announce Type: new Abstract: Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and p...

By Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin, Ziming Qin, Zheng Jiang, Wenyi Li, Calvin Xiao, Youjie Zheng, Kaisen Yang, Qinhuai Na
arXiv Computer Vision
3d ago

OpenWAM: An Open Framework for Composable World-Action Models

arXiv:2610.07922v1 Announce Type: cross Abstract: World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, su...

By Heng Yu, David D. Yuan, Juze Zhang, Changan Chen, Yao Feng, Michelle Baldonado, Steve Cousins, Li Fei-Fei, Jiajun Wu, Ehsan Adeli
arXiv Machine Learning
3d ago

Extending Music Annotation Schemas: Zero-Shot Prediction or Few-Shot Adaptation?

The paper investigates how to extend music annotation schemas when new attributes need to be added to commercial music catalogs. It compares zero‑shot prediction using audio‑language models, learning from pretrained representations, and supervised adaptation on existing annotations, using a benchmark built on the MGPHot dataset. The findings show that supervised adaptation outperforms zero‑shot prediction even with limited annotation budgets, while reusing frozen representations remains the best option for very modest budgets.

By Christos Plachouras, Emmanouil Benetos, Johan Pauwels
arXiv Machine Learning
3d ago

SkillFormer: Skill-Decomposed Adaptation for Audio Language Models

SkillFormer is a method for audio language models that decomposes audio understanding into skill‑specific low‑rank adapters and uses a learned router to activate the appropriate adapters at inference time. The router selects which adapters to engage based on the question, allowing different parameters to be used for tasks such as pitch comparison versus genre classification. An alternating training schedule updates each adapter on its own skill cluster before jointly calibrating the router, reducing gradient conflicts and adding fewer than 4% of the base model’s parameters. The approach improves average accuracy by 2.5 to 4.1 points across three distinct models on MMSU, MMAU‑Pro, and MMAR, achieving balanced gains across perception, reasoning, and semantic subcategories.

By Lee Seung-woo, Bowen Qi, Kim Min-jun, Jang Won-young
arXiv AI
3d ago

Principles that Guide, Actions that Inform: Agent Evolution via Knowledge Abstraction

The paper introduces SAGA, a framework that enables large language model agents to evolve by abstracting experiences into reusable principles, procedures, and episodic descriptions. SAGA transforms interaction trajectories into hierarchical knowledge with explicit applicability conditions, linking them back to execution evidence. Experiments on ScienceWorld and ALFWorld show that this execution–abstraction feedback loop improves task performance, and ablation studies confirm the importance of contextual instantiation and action regulation.

By Bowen Ye, Yongchao Xu, Junkai Ma, Xiang Yin, Wenzhao Li
arXiv AI
3d ago

Small Language Models for Smart Data Model Classification at the Edge: A Cost-Aware Hybrid Approach

The paper introduces a cost‑aware hybrid approach that uses small, open‑source language models to classify smart data models (SDMs) in Internet of Things (IoT) environments. It benchmarks general‑purpose, reasoning‑specialized, and code‑specialized models across domain‑specific datasets, highlighting their suitability for edge devices with limited resources. The study also compares these lightweight models to large language models and near‑zero‑cost baselines such as TF‑IDF and a lightweight sentence encoder to demonstrate practical performance gains.

By Cristian Martella, Angelo Martella, Antonella Longo, Motaz Saad
arXiv AI
3d ago

Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot

The paper presents a vision‑language navigation system that transfers from simulation to a real Ackermann‑steered mobile robot without relying on navigation graphs or panoramic views. It uses a Cross‑Modal Attention architecture trained on simulated data and fine‑tuned with limited real‑world episodes, leveraging linear photometric adjustments and a camera‑LiDAR sensor suite. Evaluation with SPL and nDTW metrics shows robust, adaptable navigation in continuous environments.

By Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara
arXiv AI
3d ago

MemMux: Runtime Verification and Honest Resource Attribution for Fleets of Parallel Coding Agents

MemMux is a local runtime designed to provide runtime verification and honest resource attribution for fleets of parallel coding agents. It emits observable signals that track per‑agent memory usage, ensure complete reclamation of terminated agents, detect escaped child processes, and keep the system from exceeding a bounded memory footprint. In benchmarks against tmux and a raw‑process baseline, MemMux keeps a fleet under a 7.5 GiB budget with zero swap, while ungoverned tools exceed the budget and spill into swap, and it achieves 100 % attribution with low overhead.

By Sumanyu Muku
arXiv AI
3d ago

Understanding and Mitigating Inference-Time Overreliance Using Agentic Memory

The paper investigates how large language model agents can over-rely on agentic memory, a phenomenon where retrieved memories distort inference even when they are correctly stored and retrieved. It shows that memory is helpful when past experience fully transfers to the current task but becomes misleading under partial query-memory overlap, a pattern confirmed by controlled experiments. To address this, the authors propose MEMTRIM, a plug‑and‑play framework that indexes memory evidence at write time and limits its reuse at read time, thereby reducing over-reliance without retraining and preserving useful memory benefits across models and memory architectures.

By Luoxi Tang, Yuqiao Meng, Nilesh Auradkar, Muchao Ye, Dazheng Zhang, Zhaohan Xi