Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,737 stories · RSS feed

arXiv AI
1d ago

MemMux: Runtime Verification and Honest Resource Attribution for Fleets of Parallel Coding Agents

MemMux is a local runtime designed to provide runtime verification and honest resource attribution for fleets of parallel coding agents. It emits observable signals that track per‑agent memory usage, ensure complete reclamation of terminated agents, detect escaped child processes, and keep the system from exceeding a bounded memory footprint. In benchmarks against tmux and a raw‑process baseline, MemMux keeps a fleet under a 7.5 GiB budget with zero swap, while ungoverned tools exceed the budget and spill into swap, and it achieves 100 % attribution with low overhead.

By Sumanyu Muku
arXiv AI
1d ago

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

Rationale-Guided Policy Optimization (RGPO) is a reinforcement‑learning framework that adaptively uses ground‑truth rationale information to scaffold a language model’s reasoning process. Instead of treating reference solutions as fixed imitation targets, RGPO temporarily incorporates rationales to help the model generate better responses, then reverts to unguided learning with higher‑reward, model‑generated solutions. Experiments in both language‑only and vision‑language tasks show that RGPO consistently outperforms RLVR baselines, with ablation studies confirming that adaptive rationale guidance is a key factor in its success.

By Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei
arXiv Computer Vision
1d ago

Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update

arXiv:2610.07031v1 Announce Type: new Abstract: Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics a...

By Sitian Shen, Jiuming Liu, Mengmeng Liu, Yian Wang, Michael Ying Yang, Francesco Nex, Hao Cheng, Daniele De Martini, Ayush Tewari, Per Ola Kristensson
arXiv Computer Vision
1d ago

World Models' Last Exam in Physics

arXiv:2610.08791v1 Announce Type: new Abstract: Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and p...

By Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin, Ziming Qin, Zheng Jiang, Wenyi Li, Calvin Xiao, Youjie Zheng, Kaisen Yang, Qinhuai Na
arXiv Computer Vision
1d ago

OpenWAM: An Open Framework for Composable World-Action Models

arXiv:2610.07922v1 Announce Type: cross Abstract: World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, su...

By Heng Yu, David D. Yuan, Juze Zhang, Changan Chen, Yao Feng, Michelle Baldonado, Steve Cousins, Li Fei-Fei, Jiajun Wu, Ehsan Adeli