Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,449 stories · RSS feed

arXiv AI
3d ago

Stateless Language Agents: Scaling Long-Horizon Automated Research

arXiv:2610.07625v1 Announce Type: cross Abstract: Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay...

By Qizheng Zhang, Changxiu Ji, Isaac Sun, Yuetai Li, Shubhangi Upasani, Sherry Ruan, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Yoonho Lee, Yuzhen Mao, Genghan Zhang, Rulin Shao, Qiuyang Mang, Andy Dimnaku, Changran Hu, Radha Poovendran, Kunle Olukotun
arXiv AI
3d ago

Monte Carlo Estimation for KV Cache Eviction

arXiv:2610.07643v1 Announce Type: cross Abstract: Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter whi...

By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Wajih Hassan Raza, Atta Ul Asad, Young D. Kwon, Michal Valko, Dean F. Hougen
arXiv AI
3d ago

CACHEFORGE: LLM-Guided End-to-End Generative Cache Replacement Policy for Performance and Hardware Efficiency

CACHEFORGE introduces a novel framework that uses a large language model (LLM) to evolve cache‑replacement policies end‑to‑end. In each iteration, the LLM generates new C++ replacement logic, which is evaluated by a trace‑based simulator and refined through reward shaping, structural checks, and mutation. The resulting policies are compact, hardware‑aware, and outperform existing CRC‑2 baselines on SPEC CPU2006, achieving significant improvements in hit rate and IPC across diverse workloads.

By Kaushal Mhapsekar, Bita Aslrousta, Brijesh Kumar Bhayana, Paula Contreras, Azam Ghanbari, Ethan Goodman, Anna Andriiko, Samira Mirbagher Ajorpaz
arXiv AI
3d ago

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

VisionWeave introduces elastic visual representation weaving, a native capability for multimodal large language models that learns where and at what granularity to encode visual information. The method combines a gated spatial pooler for coarse representations with a granularity router that allocates content‑adaptive token usage, trained end‑to‑end on large‑scale data. Experiments on Qwen3.5‑4B and Qwen3.8‑27B show that VisionWeave can save 43.0% of tokens while preserving 98.9% of performance across eight benchmarks, and delivers significant throughput gains and latency reductions when deployed on the SGLang serving engine.

By Yuan Feng, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang, Jin Xu, Xike Xie
arXiv AI
3d ago

Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices

The paper introduces a new evaluation setting called penalty‑framed no‑valid‑option MCQA, where multiple‑choice questions may contain no correct answer. By removing the correct option from the MMLU‑Pro mathematics subset and allowing models to either pick an option or abstain, the authors penalize forced‑choice responses that are invalid. Experiments reveal that even models with high standard MCQA accuracy can still produce invalid forced‑choice answers, indicating that traditional accuracy metrics miss an important aspect of model reliability.

By Jinhyeok Kim, Hye-Young Jung