Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,035 stories · RSS feed

arXiv AI
2d ago

A Validated Dataset and Benchmark for Coherent Multi-Diagram SysML Models

The paper introduces SEMAADB, a dataset comprising 3,000 engineering contexts and 15,000 SysML diagrams, each context containing five interconnected views (Requirement, Block Definition, Activity, State Machine, and Sequence). The authors verified diagram consistency and created a 100-context human‑verified benchmark. They evaluated three language models on diagram repair and cross‑diagram update tasks, finding that while syntax repair is largely solved, semantic repair and cross‑diagram consistency remain challenging.

By Ardalan Aryashad, Yan Jin
arXiv AI
2d ago

Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation

The paper introduces Physics‑Guided Visual Prompting (PG‑VP), a plug‑and‑play module that overlays a virtual obstacle onto the input of a frozen Vision‑Language‑Action model to guide navigation around invisible hazards such as radiation or temperature spikes. PG‑VP performs a physics‑based risk assessment to determine the avoidance direction and dynamically renders the same virtual obstacle across frames, allowing the existing navigation policy to detour without retraining. Experiments on OmniNav with R2R‑CE and RxR‑CE datasets show that PG‑VP steers the policy toward low‑risk actions in 84.9% and 83.2% of cases, while real‑world tests on a robot demonstrate significant safety improvements against thermal and radiation sources.

By Hojoon Son, Fan Zhang
arXiv AI
2d ago

Stateless Language Agents: Scaling Long-Horizon Automated Research

arXiv:2610.07625v1 Announce Type: cross Abstract: Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay...

By Qizheng Zhang, Changxiu Ji, Isaac Sun, Yuetai Li, Shubhangi Upasani, Sherry Ruan, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Yoonho Lee, Yuzhen Mao, Genghan Zhang, Rulin Shao, Qiuyang Mang, Andy Dimnaku, Changran Hu, Radha Poovendran, Kunle Olukotun
arXiv AI
2d ago

Monte Carlo Estimation for KV Cache Eviction

arXiv:2610.07643v1 Announce Type: cross Abstract: Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter whi...

By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Wajih Hassan Raza, Atta Ul Asad, Young D. Kwon, Michal Valko, Dean F. Hougen
arXiv AI
2d ago

CACHEFORGE: LLM-Guided End-to-End Generative Cache Replacement Policy for Performance and Hardware Efficiency

CACHEFORGE introduces a novel framework that uses a large language model (LLM) to evolve cache‑replacement policies end‑to‑end. In each iteration, the LLM generates new C++ replacement logic, which is evaluated by a trace‑based simulator and refined through reward shaping, structural checks, and mutation. The resulting policies are compact, hardware‑aware, and outperform existing CRC‑2 baselines on SPEC CPU2006, achieving significant improvements in hit rate and IPC across diverse workloads.

By Kaushal Mhapsekar, Bita Aslrousta, Brijesh Kumar Bhayana, Paula Contreras, Azam Ghanbari, Ethan Goodman, Anna Andriiko, Samira Mirbagher Ajorpaz