Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,449 stories · RSS feed

arXiv AI
3d ago

PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence

PlaySuite is a large-scale benchmark that evaluates interactive visual intelligence by using over 5,000 open-source video games from platforms like PyWeek and itch.io. The benchmark covers diverse game engines (Pygame, HTML5, Godot, Unity) and introduces a unified closed-loop interaction framework and a Video-LLM-as-a-judge protocol to standardize progress measurement. Evaluation of fourteen recent models shows a perception-action gap, with strong reasoning but poor sustained progress, spatial grounding, action execution, and self-correction.

By Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini, Michelle Lorena Acevedo Callejas, Mohammad Mahdi Derakhshani, Kristof Meding, Joaquin Vanschoren, Cees G. M. Snoek
arXiv AI
3d ago

Agentic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs

The paper introduces TraceDSE, an agentic design space exploration framework for jointly mapping AI inference workloads to heterogeneous edge SoCs and configuring each processing unit. Unlike traditional black-box optimization, TraceDSE uses a proposer‑critic loop powered by large language models and enriched with system execution traces to identify bottlenecks and refine design choices. Experiments on an Intel Meteor Lake SoC show that TraceDSE outperforms state‑of‑the‑art evolutionary and Bayesian methods, improving Pareto frontier hypervolume by up to 68% while reducing hardware evaluations by 6–9×.

By Geetha Prasuna Yarramneni, Surya Selvam, Wilfried Haensch, Anand Raghunathan
arXiv AI
3d ago

Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents

The paper proposes the Analytical Memory Unit (AMU), a memory schema that attaches a full derivation (lineage) graph to every cached result in enterprise AI agents. By gating retrieval with a policy that requires authorization for every column touched, the authors prove that sensitive columns cannot be leaked through derived results, achieving up to 90% lineage completeness to eliminate leakage. Experiments show lineage‑gated retrieval removes 18.8‑25.5% of cross‑department leakage while maintaining 81.5‑82.6% memory reuse with minimal overhead, and a real‑agent proof‑of‑concept demonstrates zero leaks over multiple interactions.

By Venkata M Sangaraju, Sudhir Vissa
arXiv AI
3d ago

Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale

The paper introduces FlowAgent, an AI agent deployed at Google to automatically repair test failures in the pre-submit continuous integration workflow. FlowAgent uses a ReAct-style generate-and-validate loop with strict latency and quality filters, and was evaluated on 195 real-world failures with a 67.18% accuracy rate. After deployment, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554, and received positive feedback from interviews.

By Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini
arXiv AI
3d ago

A Validated Dataset and Benchmark for Coherent Multi-Diagram SysML Models

The paper introduces SEMAADB, a dataset comprising 3,000 engineering contexts and 15,000 SysML diagrams, each context containing five interconnected views (Requirement, Block Definition, Activity, State Machine, and Sequence). The authors verified diagram consistency and created a 100-context human‑verified benchmark. They evaluated three language models on diagram repair and cross‑diagram update tasks, finding that while syntax repair is largely solved, semantic repair and cross‑diagram consistency remain challenging.

By Ardalan Aryashad, Yan Jin
arXiv AI
3d ago

Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation

The paper introduces Physics‑Guided Visual Prompting (PG‑VP), a plug‑and‑play module that overlays a virtual obstacle onto the input of a frozen Vision‑Language‑Action model to guide navigation around invisible hazards such as radiation or temperature spikes. PG‑VP performs a physics‑based risk assessment to determine the avoidance direction and dynamically renders the same virtual obstacle across frames, allowing the existing navigation policy to detour without retraining. Experiments on OmniNav with R2R‑CE and RxR‑CE datasets show that PG‑VP steers the policy toward low‑risk actions in 84.9% and 83.2% of cases, while real‑world tests on a robot demonstrate significant safety improvements against thermal and radiation sources.

By Hojoon Son, Fan Zhang