Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,035 stories · RSS feed

arXiv AI
2d ago

APEX: Active Protection at Execution Boundaries for LLM Agents

The paper introduces APEX, an active defense for large language model agents that protects against indirect prompt injection by enforcing safety at execution boundaries. APEX uses an evidence‑gated prevention contract and deception‑based exposure to ensure that only authorized effects, endorsed by the task, are executed. Evaluation shows APEX achieves near‑zero attack success across multiple benchmarks and capability‑unit types, outperforming 13 baseline defenses.

By Xinran Zheng, Xin Fan Guo, Zhiqiang Hao, Fan Yang, Xingzhi Qian, Jiawei Du, Jinfeng Xu, Zheng Xing, Shuo Yang, Xingjun Wang
arXiv AI
2d ago

PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence

PlaySuite is a large-scale benchmark that evaluates interactive visual intelligence by using over 5,000 open-source video games from platforms like PyWeek and itch.io. The benchmark covers diverse game engines (Pygame, HTML5, Godot, Unity) and introduces a unified closed-loop interaction framework and a Video-LLM-as-a-judge protocol to standardize progress measurement. Evaluation of fourteen recent models shows a perception-action gap, with strong reasoning but poor sustained progress, spatial grounding, action execution, and self-correction.

By Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini, Michelle Lorena Acevedo Callejas, Mohammad Mahdi Derakhshani, Kristof Meding, Joaquin Vanschoren, Cees G. M. Snoek
arXiv AI
2d ago

Agentic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs

The paper introduces TraceDSE, an agentic design space exploration framework for jointly mapping AI inference workloads to heterogeneous edge SoCs and configuring each processing unit. Unlike traditional black-box optimization, TraceDSE uses a proposer‑critic loop powered by large language models and enriched with system execution traces to identify bottlenecks and refine design choices. Experiments on an Intel Meteor Lake SoC show that TraceDSE outperforms state‑of‑the‑art evolutionary and Bayesian methods, improving Pareto frontier hypervolume by up to 68% while reducing hardware evaluations by 6–9×.

By Geetha Prasuna Yarramneni, Surya Selvam, Wilfried Haensch, Anand Raghunathan
arXiv AI
2d ago

Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents

The paper proposes the Analytical Memory Unit (AMU), a memory schema that attaches a full derivation (lineage) graph to every cached result in enterprise AI agents. By gating retrieval with a policy that requires authorization for every column touched, the authors prove that sensitive columns cannot be leaked through derived results, achieving up to 90% lineage completeness to eliminate leakage. Experiments show lineage‑gated retrieval removes 18.8‑25.5% of cross‑department leakage while maintaining 81.5‑82.6% memory reuse with minimal overhead, and a real‑agent proof‑of‑concept demonstrates zero leaks over multiple interactions.

By Venkata M Sangaraju, Sudhir Vissa
arXiv AI
2d ago

Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale

The paper introduces FlowAgent, an AI agent deployed at Google to automatically repair test failures in the pre-submit continuous integration workflow. FlowAgent uses a ReAct-style generate-and-validate loop with strict latency and quality filters, and was evaluated on 195 real-world failures with a 67.18% accuracy rate. After deployment, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554, and received positive feedback from interviews.

By Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini