arXiv:2610.07043v1 Announce Type: cross
Abstract: Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sam...
By Zhen Li, Shuai Zhang, Yanggan Gu, Yiming Zhang, Yang Yu, Mingfa Feng, Congkai Xie, Shuang Yu, Junjie Lai, Hongxia Yang
arXiv:2610.07062v1 Announce Type: cross
Abstract: Large language models are increasingly used to simulate how individuals respond to new situations, yet the behavioral reasoning behind these response...
By Yining Zhao, Bushi Liu, Haofei Yu, Zhengyang Qi, Shanyong Wang, Chuyue Li, Yuxiang Liu, Jiaxuan You
arXiv:2610.07098v1 Announce Type: cross
Abstract: Large transformer-based models increasingly depend on multi-GPU execution, which requires frequent collective communication among GPUs. Existing comm...
By Keyvan Dadashzadeh, Yuehong Zhou, Minyu Cui, Miquel Pericas
arXiv:2610.07115v1 Announce Type: cross
Abstract: The order in which candidate responses are presented can change an LLM judge's verdict. Detecting such a position flip ordinarily requires judging ea...
By Hashmath Shaik, Gnaneswar Villuri, Alex Doboli
PlaySuite is a large-scale benchmark that evaluates interactive visual intelligence by using over 5,000 open-source video games from platforms like PyWeek and itch.io. The benchmark covers diverse game engines (Pygame, HTML5, Godot, Unity) and introduces a unified closed-loop interaction framework and a Video-LLM-as-a-judge protocol to standardize progress measurement. Evaluation of fourteen recent models shows a perception-action gap, with strong reasoning but poor sustained progress, spatial grounding, action execution, and self-correction.
By Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini, Michelle Lorena Acevedo Callejas, Mohammad Mahdi Derakhshani, Kristof Meding, Joaquin Vanschoren, Cees G. M. Snoek
The paper introduces TraceDSE, an agentic design space exploration framework for jointly mapping AI inference workloads to heterogeneous edge SoCs and configuring each processing unit. Unlike traditional black-box optimization, TraceDSE uses a proposer‑critic loop powered by large language models and enriched with system execution traces to identify bottlenecks and refine design choices. Experiments on an Intel Meteor Lake SoC show that TraceDSE outperforms state‑of‑the‑art evolutionary and Bayesian methods, improving Pareto frontier hypervolume by up to 68% while reducing hardware evaluations by 6–9×.
By Geetha Prasuna Yarramneni, Surya Selvam, Wilfried Haensch, Anand Raghunathan
arXiv:2610.07207v1 Announce Type: cross
Abstract: Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability probl...
By Xin Teng, Muxiao Li, Hongyi Wen
arXiv:2610.07224v1 Announce Type: cross
Abstract: Clinical notes capture most of what is documented about a patient's care, but they cannot be used for research until protected health information (PH...
By Jose D. Posada, Somalee Datta, Priya Desai
The paper proposes the Analytical Memory Unit (AMU), a memory schema that attaches a full derivation (lineage) graph to every cached result in enterprise AI agents. By gating retrieval with a policy that requires authorization for every column touched, the authors prove that sensitive columns cannot be leaked through derived results, achieving up to 90% lineage completeness to eliminate leakage. Experiments show lineage‑gated retrieval removes 18.8‑25.5% of cross‑department leakage while maintaining 81.5‑82.6% memory reuse with minimal overhead, and a real‑agent proof‑of‑concept demonstrates zero leaks over multiple interactions.
By Venkata M Sangaraju, Sudhir Vissa
arXiv:2610.07276v1 Announce Type: cross
Abstract: Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verificatio...
By Xingru Zhou, Luis Sentis, Aarti Choudhary
The paper introduces FlowAgent, an AI agent deployed at Google to automatically repair test failures in the pre-submit continuous integration workflow. FlowAgent uses a ReAct-style generate-and-validate loop with strict latency and quality filters, and was evaluated on 195 real-world failures with a 67.18% accuracy rate. After deployment, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554, and received positive feedback from interviews.
By Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini
arXiv:2610.07298v1 Announce Type: cross
Abstract: Cyber threat analysis increasingly depends on evidence distributed across vendor advisories, vulnerability databases, and threat intelligence sources...
By Luoxi Tang, Yuqiao Meng, Ankita Patra, Weicheng Ma, Muchao Ye, Zhaohan Xi
arXiv:2610.07335v1 Announce Type: cross
Abstract: Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous a...
By Heewon Park, Somin Im, Minhae Kwon
arXiv:2610.07339v1 Announce Type: cross
Abstract: Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidan...
By Junseob Kim, Jade Chng, Ayman Ali, Victor Moas, Yichun Lee, Po-Chun Chin, Sunil Hwang, Rishikesan Kamaleswaran
The paper introduces SEMAADB, a dataset comprising 3,000 engineering contexts and 15,000 SysML diagrams, each context containing five interconnected views (Requirement, Block Definition, Activity, State Machine, and Sequence). The authors verified diagram consistency and created a 100-context human‑verified benchmark. They evaluated three language models on diagram repair and cross‑diagram update tasks, finding that while syntax repair is largely solved, semantic repair and cross‑diagram consistency remain challenging.
By Ardalan Aryashad, Yan Jin
arXiv:2610.07457v1 Announce Type: cross
Abstract: Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GP...
By Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng
arXiv:2610.07460v1 Announce Type: cross
Abstract: Inserting objects into existing 3D scenes requires more than selecting a plausible location:
the inserted object must also fit local geometry while...
By Tzu-Hsin Hsieh, Ricardo Marroquim
arXiv:2610.07532v1 Announce Type: cross
Abstract: LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks inc...
By Wonjun Lee, Kyungsik Yang, Gaeun Ji, Vaidehi Patil, Haon Park, Bumsub Ham, Mohit Bansal, Suhyun Kim
arXiv:2610.07535v1 Announce Type: cross
Abstract: Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model e...
By Dani Roytburg, Daphne Ippolito
The paper introduces Physics‑Guided Visual Prompting (PG‑VP), a plug‑and‑play module that overlays a virtual obstacle onto the input of a frozen Vision‑Language‑Action model to guide navigation around invisible hazards such as radiation or temperature spikes. PG‑VP performs a physics‑based risk assessment to determine the avoidance direction and dynamically renders the same virtual obstacle across frames, allowing the existing navigation policy to detour without retraining. Experiments on OmniNav with R2R‑CE and RxR‑CE datasets show that PG‑VP steers the policy toward low‑risk actions in 84.9% and 83.2% of cases, while real‑world tests on a robot demonstrate significant safety improvements against thermal and radiation sources.
By Hojoon Son, Fan Zhang