arXiv:2610.07335v1 Announce Type: cross
Abstract: Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous a...
By Heewon Park, Somin Im, Minhae Kwon
arXiv:2610.07339v1 Announce Type: cross
Abstract: Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidan...
By Junseob Kim, Jade Chng, Ayman Ali, Victor Moas, Yichun Lee, Po-Chun Chin, Sunil Hwang, Rishikesan Kamaleswaran
The paper introduces SEMAADB, a dataset comprising 3,000 engineering contexts and 15,000 SysML diagrams, each context containing five interconnected views (Requirement, Block Definition, Activity, State Machine, and Sequence). The authors verified diagram consistency and created a 100-context human‑verified benchmark. They evaluated three language models on diagram repair and cross‑diagram update tasks, finding that while syntax repair is largely solved, semantic repair and cross‑diagram consistency remain challenging.
By Ardalan Aryashad, Yan Jin
arXiv:2610.07457v1 Announce Type: cross
Abstract: Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GP...
By Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng
arXiv:2610.07460v1 Announce Type: cross
Abstract: Inserting objects into existing 3D scenes requires more than selecting a plausible location:
the inserted object must also fit local geometry while...
By Tzu-Hsin Hsieh, Ricardo Marroquim
arXiv:2610.07532v1 Announce Type: cross
Abstract: LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks inc...
By Wonjun Lee, Kyungsik Yang, Gaeun Ji, Vaidehi Patil, Haon Park, Bumsub Ham, Mohit Bansal, Suhyun Kim
arXiv:2610.07535v1 Announce Type: cross
Abstract: Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model e...
By Dani Roytburg, Daphne Ippolito
The paper introduces Physics‑Guided Visual Prompting (PG‑VP), a plug‑and‑play module that overlays a virtual obstacle onto the input of a frozen Vision‑Language‑Action model to guide navigation around invisible hazards such as radiation or temperature spikes. PG‑VP performs a physics‑based risk assessment to determine the avoidance direction and dynamically renders the same virtual obstacle across frames, allowing the existing navigation policy to detour without retraining. Experiments on OmniNav with R2R‑CE and RxR‑CE datasets show that PG‑VP steers the policy toward low‑risk actions in 84.9% and 83.2% of cases, while real‑world tests on a robot demonstrate significant safety improvements against thermal and radiation sources.
By Hojoon Son, Fan Zhang
arXiv:2610.07607v1 Announce Type: cross
Abstract: Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide...
By SiYuan Ma, Canran Xiao, Zikai Xiao, Albert Gao, Liang He, Xuan-Yu Wang, Shuying Cao, Xiaojun Jia
arXiv:2610.07625v1 Announce Type: cross
Abstract: Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay...
By Qizheng Zhang, Changxiu Ji, Isaac Sun, Yuetai Li, Shubhangi Upasani, Sherry Ruan, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Yoonho Lee, Yuzhen Mao, Genghan Zhang, Rulin Shao, Qiuyang Mang, Andy Dimnaku, Changran Hu, Radha Poovendran, Kunle Olukotun
arXiv:2610.07643v1 Announce Type: cross
Abstract: Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter whi...
By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Wajih Hassan Raza, Atta Ul Asad, Young D. Kwon, Michal Valko, Dean F. Hougen
arXiv:2610.07654v1 Announce Type: cross
Abstract: On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Rece...
By Jian Luo, Kehan Qi, Qingqiao Hu, Meilong Xu, Jiacheng Qiu, Weimin Lyu, Jiawei Zhou, Chao Chen
CACHEFORGE introduces a novel framework that uses a large language model (LLM) to evolve cache‑replacement policies end‑to‑end. In each iteration, the LLM generates new C++ replacement logic, which is evaluated by a trace‑based simulator and refined through reward shaping, structural checks, and mutation. The resulting policies are compact, hardware‑aware, and outperform existing CRC‑2 baselines on SPEC CPU2006, achieving significant improvements in hit rate and IPC across diverse workloads.
By Kaushal Mhapsekar, Bita Aslrousta, Brijesh Kumar Bhayana, Paula Contreras, Azam Ghanbari, Ethan Goodman, Anna Andriiko, Samira Mirbagher Ajorpaz
arXiv:2610.07676v1 Announce Type: cross
Abstract: Research on transformer expressivity shows whether a transformer is capable of solving a given task, but gives little indication of whether the solut...
By Yijia Jessica Zhu, David Chiang
arXiv:2610.07730v1 Announce Type: cross
Abstract: Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a...
By Shuyu Gan, Young-Jun Lee, Dongyeop Kang
arXiv:2610.07742v1 Announce Type: cross
Abstract: Optimized kernels such as FlashAttention and FlashDecoding are crucial for accelerating today's large models. Most of them are handwritten by experts...
By David Pissarra, Jinkun Lin, Haitian Jiang, Aurojit Panda, Jinyang Li
arXiv:2610.07817v1 Announce Type: cross
Abstract: Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually...
By Hans Schabert, Christoph Peters
arXiv:2610.07832v1 Announce Type: cross
Abstract: Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, e...
By Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang
arXiv:2610.07863v1 Announce Type: cross
Abstract: Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with s...
By Yupeng Su, Jiayi Tian, Zheng Zhang, Souvik Kundu
arXiv:2610.07913v1 Announce Type: cross
Abstract: Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whol...
By Shrihari Dumbre, Bikash Santra