arXiv:2610.07607v1 Announce Type: cross
Abstract: Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide...
By SiYuan Ma, Canran Xiao, Zikai Xiao, Albert Gao, Liang He, Xuan-Yu Wang, Shuying Cao, Xiaojun Jia
arXiv:2610.07625v1 Announce Type: cross
Abstract: Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay...
By Qizheng Zhang, Changxiu Ji, Isaac Sun, Yuetai Li, Shubhangi Upasani, Sherry Ruan, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Yoonho Lee, Yuzhen Mao, Genghan Zhang, Rulin Shao, Qiuyang Mang, Andy Dimnaku, Changran Hu, Radha Poovendran, Kunle Olukotun
arXiv:2610.07643v1 Announce Type: cross
Abstract: Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter whi...
By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Wajih Hassan Raza, Atta Ul Asad, Young D. Kwon, Michal Valko, Dean F. Hougen
arXiv:2610.07654v1 Announce Type: cross
Abstract: On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Rece...
By Jian Luo, Kehan Qi, Qingqiao Hu, Meilong Xu, Jiacheng Qiu, Weimin Lyu, Jiawei Zhou, Chao Chen
CACHEFORGE introduces a novel framework that uses a large language model (LLM) to evolve cache‑replacement policies end‑to‑end. In each iteration, the LLM generates new C++ replacement logic, which is evaluated by a trace‑based simulator and refined through reward shaping, structural checks, and mutation. The resulting policies are compact, hardware‑aware, and outperform existing CRC‑2 baselines on SPEC CPU2006, achieving significant improvements in hit rate and IPC across diverse workloads.
By Kaushal Mhapsekar, Bita Aslrousta, Brijesh Kumar Bhayana, Paula Contreras, Azam Ghanbari, Ethan Goodman, Anna Andriiko, Samira Mirbagher Ajorpaz
arXiv:2610.07676v1 Announce Type: cross
Abstract: Research on transformer expressivity shows whether a transformer is capable of solving a given task, but gives little indication of whether the solut...
By Yijia Jessica Zhu, David Chiang
arXiv:2610.07730v1 Announce Type: cross
Abstract: Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a...
By Shuyu Gan, Young-Jun Lee, Dongyeop Kang
arXiv:2610.07742v1 Announce Type: cross
Abstract: Optimized kernels such as FlashAttention and FlashDecoding are crucial for accelerating today's large models. Most of them are handwritten by experts...
By David Pissarra, Jinkun Lin, Haitian Jiang, Aurojit Panda, Jinyang Li
arXiv:2610.07817v1 Announce Type: cross
Abstract: Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually...
By Hans Schabert, Christoph Peters
arXiv:2610.07832v1 Announce Type: cross
Abstract: Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, e...
By Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang
arXiv:2610.07863v1 Announce Type: cross
Abstract: Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with s...
By Yupeng Su, Jiayi Tian, Zheng Zhang, Souvik Kundu
arXiv:2610.07913v1 Announce Type: cross
Abstract: Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whol...
By Shrihari Dumbre, Bikash Santra
arXiv:2610.07940v1 Announce Type: cross
Abstract: Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key...
By Yuhan Chen, Siyuan Zhang, Nan Wang, Feiyang Kang, Ruoxi Jia
VisionWeave introduces elastic visual representation weaving, a native capability for multimodal large language models that learns where and at what granularity to encode visual information. The method combines a gated spatial pooler for coarse representations with a granularity router that allocates content‑adaptive token usage, trained end‑to‑end on large‑scale data. Experiments on Qwen3.5‑4B and Qwen3.8‑27B show that VisionWeave can save 43.0% of tokens while preserving 98.9% of performance across eight benchmarks, and delivers significant throughput gains and latency reductions when deployed on the SGLang serving engine.
By Yuan Feng, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang, Jin Xu, Xike Xie
arXiv:2610.08093v1 Announce Type: cross
Abstract: Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-q...
By Chuan Li, Chengyu Wang, Cen Chen, Ye Lyu, Mingyuan Fan, Ming Gao
The paper introduces a new evaluation setting called penalty‑framed no‑valid‑option MCQA, where multiple‑choice questions may contain no correct answer. By removing the correct option from the MMLU‑Pro mathematics subset and allowing models to either pick an option or abstain, the authors penalize forced‑choice responses that are invalid. Experiments reveal that even models with high standard MCQA accuracy can still produce invalid forced‑choice answers, indicating that traditional accuracy metrics miss an important aspect of model reliability.
By Jinhyeok Kim, Hye-Young Jung
arXiv:2610.08155v1 Announce Type: cross
Abstract: Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by en...
By Zihan Zhou, Xinzhe Hu, Hanxu Yang, Liangjian Wen, Zhao Kang
arXiv:2610.08183v1 Announce Type: cross
Abstract: Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performa...
By Nanhe Chen, Runqiu Yang, Jiawei Tang, Sichao Liu, Yuquan Wang
arXiv:2610.08268v1 Announce Type: cross
Abstract: Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models...
By Jingpo Xu, Paul Joe Maliakel, Ivona Brandic, Shashikant Ilager
arXiv:2610.08331v1 Announce Type: cross
Abstract: The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in ex...
By Heyam Bin Jahlan Areej Alhothali Abeer Alhothali