arXiv:2609.24259v1 Announce Type: new
Abstract: The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over...
By Ruike Cao, Fanyu Zhao, Fugen Yao, Liang Dong, Jian Xu, Guanjun Jiang, Yifei Zhao, Han Zhang, Li Xiao
arXiv:2609.24298v1 Announce Type: new
Abstract: What limits KV-cache compression at extreme bit-rates? We argue that it is not the choice of compression scheme, but how its budget is allocated across...
By Sihyeon Ha, Jaeho Lee, Yo-Seb Jeon
arXiv:2609.24322v1 Announce Type: new
Abstract: We show that a quantized model that keeps its classification accuracy still changes $14$ to $46\%$ of its top-1 retrieval results, and that aggregate r...
By Luca Zhou, Alessandro Zirilli, Daniele Solombrino, Roberto Dess\`i, Emanuele Rodol\`a
arXiv:2609.24417v1 Announce Type: new
Abstract: Despite advances in long-context inference, large language models (LLMs) remain fundamentally limited by the key-value (KV) caching mechanisms that are...
By Qiuhao Zeng, Jerry Huang, Peng Lu, Ruiyi Fang, Gezheng Xu, Zihao Jing, Yufei Cui, Charles Ling, Gang Niu, Boyu Wang
arXiv:2609.24485v1 Announce Type: new
Abstract: Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction oft...
By Guangchuan Lv, Dianxing Shi, Dingjie FU
arXiv:2609.24538v1 Announce Type: new
Abstract: Functional annotation of newly sequenced proteins remains a bottleneck in molecular biology: the number of sequences in public repositories grows far f...
By Demian Pavlyshenko, Bohdan Pavlyshenko
arXiv:2609.24646v1 Announce Type: new
Abstract: On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full d...
By Ahmed Khaled Khamis, Xiaotong Ji, Hassan Jaber, Rasul Tutunov, Matthieu Zimmer, Jun Wang, Haitham Bou-Ammar
arXiv:2609.24691v1 Announce Type: new
Abstract: Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the late...
By Niklas Bubeck, Yundi Zhang, Vasiliki Sideri-Lampretsa, Julian McGinnis, Jiancheng Yang, Daniel Rueckert, Jiazhen Pan
arXiv:2609.24698v1 Announce Type: new
Abstract: Repeated execution of the target model during autoregressive decoding is a major source of LLM inference latency. Unlike linear speculation, which foll...
By Changxu Liu, Zhaogeng Li
arXiv:2507.22920v2 Announce Type: replace
Abstract: The rapid advancement of large language models (LLMs) has intensified the need for effective mechanisms to transform continuous multimodal data int...
By Jindong Li, Yali Fu, Jiahong Liu, Linxiao Cao, Wei Ji, Menglin Yang, Irwin King, Ming-Hsuan Yang
arXiv:2510.03095v4 Announce Type: replace
Abstract: Diffusion- and flow-based generative models have recently demonstrated strong performance in protein backbone generation tasks, offering unpreceden...
By Liyang Xie, Haoran Zhang, Zhendong Wang, Wesley Tansey, Mingyuan Zhou
arXiv:2512.07624v2 Announce Type: replace
Abstract: Process Model Forecasting (PMF) aims to predict how the control-flow structure of a process evolves over time by modeling the temporal dynamics of...
By Yongbo Yu, Jari Peeperkorn, Johannes De Smedt, Jochen De Weerdt
arXiv:2604.11508v3 Announce Type: replace
Abstract: Fine-tuning a pretrained classifier leaves some samples reliably learned and others cycling between correct and incorrect. Curriculum learning, dat...
By Miit Daga, Swarna Priya Ramu
The paper introduces a hybrid attention model that learns a unified time‑aware patch representation for irregular multivariate time series (IMTS) forecasting. It employs a time‑aware patch encoding to embed variable‑length intra‑patch timestamps, a time bias attention mechanism to adjust for temporal misalignment and asynchronous cross‑channel dependencies, and a hybrid causal mask on a decoder‑only Transformer to balance historical context with autoregressive forecasting. The authors also curate VersaTSA, a 30 B‑observation dataset preserving native sampling sparsity, and demonstrate state‑of‑the‑art zero‑shot performance on three IMTS benchmarks while remaining competitive on regular MTS tasks.
By Zhihao Lin, Li Lin, Qi Zhang, Kaiwen Xia, Shuai Wang, Jialin Qiao
arXiv:2609.23697v1 Announce Type: cross
Abstract: Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, howe...
By Jie Sun, Mao Zheng, Mingyang Song, Zeyuan Liu, Gengsheng Li, Houcheng Jiang, Yilin Cheng, Bichuan Feng, Yuchen Cai, Junfeng Fang, Xiang Wang
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world m...
Recent foundation models are moving toward native multimodal Vision-Language Models (VLMs), making VLMs a central form of next-generation foundation models. However, their large language backbones mak...
LiDAR point clouds acquired in underground environments exhibit severe geometric incompleteness due to occlusions and limited sensor viewpoints, making reliable point cloud completion challenging with...
Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We...