arXiv:2609.26173v1 Announce Type: new
Abstract: Many post-training quantization (PTQ) methods use layer-wise reconstruction, second-order proxy objectives, or activation-aware transformations to redu...
By Kasun Dewage, Marianna Pensky, Suranadi De Silva
arXiv:2609.26199v1 Announce Type: new
Abstract: A large graph is often available only in part: a crawl stopped by its budget, a panel, a partial dump. When the sampled fraction $s$ is known by design...
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv:2609.26216v1 Announce Type: new
Abstract: A correct teacher solution becomes useful supervision when the receiving student can continue its reasoning. We measure this compatibility with prefix...
By Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Han Wang, Jie Li, Ru Zhang
arXiv:2609.26402v1 Announce Type: new
Abstract: The discovery of novel inorganic materials drives technological breakthroughs in critical fields such as computing and energy storage. Generative AI ha...
By Thomas Egg, Harry Winston Sullivan, Ellad B. Tadmor, Stefano Martiniani
arXiv:2609.26219v1 Announce Type: cross
Abstract: Long-running LLM agent workflows often revise interior context spans while retaining long suffixes. Although suffix tokens remain unchanged, altered...
By Guotao Yang, Rui Guo, Siwei He, Sheng Chen, Yitao Hu, Keqiu Li
arXiv:2609.25008v1 Announce Type: new
Abstract: I pretrained a language model end-to-end in Rust - alone, with no team, no PyTorch, and no Python in the training path - for $164 in rented GPU time. I...
By Arif Adito
arXiv:2609.25054v1 Announce Type: new
Abstract: For a long-horizon LLM agent, the memory question is not what was once recorded but what \emph{currently holds}. Most designs answer it only indirectly...
By Bowen Qin, Yao Lu
arXiv:2609.25056v1 Announce Type: new
Abstract: A feedback-based word deduction framework based on the Jotto problem is proposed, and the problem space is represented as a weighted graph where all va...
By Dakshi Arora, Prakhar Kumar Srivastava, Ranjib Banerjee
arXiv:2609.26638v1 Announce Type: new
Abstract: Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step...
By Dohyun Kim, Sungjun Han, Hyungguk Kim, Yusik Kim, Jamin Shin, Paul Hongsuck Seo, Hongjoon Ahn
arXiv:2503.15242v3 Announce Type: replace
Abstract: We introduce BigO(Bench), a novel coding benchmark designed to evaluate the capabilities of generative language models in understanding and generat...
By Pierre Chambon, Baptiste Roziere, Benoit Sagot, Gabriel Synnaeve
arXiv:2609.25891v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) often struggle with fine-grained visual perception when processing complete images, as critical evidence may o...
By Zihan Chen, Hengguang Zhou, Yuan Kang, Yiming Zhang, Wenhui Fang, Zenghui Ding, Yining Sun, Cho-Jui Hsieh
arXiv:2605.30263v2 Announce Type: replace
Abstract: Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time intera...
By Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou, Yimin Chen, Honglie Wang, Wenqiang Sun, Zhengwei Fang, Zizheng Xun, Zihao Li, Kaiwen Zheng, Guande He, Xiao Yang, Chongxuan Li, Fan Bao, Jun Zhu
S$^4$R is a method for compressing the Key-Value cache in large language models by building low‑rank subspaces from selectively sampled tokens and performing attention over a sparsely reconstructed KV representation. It initializes key/value bases using a representative prompt subset, reducing reliance on external calibration data while avoiding the high compute cost of full‑prompt decomposition. Experiments on LongBench and RULER with Llama and Qwen models demonstrate up to 5× KV compression with near‑full‑cache accuracy, blending the efficiency of fixed compression with the adaptability of prompt‑dependent approaches.
By Jialong Han, You Wu, Kewei Tu
KwaiMind is a commercial image editing system that combines general editing capabilities with e-commerce specialization. It uses an agent-based data engine with 1.8 million editing pairs and a multimodal diffusion transformer trained through pre‑training, fine‑tuning, preference optimization, and online reinforcement learning. The system is guided by a vision‑language judge and specialized rewards for click‑through rate, text rendering, and product consistency, and it achieves top scores on ImgEdit, GEdit, REDEdit, and the new Ecom‑Bench, while improving predicted and actual CTR in offline and online experiments.
By Junlong Wu, Zijun Li, Yuting Hu, Jia Sun, Pengcheng Wei, Yimin Zhou, Honglie Wang, Huaiqing Wang, Dewen Fan, Fei Zuo, Haixuan Gao, Lihui Peng, Tingxuan She, Yuqing Li, Boheng Zhang, Fan Yang, Wenwu Ou
Long-horizon language-model agents accumulate reasoning traces, tool exchanges, and observations whose relevance changes with the current decision. Existing compression strategies often score historic...
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or freque...
PRQuant introduces a training‑free, low‑overhead method for low‑bit quantization of linear layers by permuting input channels that cause the largest quantization error into contiguous tail blocks and precomputing residual weight sub‑tensors. The approach eliminates the need for online gathering during inference, converting scattered residual compensation into a regular tail‑augmented GEMM and thereby reducing latency. Experiments show that PRQuant lowers down‑projection reconstruction error and outperforms standard MXFP4 and other post‑training quantization baselines on five downstream benchmarks, improving accuracy by up to 1.24 points on Qwen3‑4B‑Instruct‑2507.
By Peiran Wang, Anqi Wang, Jiaying Zhao, Huiwen Yang, Zhenyu Ming, Rongqian Wang, Yiwu Yao, Kun Tian, Xin Yao, Gong Zhang, Fan Yang, Zhongyi Huang
Moonworks Lunara is a text‑to‑image model that defines Artistic Intelligence as exploration‑driven world realization, preserving semantic, artistic, and compositional structure. It uses a Diffusion Mixture Transformer architecture and a training algorithm that iteratively refines the data distribution with informative samples and human‑created art. Benchmarks show Lunara ranks first in aesthetic quality and second in emotional resonance against seven other image‑generation models, while maintaining a sub‑10B parameter size and sub‑10‑second inference latency.
By Yan Wang, Yanzu Wang, Maitreyee Joshi, Samiha Sadeka, Partho Hassan, Reza Jarral, Sayeef Abdullah, Sabit Hassan
MTMed3D is a multi-task Transformer-based model that jointly performs 3D detection, segmentation, and classification in medical imaging. It uses a shared Transformer encoder to produce multi-scale features, with separate CNN decoders for each task. Evaluated on BraTS 2018 and 2019, it achieves strong results, especially in detection, while reducing computational cost and inference time compared to single-task models.
By Fan Li, Arun Iyengar, Lanyu Xu
StepKV introduces a step-aware approach to compressing the key-value cache used during large language model inference, treating reasoning steps as primary units of retention rather than individual tokens. By linking cache entries to the steps that generated them and estimating each step’s utility from trajectory signals, StepKV assigns a combined token‑ and step‑level score to guide pruning. Experiments on multi‑hop question answering and long‑horizon web reasoning show that StepKV maintains accuracy even under tight cache budgets, outperforming token‑level baselines that suffer sharp performance drops.
By Boyu Feng, Jiahong Liu, Yifan Li, Wenhao Yu, Zexuan Qiu, Yuliang Sun, Ming Shen, Xiang Li, Quanyu Dai, Irwin King