arXiv:2608. 14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.
By Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang
arXiv:2608. 13833v1 Announce Type: cross Abstract: Conversational advertising aims to deliver useful ads within multi-turn assistant interactions.
By Simiao Zuo, Chenhui Xu, Yimeng Jia, Qiang Lou, Jian Jiao, Denis Charles
arXiv:2604. 05379v2 Announce Type: replace-cross Abstract: The sequential recommendation (SR) task aims to predict the next item based on users' historical interaction sequences.
By Xing Tang, Ziqiang Cui, Jingyang Bin, Xiaokun Zhang, Fuyuan Lyu, Jingyan Jiang, Dugang Liu, Chen Ma, Xiuqiang He
arXiv:2608. 14022v1 Announce Type: cross Abstract: Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls.
By Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li, Guanchu Wang, Qingbin Liu, Xi Chen, Jiang Bian, Wai Lam
arXiv:2603. 09868v2 Announce Type: replace Abstract: Accurately quantifying terrestrial carbon exchange is essential for climate policy and carbon accounting, yet models must generalize to ecosystems underrepresented in sparse eddy covariance observations.
By Aleksei Rozanov, Arvind Renganathan, Yimeng Zhang, Vipin Kumar
arXiv:2608. 13597v1 Announce Type: cross Abstract: Coverless image steganography (CIS) synthesizes a stego image rather than modifying an existing cover image, enabling authorized recipients to reconstruct the original secret image from the stego.
By Hongxin Xu, Jianping Mei, Can Wang, Defang Chen
arXiv:2608. 02345v2 Announce Type: replace-cross Abstract: A/B testing remains the standard for rolling out new features in the technology industry.
By Stefan Hut, Lorenzo Masoero
arXiv:2608. 13844v1 Announce Type: cross Abstract: Large language models (LLMs) have become core components of cloud-based intelligent services in academia and industry, yet their training and deployment are hindered by high computational costs, data centralization, and privacy concerns.
By Qinglin Yang, Chen Qiu, Hongyuan Zhang, Pengdeng Li, Yuan Liu, Zhihong Tian
arXiv:2608. 14132v1 Announce Type: cross Abstract: Mobile GUI Agents powered by multimodal large language models (MLLMs) show promise in human-computer intelligence.
By Xiaokai Yan, Jingtao Ding, Yong Li, Zhiwen Yu
arXiv:2608. 14492v1 Announce Type: new Abstract: The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks.
By Ben Anson, Conor Houghton, Edward Milsom
We introduce ReRef-3D, a benchmark for language-guided placement in 3D scenes. It contains 33,826 instructions across 998 CLEVR-derived scenes, spanning 16 placement families and direct, one-hop, and two-hop references.
A learning system can occupy execution states that are indistinguishable under every declared present-behavior readout yet respond differently to future training. We formalize this through fiber fingerprints: controlled future-learning response laws restricted to present-behavior equivalence classes.
Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine-tune a pretrained multimodal diffusion transformer (MMDiT) on hundreds of thousands to millions of paired \emph{(reference, composed-target)} examples, where each composed target is a synthesized image of the subject in a novel scene.
Recent advances in large language models (LLMs) have reshaped semantic analysis. Opinion Extraction (OE) for Science and Technology Intelligence (STI) requires concise core opinions from large information streams.
arXiv:2506. 17337v5 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have shown promise in automating image diagnosis and interpretation in clinical settings.
By Yuan Zhong, Ruinan Jin, Qi Dou, Xiaoxiao Li
arXiv:2608. 13072v1 Announce Type: new Abstract: Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology.
By Shuailei Zhang, Muyun Jiang, Wei Zhang, Jinbo Chen, Zhiwei Guo, Yong Li, Yi Ding, Cuntai Guan
arXiv:2505. 12532v3 Announce Type: replace-cross Abstract: Efficiently adapting large pretrained models is critical under tight compute and memory budgets.
By Ahmet Bilican, M. Ak{\i}n Y{\i}lmaz, A. Murat Tekalp, R. G\"okberk Cinbi\c{s}
arXiv:2601. 21628v2 Announce Type: replace-cross Abstract: Diffusion models have achieved remarkable progress in image generation, but their increasing deployment raises serious concerns about privacy and copyright.
By Puwei Lian, Yujun Cai, Songze Li, Bingkun Bao
arXiv:2608. 13069v1 Announce Type: new Abstract: Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants.
By Lucia Mal\'i\v{c}kov\'a
arXiv:2608. 12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain interactions.
By Xuefei Wang, Jun Han, Zixuan Wang, Qingkai Zeng, Xiao Wang, Ruijie Wang, Jianxin Li