arXiv:2609.39566v1 Announce Type: new
Abstract: Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected e...
By Minye Shao, Chaohui Yu, Yixuan Wu, Fan Wang, Ling Shao, Yang Long
arXiv:2609.39657v1 Announce Type: new
Abstract: Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and...
By Adrien Ramanana Rahary, Nicolas Dufour, Patrick P\'erez, David Picard
arXiv:2609.39894v1 Announce Type: new
Abstract: Different from natural videos, Screen Content Videos (SCVs) are characterized by abrupt motion, scene switches, and high-frequency details such as text...
By Ziyin Huang, Sik-Ho Tsang, Xinyuan Qin, Yui-Lam Chan, Xueling Zhou, Feiyu Chen
arXiv:2609.39953v1 Announce Type: new
Abstract: Omni-modal large language models (OmniLLMs) enable unified audio-video understanding, but their long multimodal token sequences make deployment computa...
By Jianghao Wang, Ke Meng, Jian Li, Chi Cheng, Longyu Qi, Liyin Liang, Yifeng Qian, Chunbo Lai, Yutian Lin, Zeyu Wang
arXiv:2602.02914v4 Announce Type: replace
Abstract: Privacy-preserving face recognition (PPFR) and face anonymization have different goals, but both must retain some identity-related information for...
By Wenqi Guo, Qingyun Qian, Mohamed Shehata, Shan Du
arXiv:2602.05391v3 Announce Type: replace
Abstract: Dataset distillation seeks to synthesize a compact surrogate dataset that enables performance comparable to training on the original dataset for do...
By Qianxin Xia, Jiawei Du, Yuhan Zhang, Xin Zhang, Xuewan He, Wenbo Jiang, Jielei Wang, Tao Luo, Guoming Lu
arXiv:2605.12649v3 Announce Type: replace
Abstract: Dataset distillation aims to synthesize a compact proxy dataset that is unreadable or non-raw from the original dataset for privacy protection and...
By Qianxin Xia, Zhiyong Shu, Wenbo Jiang, Jiawei Du, Jielei Wang, Guoming Lu
arXiv:2609.34367v2 Announce Type: replace
Abstract: Learned image codecs (LICs) achieve high reconstruction quality, but their decoding speed is often insufficient for immersive virtual reality (VR)....
By Yulong Cheng, Youneng Bao, Junfeng Zhou, Mu Li, Jie Wen
arXiv:2609.34863v2 Announce Type: replace
Abstract: Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Thei...
By Cilin Yan, Yilun Qiu, Wanyang Zhang, Rui Zu, Xiaolong Jiang, Jiayin Cai, Yao Hu
arXiv:2609.34220v2 Announce Type: replace-cross
Abstract: Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as ob...
By Junqiao Fan, Yuxuan Hu, Bofan Lyu, Yanshuo Lu, Pengfei Liu, Jiarui Zhang, Fangqiang Ding, Lihua Xie, Gen Li, Jianfei Yang
arXiv:2609.40047v1 Announce Type: cross
Abstract: Gromov-Wasserstein multidimensional scaling (GW-MDS) learns low-dimensional representations from relational data but remains transductive, providing...
By Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante
arXiv:2609.38855v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models based on generative frameworks, such as Flow Matching, have recently achieved impressive performance in robotic m...
By Gongxin Yao, Yongsheng Zhao, Jiayin Deng, Deng Liang, Han Gao, Lei Zhao, Baoping Cheng
arXiv:2609.38984v1 Announce Type: cross
Abstract: World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and...
By Xinling Xie, Haodong Wang, Jiazhi Mi, Zhiming Liu, Zicong Hong, Xiaoyi Pang, Qianli Liu, Yangjia Hu, Ying Chen, Zhengyang Yan, Song Guo
arXiv:2609.39957v1 Announce Type: cross
Abstract: Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and incre...
By Jiangrui Zhao, Chenglong Li, Meng Zhang, Xiaoting Du
arXiv:2609.40055v1 Announce Type: cross
Abstract: On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for...
By Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djuri\'c, Sima Mofakham
arXiv:2601.12310v2 Announce Type: replace
Abstract: Self-training systems often degenerate due to the lack of an external criterion for judging data quality, leading to reward hacking and semantic dr...
By Jennifer Dodgson, Alfath Daryl Alhajir, Michael Joedhitya, Akira Rafhael Janson Pattirane, Surender Suresh Kumar, Joseph Lim, C. H. Peh, Adith Ramdas, Steven Zhang Zhexu
arXiv:2609.39451v1 Announce Type: new
Abstract: Progressive autoregressive image codecs provide an appealing paradigm for generative compression by quantizing continuous latents into discrete tokens,...
By Qin Yan, Ruixiao Dong, Yutao Xie, Li Li, Ying Chen, Kai Li, Daowen Li, Houqiang Li
arXiv:2609.39553v1 Announce Type: new
Abstract: 3D Gaussian Splatting (3DGS) enables real-time novel view synthesis, but existing general-purpose acceleration methods suffer severe rendering quality...
By Changbai Li, Shuo Yang, Yichen Yang, Shuwei Shao, Huobin Tan
arXiv:2609.38282v1 Announce Type: new
Abstract: Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Seq...
By Baode Wang, Zuming Huang, Kexuan Ren, Jun Huang, Wei Chu
arXiv:2609.38757v1 Announce Type: new
Abstract: Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorit...
By Chen Lu, Ke Xue, Siyuan Xu, Mingxuan Yuan, Chao Qian