arXiv:2609.25088v1 Announce Type: cross
Abstract: Survival prediction for glioblastoma multiforme (GBM) demands models that are both accurate and interpretable, yet existing approaches treat these ob...
By Mushahid Intesum
arXiv:2609.25558v1 Announce Type: cross
Abstract: Vision-language-action policies benefit from geometric supervision, but current-frame geometry alone does not explicitly describe the changes associa...
By Jinu Pahk, Jesoon Kang, Taegeon Park, Jisu An, Soo Min Kimm, Jaejoon Kim, Byoung-Tak Zhang
arXiv:2609.26638v1 Announce Type: new
Abstract: Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step...
By Dohyun Kim, Sungjun Han, Hyungguk Kim, Yusik Kim, Jamin Shin, Paul Hongsuck Seo, Hongjoon Ahn
arXiv:2609.25832v1 Announce Type: new
Abstract: Part segmentation is a fundamental problem in computer graphics and 3D vision. Recent works have expanded 3D part segmentation beyond fixed taxonomies,...
By Zhe Zhu, Yiheng Zhang, Peng Li, Zixing Zhao, Honghua Chen, Yaqing Zhang, Le Wan, Zhiyang Dou, Cheng Lin, Yuan Liu, Mingqiang Wei, Wenping Wang
arXiv:2609.25891v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) often struggle with fine-grained visual perception when processing complete images, as critical evidence may o...
By Zihan Chen, Hengguang Zhou, Yuan Kang, Yiming Zhang, Wenhui Fang, Zenghui Ding, Yining Sun, Cho-Jui Hsieh
arXiv:2609.26056v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) offer promising capabilities for automated sports coaching but face a fundamental limitation: they implicitly compare aga...
By Agamdeep Singh, Sujit PB, Mayank Vatsa
arXiv:2609.26274v1 Announce Type: new
Abstract: The rapid evolution of generative AI (e.g., Sora, Hunyuan) makes it essential to develop effective detection strategies that can generalize across ever...
By S. Hong, X. Q. Wang, C. Zhang, J. C. Wang, P. X. Duan, Y. W. Wang
arXiv:2609.23796v2 Announce Type: replace
Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open ch...
By Yang-Tian Sun, Tianjia Liu, Zehuan Huang, Yi-Hua Huang, Xiaoyang Lyu, Ziyi Yang, Zi-Xin Zou, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi
The paper investigates how reinforcement learning can unintentionally obscure the chain‑of‑thought (CoT) reasoning in vision‑language models, making their internal reasoning less traceable. By analyzing activation patterns, the authors show that template‑associated activations become less distinguishable during RL and that targeted interventions can mitigate this effect. They introduce TAME, a method that uses sparse autoencoders to suppress these problematic activations while still encouraging accurate behavior, achieving significant gains in CoT monitorability across multiple datasets and model families.
By Xutao Mao, Jianing Zhu, Jinman Zhao, Tongliang Liu, Xiaowen Chu, Cong Wang, Bo Han
SpecialEduBench is a new benchmark for vision‑language models that evaluates their pedagogical competence in language intervention for autistic children across knowledge, skill, and attitude dimensions. It contains 4,537 knowledge items, 200 skill items, and 68 attitude items derived from recorded interventions, with 192 response cells for attitude items that combine pressure and monitoring. Eight leading vision‑language models were tested, none achieving full performance, especially in honesty cells, indicating that current models still struggle with situated teaching tasks.
By Jihoi Na, Taeyeong Kim, Sungjune Kong, Jaemin Jung, Min Joung Park, Kyungtae Joo, Ahhyun Kim, Shim Jaechang, Sooyoung Joo, Dongjin Ka, SeJoong Kim, Jimin Kim, HyunJin Jung, Unggi Lee
KwaiMind is a commercial image editing system that combines general editing capabilities with e-commerce specialization. It uses an agent-based data engine with 1.8 million editing pairs and a multimodal diffusion transformer trained through pre‑training, fine‑tuning, preference optimization, and online reinforcement learning. The system is guided by a vision‑language judge and specialized rewards for click‑through rate, text rendering, and product consistency, and it achieves top scores on ImgEdit, GEdit, REDEdit, and the new Ecom‑Bench, while improving predicted and actual CTR in offline and online experiments.
By Junlong Wu, Zijun Li, Yuting Hu, Jia Sun, Pengcheng Wei, Yimin Zhou, Honglie Wang, Huaiqing Wang, Dewen Fan, Fei Zuo, Haixuan Gao, Lihui Peng, Tingxuan She, Yuqing Li, Boheng Zhang, Fan Yang, Wenwu Ou
arXiv:2605.11151v3 Announce Type: replace
Abstract: Offline-to-online reinforcement learning (RL) improves sample efficiency by leveraging pre-collected datasets prior to online interaction. A key ch...
By Andrew Choi, Wei Xu
arXiv:2609.23997v1 Announce Type: cross
Abstract: Multi-robot collaboration could enable more efficient and scalable solutions to complex robotic tasks, but collaboration under partial observability...
By Dorian Benhamou Goldfajn, Mason Nakamura, Saaduddin Mahmud, Justin Svegliato, Kyle H. Wray, Shlomo Zilberstein
An agent that interacts with users over long periods must recall facts, preferences, events, and changes from a continuously growing interaction history. Existing memory systems often compress interac...
Text-to-speech systems increasingly process user-generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical rather than surface form. We introduce UGTPhon, the f...
Foundation model embeddings of screening mammograms may encode pre-diagnostic tissue change without task-specific adaptation. We tested whether embeddings move faster along a data-derived "cancer dire...
Retraining visual perception pipelines in High-Mix, Low-Volume (HMLV) automotive manufacturing must be carried out under tight annotation, energy, and time budgets, yet most Synthetic Data Generation...
Robots operating around pedestrians often reason over a finite set of predicted human futures. Repeated online updates can concentrate this limited prediction budget on dominant destinations and leave...
The paper introduces BAS‑VLA, a task‑semantic action calibration framework for vision‑language‑action models that addresses two failure modes: unnecessary action drift under appearance changes and insufficient behavioral change under semantic alterations. BAS‑VLA uses a breaking‑centered calibration core and a selective evidence‑gated preserving auxiliary to maintain performance on clean and semantics‑preserving conditions while suppressing stale‑task behavior. Experiments on OpenPI‑pi0.5 and LIBERO‑Object Milk‑Swap show high success rates on clean and preserved tasks, a dramatic drop under target‑object swaps, and improved robustness to style shifts from 42% to 70% without harming clean performance.
By Shuaijun Liu, Feiyang You, Chengyu Wu, Shuyang Hao, Chenglong Zhang, Jingyao Cai, Xingwei Chen, Ningxin Su
The paper introduces an all‑in‑one multilingual scene text recognizer called ScriptMoE, which uses a script‑aware mixture‑of‑experts architecture to handle 10 scripts and 229 languages. It is built on a new large‑scale synthetic dataset, TextMuSS‑10M, and evaluated on the TextMuSS‑Bench, achieving 82.06% accuracy—1.31% higher than the best baseline. When integrated into the PP‑OCRv5 pipeline, ScriptMoE raises the end‑to‑end multilingual F1 score from 65.71% to 80.89%, slightly surpassing the best vision‑language model while using far fewer parameters.
By Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen