arXiv:2609.37080v1 Announce Type: new
Abstract: Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusi...
By Zhengqiang Zhang, Lingchen Sun, Rongyuan Wu, Qiaosi Yi, Xiangtao Kong, Chaodong Xiao, Lei Zhang
arXiv:2609.37495v1 Announce Type: new
Abstract: Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While exis...
By Yun Chen, Munchurl Kim, Jeonghyeok Do
arXiv:2609.37655v1 Announce Type: new
Abstract: Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA)...
By Jiayu Ying, Qijian Tian, Ruijie Xu, Xinnan Zhu, Daoguo Dong, Jiachen Xu, Xin Tan
arXiv:2609.37685v1 Announce Type: new
Abstract: Language-aligned and self-supervised visual encoders offer complementary strengths in semantic abstraction and spatial detail. Harnessing this compleme...
By Pinze Ren, Yuwei Zhang, Hao Chen, Linghao Meng, Chang Li, Qiankun Li
arXiv:2609.36944v1 Announce Type: new
Abstract: Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localiz...
By Zhongyi Li, Wan Tian, Xiang Xu, Yutian Xiao, Yikun Ban, Yijie Peng, Fuzhen Zhuang
arXiv:2609.37658v1 Announce Type: new
Abstract: LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-te...
By Min Yang, Yichen Pan, Jinghua Piao, Dandan Song, Yongshun Gong, Yong Li
arXiv:2609.35804v1 Announce Type: cross
Abstract: Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment...
By Mamehgol Yousefi, Ahmad Shahi, Mos Sharifi, Alvaro Romera, Simon Hoermann, Tham Piumsomboon
arXiv:2609.35952v1 Announce Type: cross
Abstract: We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real hum...
By Shen Yan, Duc Le, Irina-Elena Veliche
arXiv:2609.37002v1 Announce Type: cross
Abstract: High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to a...
By Xijia Tao, Yihua Teng, Xinyu Fu, Cheng Gong, Ziru Liu, Xudong Xie, Rui Liu, Lingpeng Kong
arXiv:2609.37491v1 Announce Type: cross
Abstract: Retrieval-augmented language models are expected to answer from the retrieved evidence, but in practice they often keep answering when that evidence...
By Zeyan Li, Qirong Guo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
arXiv:2602.20731v2 Announce Type: replace-cross
Abstract: Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or...
By Aram Davtyan, Yusuf Sahin, Yasaman Haghighi, Sebastian Stapf, Pablo Acuaviva, Alexandre Alahi, Paolo Favaro
arXiv:2608.00042v2 Announce Type: replace-cross
Abstract: Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrain...
By Ramesh B. Paramkusham
arXiv:2609.37675v1 Announce Type: new
Abstract: Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and str...
By Biswajit Banerjee, Claudia Alvarez Carreno, Anton S. Petrov
arXiv:2609.37861v1 Announce Type: cross
Abstract: In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world's languages are spoken, a deployed text cl...
By Bhanu Prakash Vangala, Vangala Navya
arXiv:2609.35942v1 Announce Type: new
Abstract: Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs i...
By Ting-Chih Chen, Emile van Krieken, Shujian Yu, Filip Ilievski
arXiv:2609.36893v1 Announce Type: new
Abstract: Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy,...
By Zhenwen Ji, Lei Jin, Shanyong Wang, Jiaming Lu, Chengqiang Lu, Yi Wu, Yao Hu, Lizhen Cui, Yanyu Xu
arXiv:2609.37408v1 Announce Type: new
Abstract: Narrative extraction allows us to identify online hate narratives, supporting the construction of rigorous detection systems. Existing computational ap...
By Annabelle K. L. Chua, Forster J. Khoo, Joel C. R. Tan, Huey Ting Ang, Kheng Hwee Tan, Joel Y. A. Sim, Shirley W. H. Ow, Ria Mundhra, Elsie C. K. Toh, Youfeng Xu, Lynnette H. X. Ng
arXiv:2603.10861v2 Announce Type: replace
Abstract: SiDiaC-v.2.0 is the largest comprehensive Sinhala Diachronic Corpus to date, covering a period from 1800 CE to 1955 CE in terms of publication date...
By Nevidu Jayatilleke, Nisansa de Silva, Uthpala Nimanthi, Gagani Kulathilaka, Azra Safrullah, Johan Sofalas
arXiv:2609.36194v1 Announce Type: new
Abstract: Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate. To evaluate direction reproducib...
By Muhammad Abdullahi Said, Abass Oguntade, Elisha Komolafe, Babangida Sani, Fatima Muhammad Adam, Muhammad Sammani Sani
arXiv:2609.37543v1 Announce Type: new
Abstract: Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-reso...
By Prosper Arineitwe Asiimwe, Francois Meyer, Jan Buys