arXiv:2607. 11801v1 Announce Type: cross Abstract: Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content.
By Yu-Han Huang, Chih-Kai Yang, Ke-Han Lu, An-Yu Cheng, Hung-yi Lee
arXiv:2607. 11643v1 Announce Type: cross Abstract: Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints.
By Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
arXiv:2607. 11081v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results.
By Sunyoung Jung, Jiwoo Park, Yoonseok Choi, Kyobin Choo, Ming-Hsuan Yang, Seong Jae Hwang
arXiv:2607. 10617v1 Announce Type: new Abstract: The lineage graph of open-weight language models is self-reported: Hugging Face's base_model metadata field is optional and unverified, and over 60% of Hub models document no parentage at all.
By Muhammad Awais Bin Adil, Saad Aamir
arXiv:2607. 11177v1 Announce Type: new Abstract: In this paper, we propose deep learning based NeuroMem-FHP framework for estimating the parameters of the fractional Hawkes process (FHP), a self-exciting point process that captures long-range dependence through a fractional Mittag-Leffler excitation kernel.
By Neha Gupta, Aditya Maheshwari
arXiv:2607. 09822v1 Announce Type: cross Abstract: Recognition tells an agent what is in an image; personal memory affects what is worth looking up next.
By Xiaofan Wu (Chance AI), Xi Zeng (Chance AI), Miaoxia Chen (Chance AI), Peishan Chen (Chance AI), Shuyan Li (Chance AI), Jiyun Yao (Chance AI), Hanyong Zhong (Chance AI), Jiahao Zhu (Chance AI)
arXiv:2607. 10057v1 Announce Type: cross Abstract: Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost?
By Dongping Liu, Aoyu Zhang, Luyao Zhang
arXiv:2607. 11193v1 Announce Type: cross Abstract: To ensure the overall quality of AI-enabled software, not only traditional software components but also AI components need to be tested and repaired.
By Yuta Ishimoto, Paolo Arcaini, Fuyuki Ishikawa, Masanari Kondo, Naoyasu Ubayashi, Yasutaka Kamei
arXiv:2607. 10784v1 Announce Type: cross Abstract: Deploying deep learning models for automated electrocardiogram classification on resource-constrained wearable devices remains challenging due to high computational costs.
By Yi Zhao, Jiajun Gao, Chenyang Xu, Yuxi Zhou, Hao Wang
arXiv:2607. 11138v1 Announce Type: new Abstract: The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thousands of options simultaneously.
By Prashant Devadiga, Abhishek, Adithya Mishra, Alok Singh, Amisha Sinha, Asit Desai, Gaurang Dahad, Harshit Bhushan, Mandati Pramod Reddy, Prakhar Gupta, Rupesh Patil, Siddhi Behere
arXiv:2607. 09576v2 Announce Type: replace-cross Abstract: We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse languages, totaling 160 conventional expressions, the large majority of which are idiomatic.
By Kiran Pala, Punam Silu, Luxin Yu
arXiv:2601. 22146v2 Announce Type: replace-cross Abstract: Due to limited supervised training data, large language models (LLMs) are typically pre-trained via a self-supervised "predict the next word" objective on a vast amount of unstructured text data.
By Ajay Patel, Colin Raffel, Chris Callison-Burch
arXiv:2607. 10021v1 Announce Type: new Abstract: Neural networks can learn algorithmic input-output mappings, but trusting a learned executor requires more than a correct final answer because the state transitions that produce it are usually hidden.
By Jose Luis Lima de Jesus Silva
arXiv:2604. 03922v2 Announce Type: replace Abstract: Selecting LLM-generated code candidates using LLM-generated tests is challenging because the tests themselves may be incorrect.
By Hui Sun, Yun-Ji Zhang, Zheng Xie, Ren-Biao Liu, Yali Du, Xin-Ye Li, Ming Li
arXiv:2509. 24148v3 Announce Type: replace-cross Abstract: Test-Driven Development (TDD) is a widely adopted practice that requires developers to create and execute tests alongside implementation.
By Yiran Hu, Nan Jiang, Shanchao Liang, Yi Wu, Lin Tan
arXiv:2607. 11734v1 Announce Type: cross Abstract: Differentiable simulators have advanced policy learning and model-based control, yet actuator dynamics remain an important source of sim-to-real error.
By Zhiyang Dou, John U. Onyemelukwe, Hangxing Zhang, Heng Zhang, Minghao Guo, Yunsheng Tian, Michal Piotr Lipiec, Joshua Jacob, Chao Liu, Peter Yichen Chen, Yuri Ivanov, Wojciech Matusik
arXiv:2607. 11342v1 Announce Type: cross Abstract: Despite their central role in fault detection, test oracles remain challenging to construct effectively.
By Yue Zhao, Binish Tanveer, Jelena Zdravkovic
arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.
By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
arXiv:2603. 10725v3 Announce Type: replace-cross Abstract: The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information.
By Artem Dvirniak, Evgeny Kushnir, Dmitrii Tarasov, Artem Iudin, Oleg Kiriukhin, Mikhail Pautov, Dmitrii Korzh, Oleg Y. Rogov
arXiv:2604. 22770v2 Announce Type: replace-cross Abstract: Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency.
By Nicy Scaria, Silvester John Joseph Kennedy, Deepak Subramani