arXiv:2608. 06165v1 Announce Type: cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored.
By Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoro Mo, Yaolong Ju
arXiv:2608. 05727v1 Announce Type: cross Abstract: Neural Audio Codecs are widely adopted in speech generation and editing.
By June Young Yi, Dongwook Lee, Jiheum Yeom, Sungroh Yoon
arXiv:2608. 06122v1 Announce Type: cross Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series.
By Omar Coser, Antonio Orvieto, Paolo Soda, Loredana Zollo
arXiv:2608. 05242v1 Announce Type: new Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training.
By Haoze Sun, Jiequan Cui, Qingshan Xu, Richang Hong
arXiv:2608. 05595v1 Announce Type: cross Abstract: Circuit cutting lets a large quantum neural network (QNN) run as independent subcircuits on small devices, but rebuilding its outputs by reconstruction carries a classical sampling overhead exponential in the number of cuts - the dominant runtime cost in prior work.
By Prabhjot Singh, Adel N. Toosi, Rajkumar Buyya
arXiv:2608. 06377v1 Announce Type: cross Abstract: Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong.
By Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong
arXiv:2608. 05728v1 Announce Type: cross Abstract: Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval.
By Feiyu Ji, Xiang Li, Hao Ma, Tianxiang Huang, Qingxin Lu, Mengqi Ji, Lei Han, Xiaokang Yang, Xiaoyun Yuan
arXiv:2608. 05782v1 Announce Type: cross Abstract: Wearable human activity recognition (HAR) is often limited by the scarcity of labeled sensor data, especially in low-resource, class-imbalanced, and subject-generalization settings.
By Lala Shakti Swarup Ray, Vitor Fortes Rey, Mengxi Liu, Paul Lukowicz, Bo Zhou
arXiv:2608. 05611v1 Announce Type: cross Abstract: Large Language Models (LLMs) can exhibit diverse personas, and activating expert personas has been shown to improve domain expertise and task accuracy.
By Guanyu Wang, Zidi Zhang, Xu Chu
arXiv:2608. 05166v1 Announce Type: cross Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings.
By Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs
arXiv:2602. 15649v2 Announce Type: replace Abstract: In dynamical systems reconstruction (DSR) we aim to recover the dynamical system (DS) underlying observed time series.
By Alena Br\"andle, Lukas Eisenmann, Florian G\"otz, Daniel Durstewitz
arXiv:2608. 05103v2 Announce Type: replace Abstract: Data assimilation (DA) uses Bayesian inference to update the state of a numerical forecast model with observed data.
By Dibyajyoti Chakraborty, Romit Maulik
arXiv:2608. 05673v1 Announce Type: new Abstract: Trajectory prediction has shifted toward structured formulations with explicit social modeling.
By Jiaheng Chen, Jiaxing Li, Tinghe Zhang, Chaopeng Guo
arXiv:2601. 21124v2 Announce Type: replace-cross Abstract: Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI.
By Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar
arXiv:2608. 05702v1 Announce Type: new Abstract: Scientific machine learning commonly validates models at the level of a subdomain, a benchmark split, or an explanation for one prediction.
By Gnankan Landry Regis N'guessan, Bum Jun Kim
arXiv:2608. 06253v1 Announce Type: new Abstract: Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations.
By Dohyun Ku, Min Gu Kwak, Francisco J. Pasquel, Jing Li
arXiv:2608. 06310v1 Announce Type: new Abstract: Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models.
By Chenglong Wang, Ziming Zhu, Yifu Huo, Bei Li, Qiaozhi He, Yan Ding, Xiaoyang Hao, Yuxin Gao, Tianhua Zhou, Xiaojia Chang, Tongran Liu, Jingbo Zhu
arXiv:2608. 05502v1 Announce Type: cross Abstract: In this paper, we consider a class of multiblock nonconvex nonsmooth optimization problems, which covers many applications such as the analysis of pre-earthquake anomalies and machine learning.
By Weifeng Yang
arXiv:2608. 05249v1 Announce Type: cross Abstract: Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one self-contained question.
By Xiaomin He, Dongling Xiao, Jiahao Xie, Ruiqi Lu, Qianle Wang, Zhongbin Guo, Wanxuan Sun
arXiv:2608. 06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness.
By Ro Encarnaci\'on, Tina Behzad, Emma Lurie, Dana\'e Metaxa