arXiv:2608. 04479v1 Announce Type: cross Abstract: Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions.
By Jinting Wang, Yuguang Yang, Shengyu Li, Yan Rong, Shan Yang, Xiaoda Yang, Li Liu
arXiv:2608. 04144v1 Announce Type: cross Abstract: Biomedical entity linking grounds mentions in clinical and scientific text to entities in a curated knowledge base (KB) with ontological structure, which supports downstream applications such as literature-scale information extraction and patient-record normalization.
By Yicheng Tao, Jie Liu
arXiv:2608. 04777v1 Announce Type: cross Abstract: Oscillatory signals, such as vibration, carry class-discriminative information in specific frequency bands; perturbing them in raw feature space for counterfactual analysis easily destroys their temporal structure and produces physically implausible results.
By Udo Schlegel, Julian Rakuschek, Thomas Seidl, Andreas Holzinger, Tobias Schreck, Javier Del Ser
arXiv:2608. 04840v1 Announce Type: cross Abstract: Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence.
By Jacob Arndt, Debvrat Varshney, Philipe Dias, Nivedita Nukavarapu
arXiv:2608. 05054v1 Announce Type: cross Abstract: We investigate the transferability of Earth weather foundation models to planetary atmospheres by adapting the GraphCast graph neural weather forecasting model to Mars.
By M. L. Carroll, J. Li, S. D. Guzewich, G. Villanueva, J. A. Caraballo-Vega, M. J. Frost
arXiv:2608. 04043v1 Announce Type: new Abstract: Resistive pressure arrays are the cheapest and most widely shipped tactile sensors, yet tactile representation learning has concentrated on optical sensors that image a deforming gel.
By Abdul Basit Tonmoy
arXiv:2603. 12408v3 Announce Type: replace-cross Abstract: Motion imitation learning (IL) is increasingly used in robotics and human gait modeling, yet its ability to recover biomechanically consistent joint moments without explicit kinetic information remains unclear.
By Xinyi Liu, Jangwhan Ahn, Edgar Lobaton, Jennie Si, He Huang
arXiv:2510. 21084v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong potential for clinical decision support through their advanced language understanding and reasoning capabilities.
By Juntao Li, Haobin Yuan, Ling Luo, Yuanyuan Sun, Jian Wang, Hongfei Lin
arXiv:2602. 13357v3 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art performance in high-fidelity image and video generation but suffer from expensive inference due to their iterative denoising structure.
By Dong Liu, Yanxuan Yu, Ben Lengerich, Ying Nian Wu
arXiv:2608. 05095v1 Announce Type: new Abstract: Agents for long term reasoning require a memory that can be efficiently and effectively updated over time, as new facts and external feedback continue to arrive.
By Xiawei Yue, Boran Wang, Xiaoqing Zhang, Shuxin Zheng, Ziwei Zhang
arXiv:2608. 03487v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG), which augments large language model (LLM) generation with information retrieved from databases, has become a widely used approach for knowledge-intensive applications.
By Haiqiang Zhang, Yuanqing Lei, Wanting Li, Tao Zhang, Wenqi Jiang
arXiv:2608. 04028v1 Announce Type: cross Abstract: Echo-state networks enable efficient temporal learning by fixing the recurrent dynamics and training only a linear readout.
By Jyotiranjan Beuria, Amit Shukla
arXiv:2608. 04124v1 Announce Type: cross Abstract: Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time.
By Haotian Xia, Zilin Xiao, Junbo Zou, Vicente Ordonez, Hanjie Chen
arXiv:2608. 04384v1 Announce Type: new Abstract: Neural PDE solver auto-design is fundamentally a search-space representation problem.
By Shengxin Kong, Liwen Xu, Jingwen Fu
LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowledge indicates what evidence to inspect and which failures are task-critical.
Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual searches. To overcome this limitation, we introduce Contextual Composed Image Retrieval (CoCo-IR), a novel task that enables users to progressively refine search results through interactions.
Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resources and often scarce in long-distance dependencies.
Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose steps require different reasoning skills and depend on earlier outputs.
Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models.
Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognition from CCTV-style observations.