arXiv:2607. 03994v1 Announce Type: cross Abstract: Modern language models generally represent text as sequences of discrete token embeddings, an assumption deeply rooted in current practice but rarely questioned.
By Shuyang Xiang, Hao Guan
arXiv:2607. 03644v1 Announce Type: cross Abstract: Decades of orbital missions have produced multi-modal remote sensing data for the Moon, spanning optical imagery, spectroscopy, thermal emission, radar, gravity, and elemental composition.
By Ayush Prasad, Swarnalee Mazumder
arXiv:2506. 09740v2 Announce Type: replace-cross Abstract: Diffusion models excel at image generation.
By Qin Zhou, Zhiyang Zhang, Jinglong Wang, Xiaobin Li, Jing Zhang, Qian Yu, Lu Sheng, Dong Xu
arXiv:2607. 03653v1 Announce Type: cross Abstract: Traditional malware detection methods struggle to generalize to obfuscated or previously unseen threats.
By Allyson Taylor, Prashanth BusiReddyGari
arXiv:2607. 03819v1 Announce Type: cross Abstract: Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretation.
By Zhenyu Sun, Xiaohan Zhang, Qi Liu, Huan Wang
arXiv:2607. 03663v1 Announce Type: cross Abstract: The accurate estimation of Above-Ground Biomass (AGB) in mature tropical forests remains a critical challenge in remote sensing, primarily due to the saturation of Synthetic Aperture Radar (SAR) signals in high-density areas and persistent cloud cover affecting optical imagery.
By Luiz Felipe Parente Santiago (Institute of Computing, Brazilian Army Research Institute in the Amazon), Rosiane Rodrigues de Freitas (Institute of Computing), Daniel Rodrigues dos Santos (Military Institute of Engineering), Felipe Ferrari (Military Institute of Engineering)
arXiv:2607. 03283v1 Announce Type: new Abstract: Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal observations, robot states, human demonstrations, and task contexts into structured representations, decisions, trajectories, control references, and system services.
By Junwu Xiong, Jiaxuan Gao, Wei Chai, Renxing Chen, Yuzhen Li, Yu Guo, Yucheng Guo, Mingxi Luo, Wenyang Ma, Yiyun Mou, Yifei Zhang, Chen Zhou, Yongjian Guo
arXiv:2603. 13994v2 Announce Type: replace-cross Abstract: Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties.
By Hossein Adeli, Seoyoung Ahn, Andrew Luo, Mengmi Zhang, Nikolaus Kriegeskorte, Gregory Zelinsky
arXiv:2607. 02921v1 Announce Type: cross Abstract: Quantitative 3D spatial reasoning from egocentric RGB-D video is a critical capability for next-generation wearable assistants.
By Maxwell Horton, Wei Lu, Quan Tran, Yury Astashonok, Kirmani Ahmed, Babak Damavandi, Anuj Kumar, Xiao Zhang, Seungwhan Moon
arXiv:2607. 04153v1 Announce Type: cross Abstract: Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information.
By Kai Zhao
arXiv:2606. 22574v2 Announce Type: replace-cross Abstract: While synthetic data generation resolves the manual labeling bottleneck in computer vision, minimizing the syn-to-real domain gap requires optimizing rendering variables.
By Hooman Tavakoli Ghinani, Tatjana Legler, Martin Ruskowski
arXiv:2505. 08814v3 Announce Type: replace-cross Abstract: Deep neural networks (DNNs) play a crucial role in the field of artificial intelligence, and their security-related testing has been a prominent research focus.
By Wenkai Li, Xiaoqi Li, Yingjie Mao, Yishun Wang
arXiv:2607. 04262v1 Announce Type: new Abstract: Convolutional Neural Network (CNN) and Vision Transformer (ViT) for image classification exploit a dense grid of pixels containing redundant information.
By Sarabeshwar Balaji, Shubham Mohanty, Akash Anil
arXiv:2607. 03650v1 Announce Type: cross Abstract: Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a major challenge in modern electronic health records (EHRs).
By Enshuo Hsu, Jin Zhou, Kirk Roberts
arXiv:2607. 03836v1 Announce Type: cross Abstract: Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natural Language Processing (NLP) capabilities: low-resource transliteration, archaic vocabulary, and noisy input signals.
By Nguyen Kim Hai Bui, Md. Easin Arafat, Tam\'as G\'abor Orosz, Mufti Mahmud
arXiv:2607. 03196v1 Announce Type: cross Abstract: WiFi-based human pose estimation (HPE) enables the detection and interpretation of human body positions and movements without the need for wearable devices while preserving individual privacy concerns.
By Toan D. Gian, Van-Dinh Nguyen, Vo Phi Son, Nhan Thanh Nguyen, Dinh Thai Hoang, Diep N. Nguyen, Nguyen Cong Luong, Symeon Chatzinotas
arXiv:2607. 01983v1 Announce Type: cross Abstract: Robust 3D object detection under adverse weather remains a critical hurdle for autonomous driving.
By Shuyao Li, Chuanxing Geng, Heyang Sun, Qiang Zhou, Jingjing Gu
arXiv:2607. 02611v1 Announce Type: cross Abstract: Work-related Musculoskeletal Disorders (WMSDs) require continuous ergonomic assessments.
By Xuhan Zhang, Zhuangzhuang Dai, Luis J. Mans, Victor Chang
arXiv:2603. 07866v3 Announce Type: replace-cross Abstract: Offshore inspection and maintenance have increasingly been using legged robots for routine sensing, yet many useful interventions still require physical interaction with tools, containers, and task-relevant objects.
By Dilermando Almeida, Juliano Negri, Guilherme Lazzarini, Thiago H. Segreto, Ranulfo Bezerra, Gustavo J. G. Lahr, Ricardo V. Godoy, Marcelo Becker
arXiv:2607. 04921v1 Announce Type: cross Abstract: Deep learning algorithms are notorious for their high carbon footprint and computational demands that limit their deployment on edge devices and raise concerns about their long-term sustainability.
By Manish Kolachalam, Rani Malhotra