Visual Perception Engine (VPEngine) is a modular framework that enables efficient GPU usage for robotic vision tasks by sharing a foundation model backbone across multiple specialized task heads. It eliminates redundant feature extraction, supports dynamic task prioritization, and achieves up to 3× speedup over sequential execution. The open‑source Python implementation, with ROS2 C++ bindings, delivers real‑time performance (≥50 Hz) on NVIDIA Jetson Orin AGX using TensorRT‑optimized models.
By Jakub {\L}ucki, Jonathan Becktor, Georgios Georgakis, Rob Royce, Shehryar Khattak
This paper provides a detailed overview of the Ultralytics YOLO family from YOLOv5 to YOLO27, highlighting key architectural changes, benchmarking results, and deployment considerations. It discusses the evolution of each version—YOLO27’s scale‑adaptive dual architecture, YOLO26’s loss and optimization refinements, YOLO11’s efficiency focus, YOLOv8’s anchor‑free detection, and YOLOv5’s modular ecosystem—alongside performance metrics on COCO and latency on TensorRT. The review also surveys applications in robotics, agriculture, surveillance, and manufacturing, and outlines future challenges such as dense scene handling, CNN‑Transformer integration, and hardware‑aware optimization.
By Ranjan Sapkota, Manoj Karkee
VISTA is an online self‑distillation framework that enforces consistency along a deep learning model’s optimization trajectory. It uses a validation‑informed Marginal Coverage score to identify earlier model states—called expert anchors—that retain specialized competence over distinct data regions. By integrating a coverage‑weighted ensemble of these anchors during training, VISTA regularizes the loss landscape, preserves learned knowledge, and improves robustness and generalization while cutting storage overhead by 90%.
By Eli Corn, Daphna Weinshall
The paper introduces HTML‑LM, a 154‑million‑parameter foundation model designed for Czech HTML documents. It leverages HTML‑aware training and a ModernBERT architecture, trained on 100 million web pages with objectives such as masked language modeling, bag‑of‑words prediction, and contrastive distillation from larger language models. HTML‑LM achieves state‑of‑the‑art performance on classification and regression tasks in the Czech Internet domain, outperforms larger encoders and small LLMs, and is deployed in production to process thousands of web documents per second.
By Martin Dvo\v{r}\'ak, V\'it Tlusto\v{s}, Artyom Voronin, Martin Habrovec, Kate\v{r}ina Podlesn\'a, Barbora Ri\v{s}ov\'a, Josef Von\'a\v{s}ek
The thesis presents AI frameworks that accelerate crystalline materials discovery by tackling both crystal property prediction and crystal structure generation. It introduces CrysXPP, CrysGNN, and CrysMMNet for efficient, data‑sparse property prediction using graph autoencoding, self‑supervised pretraining, and multimodal learning. For generation, TGDMat is a text‑guided diffusion model that jointly learns lattice parameters, atomic types, and coordinates, enabling valid, stable, and conditionally generated periodic materials.
By Kishalay Das
The paper introduces a novel technique called "persistence of memory" to enhance stochastic subspace methods for large‑scale optimisation. By using a weakly correlated guidance vector that is refreshed only at wide intervals, the method provides a structured direction for random subspace descent. The authors demonstrate that this guidance can be efficiently computed in sparse or minibatch settings and present the first theoretical analysis of classical SSD methods for sparse functions, showing alignment with low‑lying Hessian eigenvectors near the optimum.
By Subhroshekhar Ghosh, Clement Z. Q. Ng, Pierre-Louis Poirion, Akiko Takeda
The paper presents Real‑Time EXPO‑FT, a reinforcement learning framework that fine‑tunes large Vision‑Language‑Action models for real‑time robotic control. It separates slow, expressive action generation from fast, reactive edits, allowing a lightweight policy to adjust actions based on the latest observation. Experiments on the Kinetix benchmark and four dynamic real‑world tasks show that Real‑Time EXPO‑FT achieves superior performance, improving policy success rates from 42% to 97% with only ten minutes of online data and no human intervention.
By Perry Dong, Kuo-Han Hung, Dorsa Sadigh, Chelsea Finn
arXiv:2509.05909v3 Announce Type: replace-cross
Abstract: The reliable identification of magnetic ground states remains a major challenge in high-throughput materials databases, where density functio...
By Ahmed E. Fahmy
arXiv:2512.22143v2 Announce Type: replace-cross
Abstract: Existing Wi-Fi sensing systems rely on injecting high-rate probing packets to extract channel state information (CSI), leading to communicati...
By Gaofeng Dong, Kang Yang, Mani Srivastava
arXiv:2609.18516v1 Announce Type: new
Abstract: While Large Language Models excel in natural language processing, efficiently extending their capabilities to spoken input remains a significant challe...
By Abderrahmane Issam, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis
arXiv:2609.18529v1 Announce Type: new
Abstract: UNESCO considers the Assyrian (Syriac) language an endangered language. Although Assyrians speak the language worldwide, the speaking population is unc...
By Hadiana Sliwa, Hossein Hassani
The paper introduces a reparameterization technique that injects feature noise to jointly optimize speech model performance and computational complexity during training. Unlike traditional pruning, this method dynamically adjusts model size for a desired performance‑complexity trade‑off without heuristic weight removal. The authors validate their approach with a synthetic example and two real‑world applications—voice activity detection and audio anti‑spoofing—providing publicly available code for further research.
By Esteban G\'omez, Tom B\"ackstr\"om
The paper introduces DASH-Q, a post‑training quantization method that uses a diagonal Hessian approximation and iterative weighted least squares to reduce noise in curvature estimates. By discarding noisy cross‑channel dependencies, DASH‑Q preserves salient feature power and achieves superior performance in ultra low‑bit quantization. Across five large language models, it improves zero‑shot accuracy by an average of 7.01% and up to 14.01% over the strongest baselines, even with very small calibration datasets.
By Jaemin Kim, Sungkyun Kim, Junyeol Lee, Jiwon Seo
PhysVGGT is a feed‑forward model that predicts dense maps of friction coefficient, Shore hardness, Young's modulus, and density, along with object‑level mass, from a single RGB image in one forward pass. It treats physical property estimation as a dense per‑pixel prediction problem, using a visual geometry transformer to extract geometry‑aware tokens and separate dense and global prediction branches. A scalable pseudo‑label generation pipeline enables large‑scale weakly supervised training, and the model achieves state‑of‑the‑art performance on the ABO‑500 dataset while running 27× faster than previous methods.
By Sneha Paul, Guile Wu, Bingbing Liu, Dongfeng Bai
arXiv:2609.17548v1 Announce Type: new
Abstract: Myovox, from myo (muscle) and vox (voice), decodes open-vocabulary English text from 31-channel surface electromyography (sEMG) recorded from the muscl...
By Varshith Madishetty
arXiv:2609.18935v1 Announce Type: new
Abstract: A game character should not have to reread its entire life before every conversation. For locally deployed language-model characters, however, revising...
By Zimu Xu
arXiv:2410.22229v2 Announce Type: replace-cross
Abstract: Offloading stateful network functions to multi-threaded SoC SmartNICs promises significant performance and cost benefits. However, realizing...
By Shaoke Xi, Jiaqi Gao, Fuliang Li, Minlan Yu, Ennan Zhai
arXiv:2609.18227v1 Announce Type: new
Abstract: Wildfire smoke detection from satellite imagery is critical for early warning and rapid response. For onboard satellite deployment, detection systems m...
By Sha Lu, Yu Sun, Liang Zhao, Jixue Liu, Lin Liu, Jiuyong Li, A. K. Qin, Alejandro Mousist, Stefan Peters
arXiv:2609.18239v1 Announce Type: new
Abstract: Structured pruning is commonly formulated as ranking individual channels, although channel responses can be complementary or cancel through downstream...
By Kaixiang Shu
arXiv:2609.18955v1 Announce Type: new
Abstract: Efficient perception models are essential for real-time autonomous driving, where accuracy and computational cost must be carefully balanced. However,...
By Huy Che, Minh-Khoi Do, Dinh-Duy Phan, Duc-Khai Lam