PolarScale is a new benchmark that explicitly requires models to predict the radiometric scale needed for full Stokes reconstruction from a per-scene normalized total‑intensity image. It evaluates models on normalized Stokes components, AoLP/DoLP/DoCP, and a per‑scene scale, using metrics that include angular, self‑consistency, and physical‑bound checks. Across seven restoration‑based and generative backbones, the best restoration models achieve a 3.6‑4.3% mean relative error in scale estimation, outperforming a constant‑scale control and maintaining physical plausibility.
By Beibei Lin, Tingting Chen, Xin Zhang, Wenhao Zhao, Dongjun Li, Zifeng Yuan
The same Gaussian of a 3D Gaussian Splatting model is seen from many views, and these views do not always agree on the class it belongs to. The Gaussian may be occluded in some of them, and the confid...
Surface electromyography (sEMG) provides a wearable, noninvasive interface to neuromuscular activity for movement decoding and human-computer interaction. Population-scale decoding remains difficult b...
The article titled "Computer Vision: SIFT algorithm (Scale Invariant Feature Transform)" discusses the SIFT algorithm, a method for matching objects across different viewpoints. It highlights the elegance of this approach in handling variations in scale and orientation. The post was originally published on Towards Data Science.
By Slava Efimov
The paper introduces SciKGExtract, a schema-guided framework that uses large‑language‑model extraction, chemical normalization, and agent‑based evaluation to convert heterogeneous materials science literature into a knowledge graph. Applied to 176 atomic‑layer‑deposition papers on ZnO and IGZO, the system improves extraction F1 from 0.591 to 0.805 for ZnO with agentic refinement, while IGZO remains more challenging at 0.344. Evaluation against a detailed schema of 65 experimental properties and 155 quantitative nodes reveals segmentation and numerical assignment errors, highlighting the complementary role of chemical canonicalization and agentic verification in producing machine‑actionable experimental knowledge.
ZeBROD is a zero‑retraining framework that tackles catastrophic forgetting in object detection by combining YOLO11n for localization with DeIT and Proxy Anchor Loss for feature extraction. It classifies products using cosine similarity against embeddings stored in a Qdrant vector database, enabling accurate detection of both new and existing items without retraining. In a retail store experiment with 140 products, ZeBROD achieved high accuracy and nearly three times faster training than traditional methods, while maintaining an inference time of 580 ms per image on an edge device.
By Priyanto Hidayatullah, Nurjannah Syakrani, Yudi Widhiyasana, Muhammad Rizqi Sholahuddin, Refdinal Tubagus, Zahri Al Adzani Hidayat, Hanri Fajar Ramadhan, Dafa Alfarizki Pratama, Farhan Muhammad Yasin
The paper presents an evolutionary framework, EXAQC, that automatically discovers parameterized quantum circuits (PQCs) for use as intermediate modules in hybrid quantum‑classical neural networks for image classification. By evolving PQCs while keeping classical feature‑extraction and prediction layers fixed, the authors achieve high accuracies on MNIST, Fashion‑MNIST, and CIFAR‑10 with gate counts comparable to other quantum architecture‑search methods. The evolved hybrid models match or exceed the performance of classical networks while using far fewer trainable parameters, and the choice of encoding (rotation‑based vs amplitude) significantly impacts accuracy.
By Devroop Kar, Daniel Krutz, Travis Desell
VIDiff is a unified foundation model that uses diffusion techniques to perform a broad range of video tasks, including both understanding tasks like language‑guided video object segmentation and generative tasks such as video editing and enhancement. Unlike prior methods that focus on short clips and require time‑consuming tuning, VIDiff can edit and translate videos within seconds based on user instructions and employs an iterative auto‑regressive approach to maintain consistency in long‑form videos. The authors demonstrate convincing generative results across diverse input videos and written instructions, supported by qualitative and quantitative evidence.
By Zhen Xing, Shuyuan Tu, Qi Dai, Zihao Zhang, Hui Zhang, Han Hu, Zuxuan Wu, Yu-Gang Jiang
The paper introduces HypoDepth, an event-image monocular depth estimation framework that uses a discrete Depth Hypothesis Volume (DHV) to convert depth regression into a constrained search problem. By building a lightweight 3D cost volume between DHV features and contextual features, the method performs multi-scale correlation search for stable residual optimization, enabling efficient global-to-local refinement across resolutions. Experiments on DSEC and MVSEC show state‑of‑the‑art performance, strong zero‑shot generalization, and real‑time capability on resource‑limited devices.
By Daikun Liu, Teng Wang, Changyin Sun
UniDynamics is a diffusion-based framework that generates future 4D dynamic scenes—comprising RGB, depth, and optical flow—from a single event-RGB pair, without needing long histories or control priors. It introduces an Event Latent Enhancement module to align event data into conditioning features and a Perceptual Dynamics Space within a multi-scale U‑Net to decouple and adaptively interact depth and flow, enforcing geometric and motion constraints for coherent predictions. Experiments on VKitti2 and DSEC show state‑of‑the‑art performance, producing high‑quality, temporally coherent, and 4D‑consistent predictions even under high‑speed motion blur.
By Daikun Liu, Xin Zhan, Teng Wang, Xiaoping Wang, Changyin Sun
ForestQuery is a new framework for unified forest point cloud segmentation that incorporates boundary-aware and spatially anchored query learning. It explicitly models boundary uncertainty to improve instance query construction and uses learnable 3D anchors to encode forest vertical stratification for semantic queries. Experiments on public benchmarks and a real‑world dataset show consistent gains in both individual‑tree and semantic segmentation.
By Zhihao Zhan, Le Tao, Yifei Tian, Xin Liu, Jie Yuan
The paper presents a three‑stage pipeline for colon segmentation in 3D abdominal CT scans that preserves anatomical continuity. First, a deep‑learning model generates an initial segmentation; second, centreline bridging reconnects disjoint regions; third, a reconstruction stage refines the continuity. Experiments on the TotalSegmentator and RAOS datasets show that the method improves topological consistency while keeping segmentation accuracy high.
By Deshan Kalupahana, Sonit Singh, Praveen Ravindran, Arcot Sowmya
XGenAct is a world action model that encodes RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos using deterministic codecs. It trains a single video diffusion transformer with a unified objective, sampling different perception and action tasks to learn temporal prediction across multiple spatial modalities without separate heads. Experiments on RLBench show that structured perception training boosts closed‑loop success, with XGenAct achieving 52% success on five external tasks compared to 26% for the best baselines, and it outperforms pipelines that first generate RGB and then apply a frozen perception expert for depth and segmentation prediction.
By Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li
The paper investigates whether uncertainty can act as a proxy for semantic correctness in diffusion-based medical image synthesis, specifically for generating contrast‑enhanced CT (CECT) from non‑contrast CT (NCCT). Using the AortaDiff framework, which produces both CECT images and lumen segmentations, the authors compare six uncertainty estimation methods across pixel, region, and image levels, including their ability to detect out‑of‑distribution cases. They find that uncertainty is informative at all spatial scales, remains useful under distribution shift, and that MCDropout, in particular, offers reliable quality filtering with no extra training cost.
By Yuxuan Ou, Konstantinos Kamnitsas, OxAAA Study, AICT Consortium, Regent Lee, Vicente Grau
S2S-JEPA is a new AI weather model that focuses on predicting only the slowly varying components of the atmosphere at subseasonal-to-seasonal timescales, from two weeks to two months ahead. It adapts the Joint-Embedding Predictive Architecture (JEPA) from computer vision to discard unpredictable fine-scale details, thereby addressing the long‑known predictability desert. The model matches the skill of the ECMWF physics‑based ensemble and even outperforms it on several metrics during weeks 5 to 6.
By Chenyu Dong, Gianmarco Mengaldo
BeeWhere is an AI-assisted workflow that merges ArUco fiducial detections with deep‑learnt instance segmentations to analyze bumble bee colonies. Using high‑resolution images of Bombus impatiens microcolonies, the system annotates thousands of bee instances, pollen balls, nest structures, and chamber boundaries, and trains YOLO models for behavioral metrics such as nearest‑neighbor distance and spatial occupancy. In a validation study, BeeWhere outperformed tag‑based tracking, especially under occlusion, and revealed pesticide‑induced changes in bee spatial organization that tag‑based methods missed.
By Roberta Hunt, August Easton-Calabria, James Crall
VisionMX introduces a post‑training microscaling (MX) quantization technique for vision models, addressing three key error sources identified in direct conversion: block‑scale representation, misalignment of small convolutional weight tensors, and underuse of signed codes for nonnegative activations. The method optimizes bounded weight rounding and incorporates a foldable affine correction for activations, improving performance across image classification, object detection, semantic segmentation, and low‑light image enhancement tasks. VisionMX outperforms direct conversion and other post‑training quantization baselines, especially in architectures most sensitive to MX conversion.
By Elad Dror Cohen, Ofir Gordon, Lior Dikstein, Idan Achituve, Hai Victor Habi
The paper introduces a fully automatic pipeline for 3D dendrite instance segmentation in serial block-face scanning electron microscopy (SBF-SEM). It combines YOLOv6-guided Segment Anything Model prompting, iterative 2D mask refinement, random forest 3D linking, and high‑resolution nnU‑Net refinement to produce coherent, well‑separated dendrite reconstructions without manual prompting. Applied to hippocampal CA1 data from a control and an epileptic rat, the method achieves high semantic accuracy (Dice 0.93/0.91) and strong instance performance, while highlighting dense‑region recognition as the main challenge in epileptic tissue.
By Zewen Zhuo, Ilya Belevich, Eija Jokitalo, Alejandra Sierra, Jussi Tohka
RYOPO is an end‑to‑end query‑based RGB‑D set predictor that jointly detects, segments, and estimates 9‑DoF poses of unseen instances within known categories without relying on external instance segmentation or CAD priors. It uses shared image and scene encoding, a query‑conditioned geometry pathway, and object‑centric refinement with pose‑conditioned cross‑attention to achieve accurate pose estimation. On benchmark datasets such as NOCS, REAL275, and HouseCat6D, RYOPO outperforms published methods and runs in real time at 31.8 FPS on an RTX A6000.
By Hakjin Lee, Junghoon Seo, Jaehoon Sim
DeepStratNet introduces a context‑aware coordinate regression framework for seismic horizon tracking that directly predicts time/depth coordinates at each lateral position, avoiding the need for dense segmentation masks. The lightweight regression head, combined with an LSTM for inter‑slice context and geology‑informed regularization, outperforms traditional segmentation models on a New Zealand seismic volume, especially under sparse labeling. The method also provides a built‑in quality control by capturing local geological variations through prediction variability across traces.
By Aniq Ahmad, Musham Ahmad Malik, Ahmad Mustafa, Heather Bedle