arXiv Computer Vision

CapsuleMotion: A Lightweight Real-Time Visual Motion Predictor for Capsule Endoscopy

CapsuleMotion is a lightweight, real‑time visual motion predictor designed for video capsule endoscopy (VCE). It predicts motion between successive frames using on‑device image compression metrics, allowing the capsule to adjust its frame rate dynamically and operate in a low‑power mode before entering the small intestine. Evaluated on the Rhode Island VCE dataset and deployed on an ultra‑low‑power RISC‑V demonstrator, CapsuleMotion reduces energy consumption by up to 20.66% and improves detection of the small intestine entry point.

arXiv Machine Learning
Jul 10

Precise localization within the GI tract by combining classification of CNNs and time-series analysis of HMMs

arXiv:2310. 07895v2 Announce Type: replace Abstract: This paper presents a method to efficiently classify the gastroenterologic section of images derived from Video Capsule Endoscopy (VCE) studies by exploring the combination of a Convolutional Neural Network (CNN) for classification with the time-series analysis properties of a Hidden Markov Model (HMM).

By Julia Werner, Christoph Gerum, Moritz Reiber, J\"org Nick, Oliver Bringmann
arXiv Computer Vision
Sep 3

UnCapsTSR: An Unsupervised Transformer-based Image Super-Resolution Approach for Capsule Endoscopy Images

UnCapsTSR is an unsupervised transformer-based GAN framework designed to enhance the spatial resolution of low‑resolution wireless capsule endoscopy (WCE) images. It eliminates the need for explicit degradation modeling or paired LR‑HR data by using a Bilateral Total Variation loss to preserve spatial continuity. The authors introduce a new Kvasir Capsule dataset for training, validate generalizability on KID and GIANA datasets, and propose the Endoscopy Quality Metric (EndoQM) as a non‑reference evaluation tool, reporting 40–80% improvement in EndoQM over state‑of‑the‑art unsupervised methods.

By Anjali Sarvaiya, Shubh Kawa, Lalit Agrawal, Jagrit Joshi, Kishor Upla, Kiran Raja
arXiv AI
Aug 19

Protect the Brain When Treating the Heart: Feasibility of 2.5D U-Net for Real-Time Gaseous Microemboli Detection

The study evaluates a 2.5D U‑Net model for detecting gaseous microemboli (GME) in real‑time during cardiac surgery using transesophageal echocardiography (TEE). On a pilot dataset of eight patients, the model achieved high precision (92.55%) and recall (80.54%) with an average inference time of 0.12 s per batch, outperforming classical spot detection and 2D U‑Net while maintaining real‑time speed. External validation on a GME‑negative dataset showed few false positives, supporting the model’s feasibility for real‑time GME segmentation.

By Andrea Angino, Ken Trotti, Diego Ulisse Pizzagalli, Rolf Krause, Tiziano Torre, Stefanos Demertzis
arXiv Machine Learning
Aug 5

Dual-domain U-Nets with embedded back projection operators for motion-resolved 4D CBCT reconstruction

arXiv:2608. 03430v1 Announce Type: cross Abstract: Four-dimensional cone beam CT (4D CBCT) is important for image-guided radiation therapy of thoracic cancers, but its use is limited by long scan times, causing high patient dose and motion/sparse-sampling artifacts.

By Ivo Herzig, Pascal Paysan, Daniel Barco, Marc Andr\'e Stadelmann, Frank-Peter Schilling, Igor Peterlik, Michal Walczak, Lijin Aryananda, Woo Sang Ahn, Rudolf Marcel F\"uchslin, Lukas Lichtensteiger
Hugging Face Trending Papers
Sep 10

CEM-TUDASR: Computationally efficient multi-modality transformer based unsupervised domain adaptive super-resolution approach

CEM‑TUDASR is a lightweight, unsupervised Transformer-based super‑resolution framework designed to enhance low‑resolution images from Wireless Capsule Endoscopy (WCE). It uses a domain‑adaptive degradation network to generate realistic WCE‑like low‑resolution images from high‑resolution conventional endoscopy data, enabling effective unpaired learning. The model incorporates Deep Attention Blocks and a Fusion Attention Block to capture both global context and fine local details, achieving superior performance on WCE datasets and demonstrating cross‑domain adaptability to retinal images, all while keeping the parameter count and computational load low.

arXiv Computer Vision
Sep 11

CEM-TUDASR: Computationally efficient multi-modality transformer based unsupervised domain adaptive super-resolution approach

CEM‑TUDASR is a lightweight, unsupervised Transformer‑based super‑resolution framework designed for Wireless Capsule Endoscopy (WCE) images. It uses a domain‑adaptive degradation network to synthesize realistic low‑resolution WCE images from high‑resolution conventional endoscopy data, enabling unpaired training. The SR generator incorporates Deep Attention Blocks and a Fusion Attention Block to preserve both global context and fine local structures, achieving superior no‑reference quality metrics and improved restoration of mucosal textures, vascular patterns, and anatomical details while remaining computationally efficient.

By Anjali Sarvaiya, Jay Kadel, Kishor Upla, Kiran Raja
arXiv Computer Vision
Sep 18

AI or Real: Detecting Partially Altered Videos Under Resource-Constrained Environments

The paper introduces a lightweight full-frame detector for partially manipulated AI-generated videos, suitable for edge deployment without face-detection preprocessing. It distills a DINOv2-Base teacher into a frozen MobileNetV3-Small student using temperature-annealed soft-label transfer, attention-diversity regularization, frame-level supervision, and a residual feature adapter. The model addresses false positives on legitimate scene cuts and threshold-level miscalibration, achieving an AUC of 0.766 on a 55,393-sample spliced test set while running at 3.65 ms per 16‑frame clip with a 150.4 MB checkpoint.

By Tamoghna Chakraborty, Md Nurul Absur, Sourya Saha, Saptarshi Debroy
arXiv Machine Learning
Sep 25

Albireo: Adaptive, Energy-Efficient Inference Framework for Video Object Detection on the Edge

Albireo is an adaptive, energy‑efficient inference framework for video object detection on edge devices that wraps existing detectors without modification. It uses a 10‑dimensional Kalman filter per active object to decide when to skip detector calls, predicting bounding boxes on skipped frames at near‑zero GPU cost. Evaluated on BDD100K with YOLO and RF‑DETR detectors on NVIDIA Jetson AGX Thor and Orin, Albireo maintains AP@50 within ±1.2 pp of full‑frame inference while reducing energy consumption by 12.1–17.6 % and improving accuracy for some models.

By Amir Taherin, Jos\'e Cano, Bin Ren, Yanzhi Wang, David Kaeli
arXiv Machine Learning
Jul 8

Reliable Mislabel Detection for Video Capsule Endoscopy Data

arXiv:2602. 06938v2 Announce Type: replace-cross Abstract: The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets.

By Julia Werner, Julius Oexle, Oliver Bause, Maxime Le Floch, Franz Brinkmann, Hannah Tolle, Jochen Hampe, Oliver Bringmann
arXiv Computer Vision
Sep 18

RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment

RAUL is a reference‑assisted reconstruction framework that recovers ureteroscope trajectories from endoscopic video alone, using a high‑quality reference exploration video for each phantom. It achieves a mean translation error of 0.5 mm and increases frame‑wise localization coverage from 50.5 % to 86.1 % compared to standard Structure‑from‑Motion. The reconstructed trajectories reveal significant differences in navigation metrics between high‑ and low‑experience trainees, enabling objective skill assessment without external tracking equipment.

By Fangjie Li, Mai Bui, Charan Mohan, Michael Miga, Matthieu Chabanas, Nicholas Kavoussi, Jie Ying Wu
arXiv AI
2d ago

STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets

STATERA is a method that adapts a pretrained video backbone with mostly frozen weights and a lightweight temporal tubelet mixer to estimate the center-of-mass (CoM) of opaque, asymmetric rigid bodies from short monocular videos. It introduces the HiddenMass Benchmark, consisting of 50K simulated MuJoCo trajectories and a 63-sequence real-world test set with calibrated CoM ground truth. In simulation, STATERA reduces normalized CoM error from 41.7% to 25.2%, and in zero-shot sim-to-real transfer, its phase‑aware variant consistently predicts movement toward the true hidden offset, improving physics capture from 2.6% to 41.0%.

By Animesh Varma