arXiv Computer Vision

From the Drosophila Visual Connectome to General-Purpose Computer Vision

arXiv Computer Vision
Sep 11

Fast and Accurate Monomodal 3D High Resolution Deep Registration of Drosophila Larval Brain Volumes

The paper introduces a deep learning pipeline that rapidly and accurately registers 3D high‑resolution Drosophila larval brain volumes to a shared anatomical reference. Unlike traditional methods that require per‑case optimization and minutes per brain, the trained network performs a single forward pass, handling volumes with many more voxels and maintaining high accuracy even as image quality declines. The authors benchmarked their approach against eleven classical and seven learned baselines, achieving a 23‑percentage‑point improvement in landmark‑based mutual information and registering brains one to two orders of magnitude faster.

By Daniel Reisenb\"uchler, Yousef Sadegheih, Michael Dittrich, Pratibha Kumari, Muhammad Usman, Dorit Merhof
arXiv AI
Aug 11

Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation

arXiv:2608. 08713v1 Announce Type: cross Abstract: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges.

By Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein
arXiv Machine Learning
1d ago

HAND: A Biologically-Inspired Activation Function that Improves Generalisation and Sample Efficiency in Image Classification

The paper introduces HAND, a biologically-inspired activation function that incorporates homeostasis, accelerating nonlinearity, and divisive normalization to act as an inductive bias in deep neural networks. Experiments on image classification show that using HAND allows a ConvNeXt-tiny model to reach ImageNet1k accuracy in 25 epochs versus 200 epochs for the baseline, and yields larger accuracy gains on long-tailed and reduced-data settings. The authors report that HAND does not degrade generalisation on common corruptions and can improve the model’s ability to reject unknown classes, with benefits observed across multiple CNN architectures and datasets.

By Michael W. Spratling, Heiko H. Sch\"utt
arXiv Computer Vision
Sep 17

Lumen: Parameter-Efficient Alignment of Pretrained Vision and Language Encoders for Zero-Shot Computational Pathology

Lumen is a pathology vision‑language model that aligns frozen unimodal foundation models (Virchow2 and BioMedBERT) using rank‑4 adapters and projection heads, training only 0.40% of the total parameters on the QUILT‑1M corpus. It achieves the highest mean chance‑corrected balanced accuracy (0.546) across nine zero‑shot patch benchmarks and demonstrates strong performance on lymph‑node metastasis detection, with AUROC scores of 0.964 internally and 0.955 externally. While it ranks third in cross‑modal retrieval, Lumen’s low‑parameter training yields competitive results at both patch and slide levels.

By Kiarash Tajbakhsh, Abdelrahman Faqieh, Michael Jopiti, Javier Garcia-Baroja, Philipp Zens, Branislav Zagrapan, Yuri Tolkach, Martin D. Berger, Aurel Perren, Bastian Dislich, Inti Zlobec, Amjad Khan
arXiv Computer Vision
Sep 28

CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices

The paper introduces Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high‑resolution spatial representations from a YOLO11m‑P2 teacher to a lightweight YOLO11n student without changing the student’s inference architecture. CSCWD aligns teacher P2 features with student P3 while also applying same‑scale distillation at deeper pyramid levels, yielding a 2.92‑point mAP@0.5 improvement over the baseline and a 2.09‑point gain over same‑scale distillation alone. In zero‑shot tests on DUT‑Anti‑UAV and on a Raspberry Pi 5, the 2.58‑million‑parameter student reaches 50.32% mAP@0.5 at 82.32 ms latency (12.15 fps) with negligible runtime or memory increase.

By Amir Zamani, Zeinab Ghasemi-Naraghi