Lumen is a pathology vision‑language model that aligns frozen unimodal foundation models (Virchow2 and BioMedBERT) using rank‑4 adapters and projection heads, training only 0.40% of the total parameters on the QUILT‑1M corpus. It achieves the highest mean chance‑corrected balanced accuracy (0.546) across nine zero‑shot patch benchmarks and demonstrates strong performance on lymph‑node metastasis detection, with AUROC scores of 0.964 internally and 0.955 externally. While it ranks third in cross‑modal retrieval, Lumen’s low‑parameter training yields competitive results at both patch and slide levels.
By Kiarash Tajbakhsh, Abdelrahman Faqieh, Michael Jopiti, Javier Garcia-Baroja, Philipp Zens, Branislav Zagrapan, Yuri Tolkach, Martin D. Berger, Aurel Perren, Bastian Dislich, Inti Zlobec, Amjad Khan
arXiv:2607. 03581v1 Announce Type: cross Abstract: The widespread adoption of facial masks, accelerated by COVID-19 and mandated in security-sensitive settings, has exposed limitations of conventional face recognition systems.
By Dana A Abdullah
arXiv:2609.13237v1 Announce Type: cross
Abstract: Orthodontic report generation from intraoral data is normally cast as multimodal captioning, yet the released Bite2Text scan pairs are supplied alrea...
By Ajo Babu George, Govind Arun, Sidharth N Krishna, Uma Ranjan
arXiv:2512.15774v5 Announce Type: replace
Abstract: The absence of large-scale masked face datasets challenges masked face detection and recognition. We propose a two-step generative data augmentatio...
By Yan Yang, George Bebis, Mircea Nicolescu
arXiv:2609.15144v1 Announce Type: new
Abstract: Purpose: Paired pre- and post-operative photographs are the standard unit of evidence for plastic surgical outcomes, yet no objective metric verifies w...
By Derrick Lin, Samantha Rabinovich, Joclin Rabinovich, Kassra Garoosi, Sumun Khetpal, Evan Delanoy, Neel Bhardwaj, Jason Roostaeian
arXiv:2607. 21318v1 Announce Type: cross Abstract: Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhouette, and preservation of unrelated content.
By Jian Zhang, Zhijun Zhang
LENS‑GRF is a permutation‑invariant lesion evidence network that uses a Set‑Transformer and gated residual fusion to combine global facial context with localized lesion patches for four‑class acne severity grading. The framework integrates adaptive facial skin segmentation, a Vision Transformer prior, and a lesion set transformer that encodes spatial geometry, with a gating mechanism that modulates local residual contributions. In experiments on ACNE04 and PLSBRACNE01, the fully automated model achieved 80.82% accuracy, while using ground‑truth lesion annotations raised accuracy to 95.89% and a Quadratic Weighted Kappa of 0.9753; zero‑shot evaluation on the full cohort yielded 35.00% accuracy versus 42.50% for a global baseline, and oracle analyses on a 148‑subject cohort showed improved accuracy and QWK up to 47.97% and 0.5799.
By Muhammad Muhtasim Shahriar, M. F. Mridha
arXiv:2607. 26765v1 Announce Type: cross Abstract: Background/Objectives: Dermoscopic skin lesion classifiers often lose accuracy under domain shift across imaging devices, illumination, and capture artifacts.
By Alexander Kozachok, Ilya Latyshev, Evgeny Karpulevich, Elena Kozachok, Egor Ushakov, Oleg Samovarov
Background/Objectives: Dermoscopic skin lesion classifiers often lose accuracy under domain shift across imaging devices, illumination, and capture artifacts. We study how data augmentation improves the robustness of a binary malignant-versus-non-malignant classifier, with emphasis on out-of-domain (OOD) generalization.
Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-based methods, strong regional controllability, makeup fidelity, and identity preservation remain challenging.
The study evaluates automatic tooth segmentation on panoramic radiographs using a large annotated corpus of 1,422 images and 42,142 tooth polygons. It finds that increasing input resolution improves boundary precision (mask mAP50‑95 rises from 0.656 to 0.717) while detection performance remains unchanged, and that architectural changes have minimal impact on in‑domain accuracy. Targeted interventions such as LoRA adaptation, promptable foundation models, and anatomical label assignment provide negligible gains, indicating that resolution and acquisition diversity should be prioritized over model novelty.
By Muhammad Rehan, Moaz Amjad, Syed Danial Ahmed, Mariam Adnan, Haider Ali
CATCH is a conditional 3D diffusion model operating in an invertible Haar-wavelet domain designed to inpaint masked regions in T1‑weighted brain MRI with plausible, tumor‑free tissue while preserving observed anatomy. The model’s denoiser uses noisy target coefficients, voided‑image coefficients, and a signed mask, guided by tumor‑excluded wavelet reconstruction and a hole‑focused loss, and hard compositing ensures observed voxels remain unchanged. Experiments on BraTS data show that a weighted mixture of tumor‑derived, irregular‑blob, and ellipsoidal masks yields the best performance, achieving higher SSIM, PSNR, and lower MSE compared to fixed or random augmentation baselines.
By Simon Winther Albertsen, Hjalte Bjoernstrup, Said Djafar Said, Mostafa Mehdipour Ghazi