arXiv:2610.10303v1 Announce Type: new
Abstract: The standard Transformer architecture relies on a rigid pattern that alternates Attention and Feed-Forward Network (FFN) layers. Despite its widespread...
By Junhyeok Kim, Jinyeong Kim, Jae Wan Park, Seong Jae Hwang
arXiv:2610.10306v1 Announce Type: new
Abstract: Scientific image segmentation methods rely on extensive annotation and task-specific training, limiting adaptation across imaging modalities and experi...
By Tejaswi V. Panchagnula, Allison M. Davis, Fengqing Zhu
arXiv:2610.10324v1 Announce Type: new
Abstract: Pretrained models for cell and nuclear instance segmentation differ substantially in architecture, pretraining data and objectives, parameter count, in...
By Eiram Mahera Sheikh, Alaa Tharwat, Wolfram Schenck
arXiv:2610.10491v1 Announce Type: new
Abstract: This paper investigates AutoResearch, a protocol in which a coding language model edits a training program under a one-hour GPU budget and retains a ch...
By Justinas Lekavicius, Kursat Komurcu, Valentas Gruzauskas, Linas Petkevicius
arXiv:2610.08849v1 Announce Type: cross
Abstract: Purpose: Current abdominal ultrasound (US) simulation methods often require CT-based anatomical references for ray-casting, limiting deformation and...
By Santiago Vitale, Duilio Deangeli, Ignacio Larrabide, Jos\'e Ignacio Orlando
arXiv:2610.10359v1 Announce Type: cross
Abstract: We introduce MultiFly, a real-world, low-altitude UAV dataset for semantic perception across RGB, thermal, LiDAR, and radar modalities. MultiFly prov...
By Markus Gross, Andreas Greiner, Taehyoung Kim, Sivasubiramaniam Subbiah, Toma\v{z} Coti\v{c}, Sai Bharadwaj Matha, Conrad Christoph, Oussema Dhaouadi, Simon Zieher, Surya Vijaya Kumar, Gordon Elger, Henri Mee{\ss}, Olaf Wysocki, Paul Spannaus, Daniel Cremers
arXiv:2610.04606v2 Announce Type: replace
Abstract: We present a 4D Gaussian Splatting framework for the Sparse-View Track of the SIGGRAPH Asia 2026 Volumetric Video Challenge, which requires dynamic...
By Shengqi Wang, Zhengxian Yang, Kaiwen Tian, Yang Liu, Bowen Liu, Hua Du, Taicheng Huang, Jiamin Wu, Tao Yu
arXiv:2610.05603v2 Announce Type: replace
Abstract: Echocardiography is the most widely used cardiac imaging modality, yet interpretation demands integrating visual evidence across global anatomy, lo...
By Sheng Cheng, Donnchadh M. O'Sullivan, Daniel J. Penny, Craig G. Rusin, Minh B. Nguyen, Devika Subramanian
arXiv:1907.09194v3 Announce Type: replace-cross
Abstract: Segmentation of subcortical brain structures is fundamental to computer-aided diagnosis and treatment in neurology and related clinical field...
By Binbin Yang, Weiwei Zhang
arXiv:2602.10137v2 Announce Type: replace
Abstract: This work proposes MeCSAFNet, a multi-branch encoder-decoder architecture for land cover segmentation in multispectral imagery. The model separatel...
By Leo Thomas Ramos, Angel D. Sappa
arXiv:2603. 13357v2 Announce Type: replace-cross Abstract: Bi-CamoDiffusion is introduced, an evolution of the CamoDiffusion framework for camouflaged object detection.
By Patricia L. Suarez, Leo Thomas Ramos, Angel D. Sappa
arXiv:2610.09068v1 Announce Type: new
Abstract: Reliable robotic disassembly requires part-level representations that distinguish genuine component geometry from scanning and reconstruction artifacts...
By Zuoxu Wang, Xiao Liang
arXiv:2610.09450v1 Announce Type: new
Abstract: Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. W...
By Hanqiu Li Cai (SperidLabs), Chema Garabito (SperidLabs)
arXiv:2610.09784v1 Announce Type: new
Abstract: Reliable liver segmentation in contrast-enhanced MRI is essential for quantitative hepatic assessment, treatment planning, and longitudinal disease mon...
By Ruoshi Xu, Mingqi Gao, Shengda Luo, Jingkun Chen
arXiv:2610.09800v1 Announce Type: new
Abstract: Surgical video generation holds substantial potential for surgical education, simulation, and data augmentation, yet generating surgical videos with re...
By Tsz-Yui Qin, Siyu Zhou, Chi-Keung Tang, Yuxiang Nie, Shu Yang
arXiv:2610.10409v1 Announce Type: cross
Abstract: General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities car...
By Zhiqin Yang, Chenxin Li, Xiaomeng Hu, Yibin Liu, Weidong Huang, Jiankai Sun, Haitao Li, Zijian Wu, Yuzhi Huang, Fanding Huang, Hanwen Sun, Jiashun Liu, Jingqi Tong, Mingxin Huang, Shaoli Hu, Shijue Huang, Tianyi Bai, Xinyuan Wang, Yunlong Lin, Zhengyang Tang, Zhexin Zhang, Zhuo Chen, Xierui Song, Juntao Dai, Boyuan Chen, Jiaming Ji, Fangneng Zhan, Mengkang Hu, Wei Xue, Yonggang Zhang, Han Hu, Tsung-Yi Ho, Yike Guo
arXiv:2610.09807v1 Announce Type: new
Abstract: Camouflaged object detection requires pixel-accurate masks, but obtaining such annotations is slow and costly, making synthetic training images an attr...
By Akshat Dobhal, Sanjay Singh
arXiv:2610.09478v1 Announce Type: new
Abstract: Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insig...
By Yuchen Li, Shaoyang Zhou, Yiran Wang, Ruiyi Deng, Haoyu Wang, Ziru Wei, Zhen Zhao, Luping Zhou
arXiv:2610.10225v1 Announce Type: new
Abstract: Whole slide image (WSI) analysis in computational pathology follows a multiple instance learning (MIL) pipeline where patch embeddings are extracted in...
By Haoyu He, Basile Tessier-Cloutier, Yang Wang, Mahdi S. Hosseini
arXiv:2610.10335v1 Announce Type: new
Abstract: Visual in-context learning (ICL), well suited to label-scarce medical imaging, uses support image-label pairs to demonstrate input-output mappings, whi...
By Cheng Wan, Chenjun Li, Qingyu Zhao