Whole slide images (WSIs) in digital histopathology are acquired at discrete magnification levels encoding complementary diagnostic information from global tissue architecture to fine-grained cellular morphology. Yet, deep learning models remain sensitive to scale variation.
arXiv:2505. 04397v2 Announce Type: replace-cross Abstract: Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplored.
By Ziyuan Li, Uwe Jaekel, Babette Dellen
Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical viewpoint, a natural question is: can DNNs attain feature-learning and prediction consistency comparable to that of classical models?
arXiv:2403.04545v4 Announce Type: replace
Abstract: Scaling factors in residual branches have emerged as a prevalent method for boosting neural network performance, especially in normalization-free a...
By Zixiong Yu, Guhan Chen, Jianfa Lai, Bohan Li, Songtao Tian
GradAttn replaces fixed residual connections in deep ConvNets with attention‑controlled gradient pathways, allowing the network to dynamically weight shallow texture features and deep semantic representations. The method extracts multi‑scale CNN features at different depths and regulates them through self‑attention, leading to improved performance over ResNet‑18 on five of eight evaluated datasets, including a +11.07% accuracy gain on FashionMNIST. Analysis of gradient flow shows that controlled instabilities introduced by attention can coincide with better generalization, while positional encoding proves to be dataset‑dependent.
By Soudeep Ghoshal, Himanshu Buckchash
arXiv:2606. 29400v1 Announce Type: cross Abstract: In computer graphics, visual content is continuously warped, zoomed and resampled.
By Giulio Federico, Giuseppe Amato, Claudio Gennaro, Fabio Carrara, Marco Di Benedetto
MIMONet is a saliency detection model that uses multi‑scale inputs and outputs to better handle objects of varying sizes. It processes three differently sized images through separate encoder branches that exchange information, allowing each branch to learn size‑variation knowledge from the others. A Multi‑scale Perception module further refines features, and a Joint Saliency Loss ensures consistent, well‑preserved boundaries across the multiple saliency maps produced.
By Zhaojian Yao, Wei Gao, Tiesong Zhao, Hui Yuan, Sam Kwong
arXiv:2407.03463v2 Announce Type: replace-cross
Abstract: In the realm of self-supervised learning (SSL), conventional wisdom has gravitated towards the utility of massive, general domain datasets fo...
By Jes\'us M Rodr\'iguez-de-Vera, Imanol G Estepa, Ignacio Saras\'ua, Bhalaji Nagarajan, Petia Radeva
The paper introduces Structural Dual Super‑Resolution (SDN), a novel approach that shifts from pixel‑level super‑resolution to topological inference for trabecular bone morphology. By training on 2‑D slices and evaluating on 3‑D morphological metrics, SDN learns to predict invariant microstructures from low‑resolution CT inputs, using bidirectional modeling, a multi‑scale consistency discriminator, and four structural duality constraints. The method achieves SSIM of 0.8 and morphological parameters closely matching synchrotron micro‑CT across six metrics, demonstrating cross‑source generalization and trustworthy inference rather than mere pixel generation.
By Fan Zhang, Yi Zhang, Ling Wang
arXiv:2608. 19817v1 Announce Type: cross Abstract: Conventional convolutional kernels are typically defined on fixed discrete grids, limiting their ability to accommodate heterogeneous local structures.
By Lan Guo, Mengling Li, Haoran Li, Jun Shen, Yuanbo Jiang, Qingguo Zhou, Binbin Yong
arXiv:2608. 08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task.
By Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.
By Shaojie Li, Yunbei Xu