arXiv Computer Vision By Zixiang Tong, Lehu Bu, Jin Yang

RAFT-DVC: Resolution-Aware Machine Learning-Based Digital Volume Correlation

Read the original on arXiv Computer Vision →

RAFT-DVC is a resolution‑aware family of recurrent all‑pairs field transform (RAFT) based digital volume correlation (DVC) solvers that use encoder downsampling factors of 2, 4, and 8. The solvers localize displacement to about 0.017 feature‑grid voxels, with raw‑volume error scaling roughly as 0.017 s voxels, and exhibit complementary operating regimes determined by displacement reach and volumetric‑texture compatibility. Synthetic benchmarks show comparable performance to tuned classical DVC for fine‑texture, small‑to‑moderate displacements, while outperforming it for coarse‑texture, large‑displacement scenarios; additional tests on confocal and micro‑CT images confirm the importance of matching solver regimes to deformation magnitude and texture, and demonstrate cross‑texture transfer and improved accuracy after correcting sampler geometry.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 10

Deep Slice Interpolation for Reducing Through-Plane Anisotropy and Noise in Head CT

arXiv:2606. 09953v1 Announce Type: cross Abstract: Head computed tomography (CT) typically uses sub-millimeter in-plane resolution but 2-5 mm through-plane spacing, creating substantial anisotropy that degrades multiplanar reconstructions, volumetric measurements such as hematoma volume estimation, and downstream algorithms that assume near-isotropic voxels.

By Luis Cort\'es Ferre, Miguel A. Guti\'errez-Naranjo, Marcin Balcerzyk
arXiv Computer Vision
Aug 27

Steer the Sampling, Not the Kernel Grid: Geometry-Guided Sampling Operator for Volumetric Segmentation

The paper introduces a geometry‑guided sampling operator that directs feature sampling rather than altering convolution kernels in 3D encoder‑decoder networks. By predicting local orientations and bounded step sizes, the operator samples symmetrically around each voxel, generating compact geometric and boundary cues that improve fine‑structure segmentation. Replacing stride‑1 and stride‑2 operations in a 3D U‑Net yields consistent gains on BraTS, MSD Hepatic Vessel, and TDSC‑ABUS datasets, with better boundary metrics and fewer parameters, and the operator can be integrated into other backbones without architectural changes.

By Sizhe Wang, Himashi Peiris, Zhaolin Chen
arXiv Computer Vision
Sep 25

When Misalignment Becomes Supervision: Structured Label Noise in Supervised Synthetic CT Generation

The paper examines how residual misalignments from registration procedures introduce structured label noise in supervised synthetic CT (sCT) generation. It shows that voxel‑wise metrics are heavily influenced by the consistency between training and evaluation registrations, and that training with anatomically consistent registrations reduces variability and improves robustness. Introducing a perceptual loss based on a pretrained Segment Anything encoder yields sharper, more anatomically coherent sCT and highlights the need for anatomy‑oriented evaluation.

By Valentin Boussot, Cedric Hemon, Caroline Lafond, Jean-Claude Nunes, Jean-Louis Dillenseger
arXiv Machine Learning
Sep 10

Large-Scale Pretraining for Improving Deep Learning-Based Geometric Distortion Correction of Diffusion-Weighted Imaging

The paper explores large‑scale pretraining to enhance deep learning‑based geometric distortion correction for diffusion‑weighted imaging (DWI). By framing the task as image reconstruction, the authors compare a non‑pretrained baseline with self‑supervised and generative pretrained models, finding that the cWDM model yields the best quantitative and qualitative results. When applied to low‑resource, high‑throughput settings in a low‑ and middle‑income country, the pretrained models faced transferability issues, but aligning images to a common standard space improved predictions, indicating that harmonized preprocessing can aid cross‑domain deployment.

By Saroj Khanal, Yashawant Kumar Yadav, Kritam Bhattarai, Jeevan Neupane, Shristi Subedi, Saship Gwachha, Manish Kumar Tiwari, Dong Zhang, Confidence Raymond, Aondona Moses Iorumbur, Udunna Anazodo, Surendra Maharjan, Bishesh Khanal, Mahesh Shakya, Pralhad Kumar Shrestha
arXiv Machine Learning
Jun 17

Adaptive Volumetric Mechanical Property Fields Invariant to Resolution

arXiv:2606. 18231v1 Announce Type: cross Abstract: Accurate mechanical properties (or materials) Young's modulus ($E$), Poisson's ratio ($\nu$) and density ($\rho$) are essential for reliable physics simulation of digital worlds, but most 3D assets lack this information.

By Rishit Dagli, Donglai Xiang, Vismay Modi, Xuning Yang, Gavriel State, David I. W. Levin, Maria Shugrina
arXiv Computer Vision
2d ago

From Pixel Generation to Topological Inference: Structural Dual Super-Resolution for Trustworthy Cross-Physical-Domain Trabecular Morphology Learning

The paper introduces Structural Dual Super‑Resolution (SDN), a novel approach that shifts from pixel‑level super‑resolution to topological inference for trabecular bone morphology. By training on 2‑D slices and evaluating on 3‑D morphological metrics, SDN learns to predict invariant microstructures from low‑resolution CT inputs, using bidirectional modeling, a multi‑scale consistency discriminator, and four structural duality constraints. The method achieves SSIM of 0.8 and morphological parameters closely matching synchrotron micro‑CT across six metrics, demonstrating cross‑source generalization and trustworthy inference rather than mere pixel generation.

By Fan Zhang, Yi Zhang, Ling Wang