arXiv:2604. 09063v3 Announce Type: replace-cross Abstract: Human action recognition is pivotal in computer vision, with applications ranging from surveillance to human-robot interaction.
By Yuxi Zhou, Zhengbo Zhang, Jingyu Pan, Zhiyu Lin, Zhigang Tu
arXiv:2607. 17342v1 Announce Type: cross Abstract: Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision.
By Yuhang Wen, Mengyuan Liu, Zixuan Tang, Junsong Yuan, Sirui Li, Beichen Ding
InstEditSeg is a generative framework that treats medical segmentation as an instruction-driven image editing task. Instead of producing binary masks, it renders a color-coded overlay on the original image guided by textual instructions, leveraging latent diffusion models to align with natural image distributions and reduce domain gaps. The method incorporates a DINOv3 visual encoder and a multi-scale feature pyramid fused into the diffusion U‑Net, and uses a dual‑branch classifier‑free guidance strategy to lower inference cost, achieving competitive accuracy on polyp and skin lesion datasets while improving cross‑domain generalization and multi‑lesion segmentation.
By Ziquan Liu, Zhewei Zhu, Xuyang Shi
arXiv:2606. 07053v1 Announce Type: cross Abstract: Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios.
By Dian Gu, Zhengyi Yang
arXiv:2608.30420v1 Announce Type: cross
Abstract: Automating the analysis of whole-slide images has high clinical value, since characterizing cancers requires examining them in detail. Such analysis...
By Tiffanie Godelaine, Maxime Zanella, Karim El Khoury, Benoit Macq, Christophe De Vleeschouwer
Clear cell renal cell carcinoma (CCRCC) grading is essential for treatment planning, yet existing approaches either analyze patch-level images directly or focus solely on nuclei-level classification,...
arXiv:2607. 00716v1 Announce Type: cross Abstract: Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume complete and clean skeleton inputs.
By Yingjie Dai, Tianyang Xu, Yanglin Deng, Xiao-Jun Wu, Josef Kittler
arXiv:2609.00396v1 Announce Type: new
Abstract: Histopathological whole slide images (WSIs) are central to cancer diagnosis, but their gigapixel scale, tissue heterogeneity, weak slide-level supervis...
By Chad Wong, Sicheng Chen, Tianyi Zhang, Enhui Chai, Yueming Jin, Zeyu Liu, Fei Xia
The paper introduces Timo, a kinematics-aware multimodal diffusion transformer designed for human motion generation. Timo employs fully shared multimodal attention, flow matching, and geometric/rotational-kinematics supervision to better coordinate articulated motion, and uses a two-stage curriculum to align motion with text captions. The authors also present a new benchmark of 40,025 clips from six datasets, showing that Timo outperforms state‑of‑the‑art methods, achieving a 40.8% relative improvement over Kimodo on average.
By Zhao Wang, Jiangtao Hu, Jack Yu, Tao Yu
The paper introduces a unified conditional-flow framework that integrates text-driven motion generation, semantic editing, and intra-structural retargeting into a single rectified-flow model. By treating editing as a change in semantic condition and retargeting as a change in skeletal condition, the approach eliminates fragmented pipelines and allows a single model to perform generation, zero‑shot editing, and zero‑shot retargeting on articulated 3D motion data. Experiments on SnapMoGen and a Mixamo subset demonstrate that the model can handle all three tasks without task‑specific fine‑tuning, preserving both motion semantics and skeletal structure.
By Junlin Li, Xinhao Song, Siqi Wang, Haibin Huang, Yili Zhao
The paper introduces a semantic‑guided multimodal preprocessing technique that fuses nuclei classification maps with RGB histopathology images for Vision Transformer‑based grading of clear cell renal cell carcinoma. By concatenating classification map channels and applying multiplicative modulation, the method achieves a balanced accuracy of 0.916, markedly surpassing an RGB‑only baseline (0.707) and prior max‑voting approaches (0.427). Sensitivity analysis shows the 21‑percentage‑point improvement remains robust under simulated perturbations matching current nuclei classifier error rates, indicating effective use of imperfect nuclear‑level information.
By Fatemeh Javadian, Zhu Chen, Zahra Aminparast, Johannes Stegmaier
UniH$^3$ is a new framework for all-in-one medical image restoration that unifies hierarchical homogeneity and heterogeneity. It introduces a Hierarchical Homogeneity Memory (H2M) module to distill and retrieve shared anatomical priors, and a Hierarchical Heterogeneity Balancer (H2B) to mitigate inter- and intra-task conflicts during training. Experiments on MedIR-2D-500K and MedIR-3D-3K show that UniH$^3$ achieves state‑of‑the‑art performance for both multi‑task and single‑task restoration.
By Zhiwen Yang, Jiayin Li, Chengyu Liu, Hui Zhang, Bingzheng Wei, Yan Xu