Hugging Face Trending Papers

TEAMS: Text-prompted spatiotEmporal dual-heAd Mamba Snake

The paper introduces TEAMS, a vision‑language Mamba snake framework that enhances deep snake instance segmentation. It adds a Spatiotemporal Snake Evolution Strategy to handle complex shapes, a Contour Morphology‑Aware Mamba to improve fine‑grained detail capture, and a Text‑prompted Collaborative Dual‑Head Snake to integrate textual cues and reduce detection errors. Experiments on five medical imaging datasets show TEAMS surpasses existing methods, achieving significant gains in mDice and mBF metrics.

arXiv Computer Vision
Sep 25

MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting

MEVL-STP introduces a two‑stage pipeline for spotting arbitrarily shaped scene text. The detection stage fuses features from six frozen vision encoders via a hierarchical Feature Pyramid Network and a Progressive Scale Expansion network to produce precise polygon masks. The recognition stage then crops these masks and feeds them to a fine‑tuned Qwen3‑VL‑8B‑Instruct model, achieving state‑of‑the‑art detection and end‑to‑end performance on CTW1500, Total‑Text, and ICDAR 2015 without synthetic pretraining.

By Aman Anand, Partha Pratim Roy, Shivakumara Palaiahnakote
arXiv Computer Vision
Sep 10

Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical Tasks

arXiv:2505.03380v2 Announce Type: replace Abstract: Accurate delineation of tumors and surrounding organs-at-risk is essential for radiotherapy, surgery and treatment response assessment, yet remains...

By Haonan Wang, Jiaji Mao, Lehan Wang, Qixiang Zhang, Marawan Elbatel, Yi Qin, Huijun Hu, Baoxun Li, Wenhui Deng, Weifeng Qin, Hongrui Li, Jialin Liang, Jun Shen, Xiaomeng Li
arXiv AI
Sep 3

InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation

InstEditSeg is a generative framework that treats medical segmentation as an instruction-driven image editing task. Instead of producing binary masks, it renders a color-coded overlay on the original image guided by textual instructions, leveraging latent diffusion models to align with natural image distributions and reduce domain gaps. The method incorporates a DINOv3 visual encoder and a multi-scale feature pyramid fused into the diffusion U‑Net, and uses a dual‑branch classifier‑free guidance strategy to lower inference cost, achieving competitive accuracy on polyp and skin lesion datasets while improving cross‑domain generalization and multi‑lesion segmentation.

By Ziquan Liu, Zhewei Zhu, Xuyang Shi
arXiv AI
Aug 28

From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.

By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao