arXiv Computer Vision
Sep 25

MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting

MEVL-STP introduces a two‑stage pipeline for spotting arbitrarily shaped scene text. The detection stage fuses features from six frozen vision encoders via a hierarchical Feature Pyramid Network and a Progressive Scale Expansion network to produce precise polygon masks. The recognition stage then crops these masks and feeds them to a fine‑tuned Qwen3‑VL‑8B‑Instruct model, achieving state‑of‑the‑art detection and end‑to‑end performance on CTW1500, Total‑Text, and ICDAR 2015 without synthetic pretraining.

By Aman Anand, Partha Pratim Roy, Shivakumara Palaiahnakote
Hugging Face Trending Papers
Jun 23

Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods

WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially more challenging than general Scene Text Recognition (STR). Existing STR datasets and methods, typically built around regular scene text and fixed-template inputs, struggle to scale to WATER.

arXiv Computer Vision
Sep 3

Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling

The paper investigates how much clinical text influences pixel‑level predictions in multimodal medical image segmentation. It shows that segmentation performance is largely insensitive to the choice of fusion module, but that the impact of text varies across datasets: removing text severely degrades performance on BUSI and BTMRI, while it has only a marginal effect on ISIC and Kvasir‑SEG. Using an Evidence Decoupling Decoder, the authors reveal that text mainly modulates global semantic context rather than spatial localization, and that the specific semantic components driving sensitivity differ by dataset.

By Ziquan Liu, Zhewei Zhu, Xuyang Shi