Confidence-Aware Ensemble and Long-Word Refinement for Artistic Text Recognition
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially more challenging than general Scene Text Recognition (STR). Existing STR datasets and methods, typically built around regular scene text and fixed-template inputs, struggle to scale to WATER.
arXiv:2609.00816v1 Announce Type: new Abstract: Scene text recognition is reported as 89--97% accurate on the six standard benchmarks, and the problem is widely treated as saturated. We present an al...
SpanCalib-VLM is a hybrid system for detecting hallucinated text spans in Vision‑Language Models. It combines a multimodal sequence tagger (XLM‑RoBERTa‑Large + SigLIP) with a fine‑tuned generative VLM (Qwen3.5‑4B‑SHROOM‑SFT) and uses a Union‑Calibrated Fusion strategy to re‑score candidate spans. On the SHROOM‑Visions English evaluation split, the ensemble achieves a Pearson calibration correlation of 0.41, an overall IoU of 0.39, a clean‑response IoU of 0.91, and a detection accuracy of 70.7%.
RefLAM is a pipeline that converts manuscript page images and clean transcriptions into validated, line-level ground truth for Arabic handwritten text recognition. It combines a deep‑learning page‑segmentation model, a multimodal large language model for structured OCR, and a diacritic‑agnostic fuzzy alignment engine that assigns a confidence score to each line, with a provable correctness guarantee for perfect scores. Using RefLAM, the authors achieved a 75× speedup over manual annotation and released AraMS‑28k, a dataset of 14 historical Arabic manuscripts with detailed annotations.
arXiv:2608. 19637v1 Announce Type: new Abstract: Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition.
arXiv:2607. 00684v1 Announce Type: new Abstract: The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts.