arXiv:2608.22922v1 Announce Type: new
Abstract: We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourc...
By Thisen Ekanayake, Nisansa de Silva
arXiv:2608.29970v1 Announce Type: new
Abstract: Artistic Text Recognition (ATR) remains challenging because word images often combine decorative fonts, curved layouts, object-like characters, clutter...
By Lucas A. Dias, Henrique A. Schulz, Rafaela de Miranda, Guilherme L. Peres, Pedro L. Bittencourt, Rayson Laroca
MIL-BERT is a neural network algorithm that classifies large texts by selecting relevant excerpts, inspired by multiple instance learning. It scales to samples with nearly 1 million tokens and has been evaluated on seven datasets, achieving state‑of‑the‑art results on three long‑text tasks such as political bias detection, trigger warning identification, and author demographic inference. The model also generalizes from weakly‑labeled text bags to accurately classify smaller instances.
By John Cadigan, Dayne Freitag, Eric Yeh
arXiv:2311. 17633v2 Announce Type: replace-cross Abstract: Transformers have dominated empirical machine learning models of natural language processing.
By Tong Xiao, Jingbo Zhu
WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially more challenging than general Scene Text Recognition (STR). Existing STR datasets and methods, typically built around regular scene text and fixed-template inputs, struggle to scale to WATER.