arXiv Computer Vision
Sep 11

TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

TeleOCR is a unified framework for document parsing that tackles challenges in both decoupled and end-to-end Vision‑Language Models. It introduces deformation‑aware learning to handle geometric distortions, an adaptive sampling mechanism for complex layouts, and a content‑structure decoupled strategy to model formula grammars and table structures. The approach achieves state‑of‑the‑art results on multiple benchmarks, including top placement in the ICDAR 2026 Sci‑ImageMiner Challenge.

By Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, MingKun Jiang, Zhongjiang He, Hao Sun
arXiv Computation and Language
Aug 28

TabuLM: Morphology-Aware Tabular Pre-training for Low-Resource Languages

TabuLM is a new language model pre‑trained on Kinyarwanda tabular data, extending KinyaBERT‑large with row, column, and cell‑type embeddings and a table‑structure attention bias. It introduces two pre‑training objectives—Masked Cell Recovery and Column Type Prediction—and is trained on 172 Rwandan government tables. On the TabQA‑kin benchmark, TabuLM achieves 62.0% exact match, outperforming KinyaBERT‑large and multilingual baselines by significant margins.

By Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre