arXiv Machine Learning

PA-CDM: Position-Aware Character Detection Matching for Evaluating Handwritten Mathematical Expression Recognition

The paper introduces PA-CDM, a new position‑aware character detection matching metric for handwritten mathematical expression recognition that improves on existing render‑based and tree‑edit metrics by incorporating position‑forest encoding and divergence‑level weighting. It also presents StructPerturb v2.0, a benchmark of 1,340 controlled perturbation pairs, and a cross‑metric consistency protocol that includes a sensitivity matrix, a human study, and LLM‑judge calibration. In a six‑annotator study, PA‑CDM achieves the highest correlation with human judgments (Spearman rho = 0.9535) among seven automatic metrics, approaching the performance of a costly LLM judge while remaining deterministic and cost‑free.

arXiv AI
Aug 20

OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios

OmniHandwritingOCR is a diagnostic benchmark designed to evaluate multimodal large language models (MLLMs) and OCR systems on handwritten text and mathematical expression recognition. It comprises 77.57K labeled images across six subtasks and twelve subsets, including a difficulty‑stratified multi‑line formula corpus that tests robustness to increasing structural complexity. The benchmark reveals that current systems perform poorly on complex multi‑line formulas, exhibit variable rankings across languages and formula settings, and sometimes hallucinate corrections that are not visually supported.

By Zinuo Guo, Min Zhang, Bo Jiang
arXiv Computer Vision
Sep 2

Can Scene Text Recognition Read Rare Compositions?

The paper reports that scene text recognition models, while achieving 89–97% accuracy on standard benchmarks, perform significantly worse on rare word–trigram combinations, with a 10–18 point drop in accuracy at the rare‑word/rare‑trigram corner across multiple languages and models. Scaling the vision backbone improves overall accuracy but does not alleviate this corner‑specific deficit. The authors identify the autoregressive decoder’s lexical prior as the root cause and show that architectural changes—specifically moving from autoregressive to CTC decoding—yield the largest improvement for these rare compositions.

By Genpei Zhang
arXiv Computer Vision
Sep 4

OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement

OCR-EDR is a rendering‑aware framework that diagnoses and repairs OCR errors by comparing an editable OCR prediction with its rendered image. It jointly assesses consistency, localizes genuine errors, and applies executable edits, optionally requesting updated renderings for iterative reassessment. On the newly constructed OCRErrBench, the DocEDR model achieves 94.78% diagnostic accuracy and repairs 86.23% of erroneous inputs, improving formula metrics by up to 30.99 percentage points on benchmark datasets.

By Linnan Zhao, Kang Liu, Hao Yu, Jiabo Zhan, Chong Sun, Chen Li
Hugging Face Trending Papers
Sep 3

OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement

OCR-EDR is a rendering‑aware framework that diagnoses and repairs OCR errors by jointly evaluating an OCR prediction, its editable form, and the rendered image of the source document. It localizes genuine mistakes while preserving valid or rendering‑equivalent predictions, then applies executable edits and iteratively reassesses with updated renderings. On the newly constructed OCRErrBench, the DocEDR model achieves 94.78% diagnostic accuracy and repairs 86.23% of errors, boosting formula metrics by up to 30.99 percentage points and improving CDM scores on several OCR systems.

arXiv AI
2d ago

On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence

The paper evaluates nine on‑device named‑entity recognition models ranging from classical taggers to large language models, measuring not only accuracy but also latency and output validity. Using a silver‑gold benchmark derived from an LLM judge panel and a human‑validated corpus, the study shows that encoder‑based models achieve comparable accuracy to a 4 B instruct LLM while being much smaller, faster, and producing no malformed output. Confidence calibration of GLiNER is analyzed, revealing over‑confidence but improved reliability after temperature scaling and thresholding.

By Vinay Kumar Chaganti
Hugging Face Trending Papers
Aug 6

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at.