arXiv Machine Learning By Manasi Waghe, Danish Chandargi, Mohammad Aamir Rayyan, Raviraj Joshi, A. R. Deshpande

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi

Read the original on arXiv Machine Learning →

arXiv:2606. 28796v1 Announce Type: cross Abstract: Government documents in India are predominantly issued in regional languages such as Marathi, creating substantial accessibility barriers for non-native readers, interstate administrative bodies, and policy analysts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 19

Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts

The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline improves transcription accuracy across subsequent pages, reducing the need for costly human annotation. The authors apply the method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the results against leading multimodal large language models.

arXiv AI
Aug 20

Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts

The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline reduces human annotation effort and improves transcription accuracy across subsequent pages. The authors apply this method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the pipeline against leading multimodal large language models.

By Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis
arXiv Computation and Language
Aug 28

STAR : Sentence Translation Alignment Rate for Document-to-Document Machine Translation

The paper introduces STAR, a metric that measures sentence-level alignment between source and target documents in document-to-document machine translation. Using STAR, the authors develop StarPO, a preference‑optimization framework that ranks translation hypotheses by structural quality and applies a dynamic alignment mask to focus learning on misaligned segments. Experiments on news and literary data show that StarPO improves both translation quality and structural integrity, enabling small models to outperform large proprietary systems such as GPT‑4o while remaining more token‑efficient.

By Yichen Dong, Hao Wang, Junhui Li, Linlong Xu, Longyue Wang, Weihua Luo