arXiv Machine Learning By Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Ravindra Murumkar, Raviraj Joshi

Named Entity Recognition using Sliding Window Approach

Read the original on arXiv Machine Learning →

The paper presents an inference-only pipeline that extends the frozen NER model MahaNER‑BERT to document‑level prediction using overlapping sliding windows, eliminating the need for retraining or architectural changes. The approach is evaluated on six document‑level corpora derived from the MahaNER test set, employing two repetition strategies (Normal Repeat and Random Repeat) at three length levels and various window configurations. Results show the model maintains a macro F1‑score of up to 0.8902 with minimal variation, outperforming non‑windowed methods by avoiding boundary‑fragmentation errors and achieving more stable document‑level performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 14

Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking

Doc2FRC introduces Fixed-Range Chunking (FRC), a dynamic programming method that partitions documents into chunks of a predefined length interval, ensuring consistent length distributions during training and inference. This approach reduces train-test length mismatch, mitigates n-gram repetition, and improves translation quality for 7B LLMs compared to direct Doc2Doc fine-tuning. Experiments on IWSLT2017 and a new 10-language test set, GlobVDoc, demonstrate that FRC outperforms existing document-level machine translation methods and enhances out-of-distribution translation performance.

By Xiaotian Wang, Youyuan Lin, Zhan Shen, Hitomi Yanaka
arXiv Machine Learning
Jun 2

GottBERT: a pure German Language Model

arXiv:2012. 02110v2 Announce Type: replace-cross Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimized version, RoBERTa.

By Raphael Scheible, Johann Frei, Fabian Thomczyk, Henry He, Patric Tippmann, Jochen Knaus, Victor Jaravine, Frank Kramer, Martin Boeker