Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Doc2FRC introduces Fixed-Range Chunking (FRC), a dynamic programming method that partitions documents into chunks of a predefined length interval, ensuring consistent length distributions during training and inference. This approach reduces train-test length mismatch, mitigates n-gram repetition, and improves translation quality for 7B LLMs compared to direct Doc2Doc fine-tuning. Experiments on IWSLT2017 and a new 10-language test set, GlobVDoc, demonstrate that FRC outperforms existing document-level machine translation methods and enhances out-of-distribution translation performance.
Intent-Driven Dynamic Chunking (IDC) segments documents by predicting user queries with a Large Language Model and then applying dynamic programming to find optimal chunk boundaries. This method outperforms traditional fixed-length or coherence-based segmentation on five out of six question-answering datasets, improving top-1 retrieval accuracy by 5% to 67% and reducing the number of chunks by 40–60% while maintaining 93–100% answer coverage. IDC demonstrates that aligning document structure with anticipated information needs can significantly boost retrieval performance for long and heterogeneous documents.
The paper presents an inference-only pipeline that extends the frozen NER model MahaNER‑BERT to document‑level prediction using overlapping sliding windows, eliminating the need for retraining or architectural changes. The approach is evaluated on six document‑level corpora derived from the MahaNER test set, employing two repetition strategies (Normal Repeat and Random Repeat) at three length levels and various window configurations. Results show the model maintains a macro F1‑score of up to 0.8902 with minimal variation, outperforming non‑windowed methods by avoiding boundary‑fragmentation errors and achieving more stable document‑level performance.
arXiv:2608. 10939v1 Announce Type: cross Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages.
arXiv:2608.27658v1 Announce Type: new Abstract: Subword tokenization hinders low-resource language processing by imposing frequency patterns from dominant languages onto script-sharing variants. Byte...
ChunkRank is an open‑source Python library that automatically determines chunk boundaries based on a target model’s tokenizer and context window, and then selects an answer from independently produced chunk candidates. It includes a registry of 90 models from 15 providers and six answer‑selection methods, and requires only three core dependencies. Experiments show that token‑exact budgeting is important across 11 languages, and that for several datasets no content‑based ranker outperforms simply taking the first non‑empty answer due to reader abstention on chunks lacking the answer.