arXiv AI By Nikhil Reddy Pottanigari, Sepideh Kharaghani, Saverio Vadacchino, Alejandro Posada, Ying Zhang

Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Sep 14

Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking

Doc2FRC introduces Fixed-Range Chunking (FRC), a dynamic programming method that partitions documents into chunks of a predefined length interval, ensuring consistent length distributions during training and inference. This approach reduces train-test length mismatch, mitigates n-gram repetition, and improves translation quality for 7B LLMs compared to direct Doc2Doc fine-tuning. Experiments on IWSLT2017 and a new 10-language test set, GlobVDoc, demonstrate that FRC outperforms existing document-level machine translation methods and enhances out-of-distribution translation performance.

By Xiaotian Wang, Youyuan Lin, Zhan Shen, Hitomi Yanaka
arXiv AI
Aug 19

Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs

Intent-Driven Dynamic Chunking (IDC) segments documents by predicting user queries with a Large Language Model and then applying dynamic programming to find optimal chunk boundaries. This method outperforms traditional fixed-length or coherence-based segmentation on five out of six question-answering datasets, improving top-1 retrieval accuracy by 5% to 67% and reducing the number of chunks by 40–60% while maintaining 93–100% answer coverage. IDC demonstrates that aligning document structure with anticipated information needs can significantly boost retrieval performance for long and heterogeneous documents.

By Christos Koutsiaris
arXiv Machine Learning
5d ago

Named Entity Recognition using Sliding Window Approach

The paper presents an inference-only pipeline that extends the frozen NER model MahaNER‑BERT to document‑level prediction using overlapping sliding windows, eliminating the need for retraining or architectural changes. The approach is evaluated on six document‑level corpora derived from the MahaNER test set, employing two repetition strategies (Normal Repeat and Random Repeat) at three length levels and various window configurations. Results show the model maintains a macro F1‑score of up to 0.8902 with minimal variation, outperforming non‑windowed methods by avoiding boundary‑fragmentation errors and achieving more stable document‑level performance.

By Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Ravindra Murumkar, Raviraj Joshi
arXiv AI
Aug 12

A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models

arXiv:2608. 10939v1 Announce Type: cross Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages.

By Wajdi Ben Saad, Safa Madiouni
arXiv Computation and Language
5d ago

ChunkRank: Model-Aware Text Chunking and Abstention-Aware Answer Selection for LLM Pipelines

ChunkRank is an open‑source Python library that automatically determines chunk boundaries based on a target model’s tokenizer and context window, and then selects an answer from independently produced chunk candidates. It includes a registry of 90 models from 15 providers and six answer‑selection methods, and requires only three core dependencies. Experiments show that token‑exact budgeting is important across 11 languages, and that for several datasets no content‑based ranker outperforms simply taking the first non‑empty answer due to reader abstention on chunks lacking the answer.

By Amit Nautiyal, Ayush Bhatt, Gaurav Nautiyal