arXiv AI By Christos Koutsiaris

Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs

Read the original on arXiv AI →

Intent-Driven Dynamic Chunking (IDC) segments documents by predicting user queries with a Large Language Model and then applying dynamic programming to find optimal chunk boundaries. This method outperforms traditional fixed-length or coherence-based segmentation on five out of six question-answering datasets, improving top-1 retrieval accuracy by 5% to 67% and reducing the number of chunks by 40–60% while maintaining 93–100% answer coverage. IDC demonstrates that aligning document structure with anticipated information needs can significantly boost retrieval performance for long and heterogeneous documents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 25

ChunkRank: Model-Aware Text Chunking and Abstention-Aware Answer Selection for LLM Pipelines

ChunkRank is an open‑source Python library that automatically determines chunk boundaries based on a target model’s tokenizer and context window, and then selects an answer from independently produced chunk candidates. It includes a registry of 90 models from 15 providers and six answer‑selection methods, and requires only three core dependencies. Experiments show that token‑exact budgeting is important across 11 languages, and that for several datasets no content‑based ranker outperforms simply taking the first non‑empty answer due to reader abstention on chunks lacking the answer.

By Amit Nautiyal, Ayush Bhatt, Gaurav Nautiyal
arXiv AI
Sep 16

ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation

The paper introduces ORDER, a task‑conditioned retrieval‑augmented generation framework that dynamically adapts both indexing and retrieval strategies to each incoming query. It first clusters questions to learn cluster‑specific chunking, metadata filtering, and reranking settings, then routes queries to the appropriate pre‑built index via nearest‑centroid assignment. Additionally, a supervised query router predicts relevant collections and a Uniform Multi‑source Sampler distributes the retrieval budget evenly across selected sources, yielding superior performance on heterogeneous historical archives compared to existing RAG systems.

By Aur\'elien Pellet (LRE), Julien Perez, Marie Puren
arXiv AI
Sep 4

STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

The paper introduces STAIR, a retrieval system that uses a document’s Table of Contents to guide large language models in accessing global structure, thereby reducing hallucinations in Retrieval Augmented Generation. Experiments with a fine‑tuned Differentiable Search Index show that ToC‑based retrieval yields a low hallucination rate (<0.05%) and improves Recall@1 to 82.6% on the newly released SearchTome benchmark, outperforming baselines like BM25, DPR, and Mistral. The authors also release SearchTome, a diverse dataset of 18 books across six domains, to encourage further research in ToC‑based retrieval.

By Vineet Kumar, Meghanadh Pulivarthi, vishwajeet kumar, Jaydeep Sen, Riyaz Ahmad Bhat, Sachindra Joshi