arXiv AI

Improving Attributed Long-form Question Answering with Intent Awareness

arXiv:2603. 27435v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports.

Hugging Face Trending Papers
Jul 23

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning

Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context when evidence is sparse, noisy, or in conflict with parametric knowledge. Existing grounding methods either attach citations after generation or encourage evidence retrieval inside the trace, but they often do not ensure that cited content is sufficient for the local inference and final answer.

arXiv AI
Aug 19

Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs

Intent-Driven Dynamic Chunking (IDC) segments documents by predicting user queries with a Large Language Model and then applying dynamic programming to find optimal chunk boundaries. This method outperforms traditional fixed-length or coherence-based segmentation on five out of six question-answering datasets, improving top-1 retrieval accuracy by 5% to 67% and reducing the number of chunks by 40–60% while maintaining 93–100% answer coverage. IDC demonstrates that aligning document structure with anticipated information needs can significantly boost retrieval performance for long and heterogeneous documents.

By Christos Koutsiaris
arXiv Computation and Language
Sep 17

Reading Between the Lines: Can LLMs Discover the Question Behind the Text?

The paper introduces "question archaeology," an evaluation task that asks models to infer the single, authentic question that motivated a text. It presents a new dataset of commissioned texts paired with their original research questions and distractors, and evaluates both proprietary and open‑source LLMs. Results show newer models outperform older ones, with BERT-based models lagging, and current LLMs even surpassing human performance on this task.

By Claudiu Creanga, Liviu P. Dinu
arXiv AI
Sep 25

Learning to Ideate for Scientific Impact

The paper "Learning to Ideate for Scientific Impact" explores using delayed signals of scientific uptake—specifically citation-normalized impact—as feedback to steer large language models toward generating high‑impact research ideas. The authors build a dataset of over 100,000 computer science papers, train a reward model to predict citation impact from goal‑idea pairs, and align an idea generator via supervised fine‑tuning and reinforcement learning. Evaluation with a reference‑grounded protocol shows that the RL‑tuned model consistently produces ideas with higher estimated impact than baseline models.

By Shubham Kale, Aniketh Garikaparthi, Manasi Patwardhan
arXiv Computation and Language
Aug 25

ConvergeWriter: Data-Driven Bottom-Up Article Construction

ConvergeWriter introduces a bottom‑up, data‑driven framework for long‑form document generation that first retrieves exhaustive knowledge from a source corpus and clusters it into distinct knowledge groups. These clusters then guide the creation of a hierarchical outline and the final text, ensuring the output is strictly grounded in the retrieved material and traceable to its sources. Experiments on 14B and 32B LLMs show that this approach matches or surpasses state‑of‑the‑art baselines, especially in scenarios requiring high factual fidelity and structural coherence.

By Binquan Ji, Jiaqi Wang, Ruiting Li, Xingchen Han, Yiyang Qi, Shichao Wang, Yifei Lu, Yuantao Han, Feiliang Ren
Hugging Face Trending Papers
Aug 19

DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering

DeepWeaver addresses the evidence synthesis gap in open‑ended question answering by weaving noisy retrieved evidence into comprehensive answers. It introduces Thought Block Chains (TBCs) that organize claims, key information, and citations, allowing the system to revise and expand evidence before final generation. Evaluations on LoQA and DeepResearch Bench show improved content sufficiency, citation grounding, and detail preservation across multiple LLMs.