arXiv Computation and Language By Edward Xi Yang (Ertas AI)

Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum

Read the original on arXiv Computation and Language →

The paper introduces a retrieved‑span training approach for query‑focused meeting summarization on the QMSum benchmark. By fine‑tuning a 406 M Fusion‑in‑Decoder model on 2,000‑word retrieved spans, the authors recover a 6.30 ROUGE‑1 loss incurred when moving from capped long input and achieve a test score of 36.33 ROUGE‑1, comparable to a larger 1.2 B system. The smaller model uses roughly one‑third the parameters and less than half the peak inference memory, while span‑regime fine‑tuning adds significant gains over the baseline. "whyItMatters":"The study demonstrates that efficient, smaller models can match or exceed larger systems on QMSum using span‑based fine‑tuning, offering a practical path for scalable query‑focused meeting summarization."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 22

BudgetMem: Training-Free Selective Memory for Cost-Efficient Long-Context Processing in Language Models

arXiv:2511. 04919v3 Announce Type: replace Abstract: Processing long documents with large language models (LLMs) is expensive: a single query over a 100K-token document can cost from tens of cents to over a dollar in API fees, depending on the model, and memory grows linearly with context length.

By Chandra Vamsi Krishna Alla, Harish Naidu Gaddam, Manohar Kommi, Sheikh Nazib Ahmed
arXiv Machine Learning
Sep 18

Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute

The paper investigates whether confidence signals from fine‑tuned large language models can improve extractive question answering that relies heavily on retrieval. Experiments on four 7‑9B model families show that retrieval alone recovers 92–99.8% of the best possible accuracy, leaving little room for confidence‑based routing or adaptation to help. The sequence‑likelihood confidence metric, even after recalibration or temperature scaling, fails to provide a statistically significant benefit across different correctness criteria and answer lengths, and the study ultimately offers a set of pre‑specified negatives with explicit dependencies as its main contribution.

By Gunwoo Lee, Changmin Sung, Sang-Hwan Gwak, Ina Kim, Ji-Young Choi, Kyong-Ha Lee