arXiv Computation and Language

Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum

The paper introduces a retrieved‑span training approach for query‑focused meeting summarization on the QMSum benchmark. By fine‑tuning a 406 M Fusion‑in‑Decoder model on 2,000‑word retrieved spans, the authors recover a 6.30 ROUGE‑1 loss incurred when moving from capped long input and achieve a test score of 36.33 ROUGE‑1, comparable to a larger 1.2 B system. The smaller model uses roughly one‑third the parameters and less than half the peak inference memory, while span‑regime fine‑tuning adds significant gains over the baseline. "whyItMatters":"The study demonstrates that efficient, smaller models can match or exceed larger systems on QMSum using span‑based fine‑tuning, offering a practical path for scalable query‑focused meeting summarization."

arXiv Computation and Language
Sep 22

BudgetMem: Training-Free Selective Memory for Cost-Efficient Long-Context Processing in Language Models

arXiv:2511. 04919v3 Announce Type: replace Abstract: Processing long documents with large language models (LLMs) is expensive: a single query over a 100K-token document can cost from tens of cents to over a dollar in API fees, depending on the model, and memory grows linearly with context length.

By Chandra Vamsi Krishna Alla, Harish Naidu Gaddam, Manohar Kommi, Sheikh Nazib Ahmed
arXiv Machine Learning
Sep 18

Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute

The paper investigates whether confidence signals from fine‑tuned large language models can improve extractive question answering that relies heavily on retrieval. Experiments on four 7‑9B model families show that retrieval alone recovers 92–99.8% of the best possible accuracy, leaving little room for confidence‑based routing or adaptation to help. The sequence‑likelihood confidence metric, even after recalibration or temperature scaling, fails to provide a statistically significant benefit across different correctness criteria and answer lengths, and the study ultimately offers a set of pre‑specified negatives with explicit dependencies as its main contribution.

By Gunwoo Lee, Changmin Sung, Sang-Hwan Gwak, Ina Kim, Ji-Young Choi, Kyong-Ha Lee
arXiv Machine Learning
Sep 14

What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework

The paper introduces LAST-CQ, a five-agent, training‑free, execution‑grounded framework for Text‑to‑Cypher that evaluates which components of an agentic pipeline contribute most to performance. Experiments on 2,471 live‑database queries across six backbones show that removing correction reduces execution‑BLEU by 3.1–12.3%, while substituting schema‑grounded feedback with raw error strings has negligible impact. Parallel sampling degrades quality by 10–11%, whereas failure detection and retry routing recover 91.7% of initially failed queries, highlighting that simple failure handling is more effective than sophisticated feedback or increased sampling.

By Ioannis Prokopiou, Athanasios Aidinis, Panagiotis-Christos Kyrmpatsos, Pantelis Vikatos
arXiv AI
Aug 26

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

RENDER is a benchmark that controls the reader‑facing artifact in memory and RAG evaluations while keeping the conversation fixed. It introduces a five‑level packet ladder and deterministic templates that mimic ChatGPT‑style entries, LangChain summaries, MemGPT‑style typed records, and raw conversation. Experiments on 500 LongMemEval questions across nine models show that matched‑budget packets outperform raw dialogue by 42.4–72.6 points, and that ChatGPT‑style entries often score higher than raw conversation, with effects persisting under retrieval noise and transferring to HotpotQA.

By Yuan Si, Simeng Han, Daming Li, Jialu Zhang
arXiv AI
Sep 28

HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents

The paper introduces Hard-Origin Adaptively Softened Memory (HasMem), a memory system for large language model agents that combines frozen hard‑prompt embeddings with a controller, writer, reader, and global module to adaptively resize and re‑encode memory entries. On a reconstruction probe of 535 questions, HasMem achieves a lexical F1 of 95.3, outperforming the hard reference by 4.4 percentage points while maintaining 93.6% of the reference’s memory positions. Across six configurations with similar per‑question budgets, the system surpasses rule‑based re‑encoding by 8.0–23.6 exact‑match points, and on LongMemEval‑S it improves local lexical F1 from 3.4 to 8.9 and reduces answer negative log‑likelihood from 12.257 to 5.274.

By Zihong He, Junxiao Shen, Chen Liang, Hai-Ning Liang