arXiv Machine Learning By Aleesha Zainab, Asifullah Khan, Muhammad Ahmed Khalid, Faheem Ullah Khan

A Channel-Boosted Multi-Agent System with Iterative Consultation for Document Sensitivity Classification

Read the original on arXiv Machine Learning →

The paper introduces a Channel‑Boosted Multi‑Agent System (CB‑MAS), instantiated as IC‑MAS, to classify document sensitivity without the input‑length truncation problem of standard transformers. IC‑MAS uses a Channel Critic Agent to adaptively weight two first‑window encoders and Consultation Agents to exchange belief states, achieving 90.72% accuracy and 91.23% F1‑score while reducing computation by ~54% compared to a fixed‑round baseline. The authors provide explainability via LIME/SHAP, multi‑agent evaluation, and a transparent discussion of limitations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 4

Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning

The paper introduces LRE (Learned Relevance Eviction), a lightweight, CPU‑only, language‑model‑free scorer that learns which parts of an agent’s interaction history are task‑critical and preserves them verbatim. In experiments, LRE matches or surpasses baseline eviction policies on accuracy‑cost trade‑offs, recovers 93% of full‑history accuracy, reduces worst‑case prompt size by 52%, and outperforms dense and token‑pruning encoders in conversational memory while being 295–1569× smaller. The method also achieves superior budgeted answer quality on LoCoMo reading and can be trained annotation‑free, recovering 95% of supervised scorer performance.

By Nusrat Jahan Lia, Aritra Mazumder
arXiv Computation and Language
Sep 11

DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

Deep Research Bench II is a new benchmark designed to evaluate Deep Research Agents (DRAs) by requiring them to produce research reports for 132 grounded tasks across 22 domains. Each report is assessed using 9,430 fine‑grained binary rubrics that cover information recall, analysis, and presentation, all derived from expert‑written investigative articles through a rigorous LLM‑plus‑human pipeline. Evaluation of current state‑of‑the‑art DRAs shows that even the best models satisfy fewer than 50% of these rubrics, highlighting a significant gap between automated agents and human experts.

By Ruizhe Li, Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao