arXiv:2609.00756v1 Announce Type: new
Abstract: The Mutual Reinforcement Effect (MRE) asks whether a fine, span-level and a coarse, document-level task help each other when one model handles both. We...
By Chengguang Gan, Yunhao Liang, Hanjun Wei, Qinghao Zhang, Shiwen Ni
The paper introduces GRADE, a graph-based representation of large language model (LLM) agent executions that captures both execution steps and their dependencies. By adding graded dependency edges—observed, declared, or inferred—to the trace, the authors evaluate how this dependency layer affects failure prediction across six corpora involving tool use, coding, and web tasks. Experiments show that the dependency block can improve prediction in some settings, but its effectiveness varies with the evaluation probe and corpus, and controlled experiments demonstrate that the observed structure is not merely a degree-matched artifact.
By Yue Zhao
The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.
By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
arXiv:2607. 21692v2 Announce Type: replace Abstract: Sparse attention prunes a long context to the blocks a model needs, and the usual selector is distilled from a dense teacher's attention.
By Jim Allchin
arXiv:2607. 13188v1 Announce Type: new Abstract: Human cognition does not separate understanding and generation.
By Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro V\'elez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
arXiv:2607. 18293v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) teaches large language models new skills through a teacher that shares the student's backbone and supervises its own rollouts.
By Yingzi Ma, Zichen Zhu, Ming Jiang, Chaowei Xiao
arXiv:2607. 16451v1 Announce Type: cross Abstract: Chat models sometimes commit to an answer and then produce reasoning that justifies it rather than deriving it -- even when the answer contradicts a task premise.
By Heejin Jo
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws $\textit{together}$, each modality reshapes the other.
The paper reports on a deployed multi‑agent tender‑response system that uses an open‑weights language model under sovereignty constraints. In a blind comparison, the system’s answers were judged at least as good as human‑written bids in 40 of 55 sections, with only a few gaps attributable to missing knowledge rather than writing quality. The study also demonstrates an asymmetry in conditioning: while structural markup improves reading tasks, converting instruction material from prose to nested XML degrades answer quality, and naming forbidden constructions concentrates defects.
By Cheng Yu, Nikhil Mathew, Zhengjie Wang
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
arXiv:2606. 11198v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) systems inject external knowledge to improve LLM outputs, yet the format of injected content -- distinct from its semantic relevance -- can independently distort the model's attention distribution.
By Yuqi Zhang, Di Zhang
The paper investigates neural text degeneration by measuring the fixed‑point structure of short‑window argmax maps across 17 pretrained models, using 96 random two‑token starts without prompts. It finds a stable four‑way classification that varies across model families and scales, with some models funneling to a single endpoint token while others do not, and shows that this behavior is not solely determined by training data or corpus frequency. The study demonstrates that repetition phenomena are not uniformly explained by either training data or network architecture alone, highlighting the complexity of neural text generation dynamics.
By Nicol\'as Vera Z\'u\~niga