DocAtlas: Long-Document Understanding as Mutable-State Interaction
arXiv:2608. 07527v1 Announce Type: cross Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts.
arXiv:2607. 22644v1 Announce Type: new Abstract: Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or type.
arXiv:2608. 07527v1 Announce Type: cross Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts.
arXiv:2607. 18358v1 Announce Type: cross Abstract: Document classification is a solved problem in the laboratory and an unsolved one in the enterprise.
PaperScout is an autonomous agent that treats academic paper search as a sequential decision-making process, dynamically deciding when and how to use search and expansion tools based on accumulated context. The authors identify a granularity mismatch in standard reinforcement learning for multi-turn tasks and propose Proximal Sequence Policy Optimization (PSPO), a sequence-level policy optimization method that aligns learning with agent–environment interactions. Experiments on synthetic and real-world benchmarks show that PaperScout outperforms structured retrieval and RL baselines in recall and relevance, demonstrating the effectiveness of its adaptive agentic framework and optimization strategy.
arXiv:2602. 10238v2 Announce Type: replace-cross Abstract: The growing size of Large Language Models (LLMs) makes efficient inference challenging, primarily due to the memory demands of the autoregressive Key-Value (KV) cache.
arXiv:2607. 20497v1 Announce Type: new Abstract: Prompt optimization for text classification spans diverse approaches, from demonstration selection to exploration-based search to error-driven diagnosis, each with known but incompletely characterized strengths and limitations.
arXiv:2605.29307v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) search agents have shown strong promise on knowledge-intensive tasks through iterative reasoning and retrieval. Mo...
GrammarRL introduces a label‑free reinforcement learning approach that adapts language models to grammar constraints without annotated data. It optimizes two self‑supervised rewards—direct and reverse—using a Reinforce Leave‑One‑Out objective over grammar‑constrained rollouts, and regularizes toward a frozen base model. Experiments on sign‑language gloss translation, hierarchical text classification, and named entity recognition with Llama models show consistent gains over constrained greedy decoding and competitive performance to beam search while keeping inference cost low.
arXiv:2607. 13639v1 Announce Type: cross Abstract: We introduce OvisOCR2, a 0.
arXiv:2607. 10694v1 Announce Type: cross Abstract: We study the problem of optimal continual fine-tuning for a pre-trained Foundation Model deployed at a resource-limited device.
arXiv:2608. 16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult.
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approache...
arXiv:2607. 20083v1 Announce Type: cross Abstract: Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models.