AgentOCR: Reimagining Agent History via Optical Self-Compression
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608.29897v1 Announce Type: new Abstract: Long-horizon agents need a context manager to compress growing interaction histories into a bounded working context, via passive strategies or active s...
arXiv:2608. 08960v1 Announce Type: new Abstract: Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs.
arXiv:2609.11899v1 Announce Type: new Abstract: Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses t...
arXiv:2606. 30296v1 Announce Type: new Abstract: Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many reflection rounds on one task are discarded before the next begins.
arXiv:2605. 14211v3 Announce Type: replace Abstract: Long-horizon embodied tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrations, neither of which scales.
The paper introduces DynaTokens, a transformer-based token generator that produces fine‑tuning tokens on demand for continual VideoQA with multimodal large language models. It uses shared generation weights and meta‑learning‑inspired regularisers to reduce task interference and forgetting, connecting the objective to sharpness‑aware optimisation for flatter cross‑task minima. Experiments on standard continual VideoQA benchmarks show that DynaTokens achieves higher average accuracy, lower forgetting, better zero‑shot generalisation, and robust cross‑modal transfer in a new ImageQA→VideoQA protocol.