arXiv AI
Jun 30

ManimAgent: Self-Evolving Multimodal Agents for Visual Education

arXiv:2606. 30296v1 Announce Type: new Abstract: Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many reflection rounds on one task are discarded before the next begins.

By Wenjia Jiang, Zongyuan Cai, Yuanhang Shao, Chenru Wang, Boyan Han, Zhixue Song, Keyu Chen, Shengwei An, Xu Yang, Zhou Yang
arXiv Machine Learning
Sep 3

DynaTokens: Controlling Token Dynamics for Continual Video-Language Understanding

The paper introduces DynaTokens, a transformer-based token generator that produces fine‑tuning tokens on demand for continual VideoQA with multimodal large language models. It uses shared generation weights and meta‑learning‑inspired regularisers to reduce task interference and forgetting, connecting the objective to sharpness‑aware optimisation for flatter cross‑task minima. Experiments on standard continual VideoQA benchmarks show that DynaTokens achieves higher average accuracy, lower forgetting, better zero‑shot generalisation, and robust cross‑modal transfer in a new ImageQA→VideoQA protocol.

By Toan Nguyen, Yang Liu, Celso De Melo, Flora D. Salim