VISA (Visual Instruction Synthesis Agent) is an agentic framework that transforms multimodal instruction synthesis into a self‑evolving loop. Each cycle analyzes images to filter constraints, samples new constraint sets, generates candidate instructions, and verifies them using executable tools and large language model judges. Failed samples trigger diagnostic recovery, while accepted samples are evaluated against the target model to estimate difficulty, with all feedback written back to memory to adapt future rounds and provide reward signals for reinforcement learning.
By Min Zeng, Guanxin Tan, Libin Cen, Yawei Wen, Rui Hu, Liuyang Bian, Xiaolong Chen, Xiaoxin Chen
OmniHarness is a framework that enables generalizable visual generation by learning symbolic policies from verified executions. It abstracts shared procedures and applicability conditions, allowing these policies to be instantiated, adapted, and composed for new tasks while keeping model parameters fixed. The system uses intermediate verification for refinement, self-directed inquiry to generate practice tasks, and continuous feedback to expand capabilities, achieving strong results on multiple benchmarks and outperforming baselines on Creative tasks.
By Xu Xu (Beihang University), Jinxiu Liu (The Chinese University of Hong Kong), Zhangbo Qiao (Beihang University), Jiaxing Lu (Beihang University), Xiangyu Zhang (Beihang University), Yubin Gu (National University of Singapore), Fangwei Ning (Beihang University), Yan Shi (Beihang University)
arXiv:2609.37923v1 Announce Type: new
Abstract: Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce E...
By Ziyun Zeng, Hang Hua, Shaden Alshammari, Rogerio Feris, William T. Freeman, Jiebo Luo
arXiv:2608.31022v1 Announce Type: new
Abstract: AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However,...
By Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria
arXiv:2606. 16496v1 Announce Type: cross Abstract: Large multimodal language models (LLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies.
By Pan Wang
arXiv:2607. 01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance.
By Yuanzhi Liu, Shousheng Zhao, Bo Zhou, Kongming Liang, Zhanyu Ma