arXiv AI By Xuehang Guo, Pengyuan Li, Tom Hope, Tirthankar Ghosal, Manling Li, Qingyun Wang

Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning

Read the original on arXiv AI →

arXiv:2608. 04926v1 Announce Type: cross Abstract: As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross-representation understanding across these modalities poses fundamental challenges for AI systems: the relationships across representations are inherently \textit{one-to-many}, supervision is ambiguous and costly, and model optimization lacks a principled signal that is both direction-adaptive and representation-generalizable beyond task-specific objectives.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models

The paper introduces MATE, a reinforcement‑learning‑based post‑training framework for unified multimodal models that lets the generation and understanding branches challenge each other instead of cooperating. In MATE, each branch proposes candidate outputs that the other must reproduce, and the solver is trained on the worst‑handled candidate, creating an evolving adversarial loop without a separate adversary. Experiments on Janus‑Pro‑1B show that MATE improves generation and understanding metrics, including GenEval (+2.4), DPG‑Bench (+1.7), and an average of nine understanding benchmarks (+0.7), while enhancing consistency across image‑text cycles.

By Wentao Zhou, Weijie Gan, Jiayun Wang
arXiv Machine Learning
Jun 25

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning

arXiv:2507. 16518v3 Announce Type: replace-cross Abstract: Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities.

By Xiuwei Chen, Wentao Hu, Hanhui Li, Yongxin Wang Jun Zhou, Zisheng Chen, Meng Cao, Yihan Zeng, Kui Zhang, Yu-Jie Yuan, Jianhua Han, Hang Xu, Xiaodan Liang
arXiv AI
Sep 7

Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study

The study investigates how unified vision‑language models (VLMs) can simultaneously support visual understanding and generation. Using controlled benchmarks (SmartWatch and modified CelebA) that pair VQA, captioning, and text‑to‑image tasks, the authors evaluate several LLM‑based architectures built on SigLIP and VQ‑VAE visual spaces. Results show that mixed training can improve both understanding and generation, but the gains depend on how well the visual input and output spaces are aligned; misaligned or distorted visual spaces can weaken or reverse these benefits. The paper also demonstrates that balancing data across tasks and controlling attribute frequencies can help recover underrepresented visual concepts, and that the transfer is driven more by the base language model’s learned relationships than by visual adapters.

By Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang, Yu Cheng
arXiv AI
Aug 12

VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

arXiv:2511. 19436v2 Announce Type: replace-cross Abstract: Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary models, creating a dependency on external supervision.

By Qiang Wang, Xinyuan Gao, Yuhang He, Jizhou Han, Jiangyang Li, SongLin Dong, Zhiheng Ma, Yihong Gong