arXiv Machine Learning By Zewen Liu

Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents

Read the original on arXiv Machine Learning →

arXiv:2606. 16682v1 Announce Type: new Abstract: When AI agents use language models to evaluate their own outputs in a feedback loop, systematic biases emerge.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration

ReNFT is a method that repairs mode collapse in diffusion generators after reward post‑training by internally recalibrating probability mass. It identifies suppressed alternatives through anti‑hub prompts and uses two policy‑dominated routes to generate counterfactual proposals, then applies reward‑based ranking and a joint‑and‑paired NFT update to restore diversity while preserving reward. Experiments on PickScore and GenEval show that ReNFT retains almost all of the original reward while significantly boosting diversity metrics.

By Yuchen Bao, Chao Wen, Haowei Wang, Ruoxin Chen, Donghao Luo, Jiahui Zhan, Wenjian Huang, Shen Chen, Yiting Wang, Taiping Yao, Chengjie Wang, Shouhong Ding, Jianguo Zhang