Pareto-Guided Optimal Transport for Multi-Reward Alignment
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2604. 17415v3 Announce Type: replace-cross Abstract: Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model.
arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.
arXiv:2608.30194v1 Announce Type: new Abstract: Diffusion models have recently advanced text-to-video (T2V) generation, yet they still struggle with fine-grained compositional alignment, such as attr...
arXiv:2601.03468v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a standard approach for post-training large language models and, more recently, for improving image generati...
arXiv:2607. 15740v1 Announce Type: cross Abstract: As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI.
arXiv:2608.29647v1 Announce Type: new Abstract: To mitigate the time complexity of generative models, one-step generative models have recently emerged through direct mapping from noise to data in a s...