arXiv Machine Learning

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

arXiv:2607. 13188v1 Announce Type: new Abstract: Human cognition does not separate understanding and generation.

Hugging Face Trending Papers
Aug 10

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models.