arXiv Machine Learning

Coarse-to-Fine Compositional Diffusion for Long-Horizon Planning

arXiv:2606. 00837v1 Announce Type: cross Abstract: Diffusion models provide strong priors for generating structured data, but many tasks require outputs beyond the scale on which these models are typically trained.

Hugging Face Trending Papers
Jul 9

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasize task-critical evidence and ignore underlying factors.