Hugging Face Trending Papers

Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation

Read the original on Hugging Face Trending Papers →

While vision-language models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this to a perception-reasoning modality gap: visual planning requires models to infer latent state structures from pixels and then reason over the recovered structure to produce valid actions, whereas symbolic planning directly leverages explicit objects and constraints.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.