Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,971 stories · RSS feed

arXiv AI
Sep 30

Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning

arXiv:2609.37519v1 Announce Type: cross Abstract: Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched...

By Merve Atasever, Keyan Azbijari, Cagan Bakirci, Bo-Ruei Huang, Tolga Izdas, Zahra Shahrooei, Richard Yang, Erdem Biyik, Jyotirmoy V. Deshmukh