Hugging Face Blog
Apr 11, 2024
The paper introduces Alignment‑Guided Flow Transformer (AGFT), a framework for Vision‑Language‑Action (VLA) models that explicitly enforces tri‑modal alignment among vision, language, and action through a dedicated alignment loss. AGFT bridges representational gaps across modalities, improving task adaptation and robustness, and employs a flow‑matching objective to reduce inference steps compared to diffusion‑based policies. Experiments on a large benchmark demonstrate that AGFT achieves higher success rates and lower inference latency than state‑of‑the‑art baselines, highlighting tri‑modal alignment as crucial for scalable VLA manipulation.