DOPD: Dual On-policy Distillation
arXiv:2606. 30626v1 Announce Type: new Abstract: On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals.
Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.
arXiv:2606. 30626v1 Announce Type: new Abstract: On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals.
arXiv:2606. 30190v1 Announce Type: cross Abstract: Existing domain-incremental learning (DIL) strategies call for massive amounts of data to adapt to new domains and suffer from the overfitting problem in the case of data scarcity.
arXiv:2606. 30111v1 Announce Type: cross Abstract: Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules.
arXiv:2606. 30296v1 Announce Type: new Abstract: Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many reflection rounds on one task are discarded before the next begins.
arXiv:2606. 29879v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning.
arXiv:2606. 29431v1 Announce Type: new Abstract: Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image.
arXiv:2507. 00263v2 Announce Type: replace-cross Abstract: The rapid growth of vacation rental (VR) platforms has led to an increasing volume of property images, often uploaded without structured categorization.
arXiv:2606. 29667v1 Announce Type: cross Abstract: The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inaccessible to AI at scale.
arXiv:2604. 07753v2 Announce Type: replace-cross Abstract: Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient conflicts.
arXiv:2606. 30319v1 Announce Type: cross Abstract: Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience.
arXiv:2606. 29719v1 Announce Type: new Abstract: Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it.
arXiv:2603. 13326v2 Announce Type: replace-cross Abstract: Multimodal Transformers often produce predictions without clarifying how different modalities jointly support a decision.
arXiv:2606. 30113v1 Announce Type: cross Abstract: Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actions from discrete codes remains challenging.
arXiv:2509. 12159v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models have demonstrated exceptional performance in UI2Code tasks, significantly enhancing website development efficiency.
arXiv:2604. 28123v3 Announce Type: replace-cross Abstract: The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiable rewards (RLVR).
arXiv:2606. 29888v1 Announce Type: new Abstract: Vision-language models map images and text into a joint embedding space.
arXiv:2606. 29788v1 Announce Type: new Abstract: When a multimodal AI agent is asked to forget a fact, current memory systems usually delete the text entry and report success.
arXiv:2604. 02327v2 Announce Type: replace-cross Abstract: Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification, and segmentation.
arXiv:2606. 29949v1 Announce Type: cross Abstract: H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-seq provides transcriptome-wide resolution at high cost with limited archival availability.
arXiv:2606. 30645v1 Announce Type: cross Abstract: Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion.