Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution
arXiv:2608. 03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.
Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.
arXiv:2608. 03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.
arXiv:2608. 03023v1 Announce Type: cross Abstract: Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods.
arXiv:2608. 03150v1 Announce Type: new Abstract: Generative retrieval (GR) is a promising paradigm for industrial search advertising, yet its deployment is constrained by strict relevance and latency requirements.
arXiv:2608. 03079v1 Announce Type: cross Abstract: Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions.
arXiv:2607. 27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning.
arXiv:2608. 03450v1 Announce Type: cross Abstract: Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction.
arXiv:2608. 03979v1 Announce Type: cross Abstract: We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration.
arXiv:2608. 03451v1 Announce Type: new Abstract: Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia.
arXiv:2608. 03050v1 Announce Type: cross Abstract: What is music style?
arXiv:2608. 03327v1 Announce Type: new Abstract: Hybrid computer-use agents can act through screenshots or call text tools.
arXiv:2508. 05149v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks.
arXiv:2608. 03733v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate.
arXiv:2608. 02688v1 Announce Type: cross Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses.
arXiv:2608. 02958v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress.
arXiv:2608. 03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes.
arXiv:2608. 03817v1 Announce Type: cross Abstract: Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence.
arXiv:2608. 02830v1 Announce Type: cross Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied.
arXiv:2608. 03207v1 Announce Type: cross Abstract: Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs.
arXiv:2608. 03918v1 Announce Type: cross Abstract: Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence.
arXiv:2608. 02769v1 Announce Type: cross Abstract: Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance.