From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds
arXiv:2606. 03557v1 Announce Type: new Abstract: As generative AI capabilities expand, AI-driven virtual worlds face a growing architectural challenge.
Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.
arXiv:2606. 03557v1 Announce Type: new Abstract: As generative AI capabilities expand, AI-driven virtual worlds face a growing architectural challenge.
arXiv:2606. 03273v1 Announce Type: cross Abstract: Visual DeepSearch requires multimodal large reasoning model (MLRM) agents to answer complex visual queries by repeatedly inspecting image regions, grounding intermediate reasoning in visual evidence, and connecting fine-grained clues across long reasoning chains.
arXiv:2606. 03954v1 Announce Type: cross Abstract: As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible consequences that digital errors do not.
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.
arXiv:2601. 14569v2 Announce Type: replace-cross Abstract: Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions.
arXiv:2606. 02892v1 Announce Type: new Abstract: Breast cancer recurrence, a leading cause of long-term mortality among survivors, requires timely and accurate risk assessment to guide follow-up care and treatment planning.
arXiv:2606. 03626v1 Announce Type: cross Abstract: Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks.
arXiv:2606. 03100v1 Announce Type: cross Abstract: Recently, zero-shot 3D scene understanding via 2D Vision-Language Models (VLMs) has gained increasing research interest due to their promising spatial reasoning capabilities.
arXiv:2606. 03322v1 Announce Type: cross Abstract: The graphical representation of the brain offers critical insights into diagnosing and prognosing neurodegenerative disease via relationships between regions of interest (ROIs).
arXiv:2606. 03223v1 Announce Type: cross Abstract: Robot storytelling offers a unique blend of technological innovation and creative expression that engages children in unprecedented ways.
arXiv:2606. 02739v1 Announce Type: cross Abstract: Audio tokenizers serve as the discrete interface between continuous audio and Audio Language Models (ALMs), but existing tokenizers often struggle to support both understanding and generation.
arXiv:2606. 03093v1 Announce Type: new Abstract: Prompting steers large language models (LLMs) and vision-language models (VLMs) without weight updates, but it remains unclear how instruction changes reshape internal representations to produce behavior.
arXiv:2606. 03066v1 Announce Type: new Abstract: The rapid rise of generative AI has made multimodal fake news increasingly realistic and pervasive, posing severe threats to public trust and social stability.
arXiv:2606. 02914v1 Announce Type: new Abstract: Background: Oral diseases affect nearly 3.
arXiv:2606. 03967v1 Announce Type: cross Abstract: We describe AlignAtt4LLM, an IWSLT 2026 simultaneous speech translation system for English to German, Italian, and Chinese.
arXiv:2606. 03876v1 Announce Type: cross Abstract: With the growing prevalence of modern ubiquitous computing technologies, multi-modal tracking systems hold promise for providing timely awareness and reassurance to stakeholders such as remote family members (RFMs) of older adults, who play a central role in care coordination.
arXiv:2606. 03180v1 Announce Type: cross Abstract: Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows.
arXiv:2606. 03486v1 Announce Type: cross Abstract: Large language models remain vulnerable to jailbreak attacks that hide harmful intent behind seemingly ordinary requests such as role-play, translation, encoding, adversarial suffixes, and multi-turn buildup.
arXiv:2604. 18572v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.
arXiv:2512. 03627v2 Announce Type: replace Abstract: Despite rapid progress in large-scale language and vision models, AI agents still suffer from a fundamental limitation: they cannot remember.