Did the Grid Erase the Event? EndoClock for Auditing Medical World-Model Pipelines
arXiv:2608. 09266v1 Announce Type: cross Abstract: Medical world models commonly learn from multimodal recordings synchronized onto a fixed-rate grid.
Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.
arXiv:2608. 09266v1 Announce Type: cross Abstract: Medical world models commonly learn from multimodal recordings synchronized onto a fixed-rate grid.
arXiv:2608. 07565v1 Announce Type: cross Abstract: Conversational assistants increasingly recommend follow-up edits to help users continue a task.
arXiv:2604. 11741v2 Announce Type: replace Abstract: Vision-language models (VLMs) have shown impressive capabilities in perceptual tasks, yet they degrade in complex multi-hop reasoning under multiplayer game settings with imperfect and deceptive information.
arXiv:2608. 07543v1 Announce Type: cross Abstract: Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis.
arXiv:2608. 07533v1 Announce Type: new Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body.
arXiv:2608. 09053v1 Announce Type: cross Abstract: Cardiologists interpret electrocardiograms by localizing waveform components, measuring rhythm and interval patterns, and translating these structured observations into diagnostic evidence.
arXiv:2606. 27826v3 Announce Type: replace Abstract: Embodied agents driven by multimodal large language models (MLLMs) can often complete everyday tasks from visual observations, but goal achievement does not establish whether they proactively respect unstated social norms.
arXiv:2608. 08009v1 Announce Type: cross Abstract: Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for Detecting and Grounding Multi-Modal Media Manipulation (DGM4).
arXiv:2512. 01045v2 Announce Type: replace Abstract: Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the construction of such datasets often remains labor-intensive, weakly traceable, and difficult to configure.
arXiv:2605. 15532v3 Announce Type: replace-cross Abstract: Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically chosen via simple heuristics or aggregated from off-the-shelf datasets.
arXiv:2608. 09270v1 Announce Type: cross Abstract: Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation.
arXiv:2605. 30716v2 Announce Type: replace-cross Abstract: Generating clinically useful pathology reports for pathology cases from whole-slide images (WSIs) is challenging due to gigapixel resolution, long visual-token sequences, and the complexity of case-level reasoning, where a single case may contain multiple WSIs with heterogeneous tissues and ambiguous findings.
arXiv:2608. 08557v1 Announce Type: cross Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding.
arXiv:2512. 08462v2 Announce Type: replace Abstract: Decoding brain states from functional magnetic resonance imaging (fMRI) data is vital for advancing neuroscience and clinical applications.
arXiv:2608. 09482v1 Announce Type: cross Abstract: All-in-one image restoration is a unified low-level vision task that aims to effectively recover high-quality images from inputs degraded by various types and levels of corruption using a single model.
arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.
arXiv:2608. 09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image.
arXiv:2608. 08852v1 Announce Type: new Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content.
arXiv:2608. 07525v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications.
arXiv:2608. 08794v1 Announce Type: new Abstract: Omni-modal LLMs jointly process audio, video, and text, but long multimodal sequences incur substantial prefill and KV-cache costs.