Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models
arXiv:2608. 05960v1 Announce Type: cross Abstract: Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume.
Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.
arXiv:2608. 05960v1 Announce Type: cross Abstract: Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume.
arXiv:2608. 06110v1 Announce Type: new Abstract: This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management.
arXiv:2608. 06075v1 Announce Type: cross Abstract: Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems.
arXiv:2608. 05729v1 Announce Type: new Abstract: As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time.
arXiv:2506. 02260v5 Announce Type: replace-cross Abstract: Wearable devices enable continuous multi-modal physiological and behavioral monitoring, yet analysis of these data streams faces fundamental challenges including the lack of gold-standard labels and incomplete sensor data.
arXiv:2608. 05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.
arXiv:2608. 06154v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception.
arXiv:2608. 05798v1 Announce Type: cross Abstract: Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation.
arXiv:2608. 05381v1 Announce Type: new Abstract: Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning.
arXiv:2608. 05341v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings are left unreported due to the omission of subtle findings.
arXiv:2608. 05833v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images.
arXiv:2608. 05792v1 Announce Type: new Abstract: Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer technology into a goal-driven, closed-loop intelligent system, a paradigm we term AISAC.
arXiv:2608. 06023v1 Announce Type: new Abstract: To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD.
arXiv:2608. 05238v1 Announce Type: new Abstract: Training multimodal models to align time series with language runs into a self-supervision trap.
arXiv:2608. 05454v1 Announce Type: new Abstract: Probabilistic prediction heads in neural networks typically output either a Gaussian mixture or a single conformal region.
arXiv:2608. 05668v1 Announce Type: cross Abstract: With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-quality financial trading.
Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural question: do they reduce the need for classic, physics-informed low-level vision?
To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD. The proposed framework leverages physiological signals as privileged information during training to guide a video-based student model in learning deep affective representations, while relying solely on non-intrusive video inputs at inference time.
Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires accurate instantaneous channel state information (CSI) and iterative optimization, which incur substantial pilot overhead and computational complexity.
Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings.