arXiv:2608. 04381v1 Announce Type: cross Abstract: Self-supervised learning on graphs is largely shaped by contrastive methods that depend on carefully designed augmentations, and by generative methods that reconstruct node attributes in the input space.
By Tinghe Zhang, Jian Xu, Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Qiang Wang
arXiv:2507. 03897v3 Announce Type: replace Abstract: We introduce GenAI-Powered Inference (GPI), a statistical framework for both causal and predictive inference using unstructured data, including text and images.
By Kosuke Imai, Kentaro Nakamura
arXiv:2512. 23818v2 Announce Type: replace-cross Abstract: Denoising and score estimation are classically linked through Tweedie's formula, which relates the posterior mean under Gaussian noise to the Stein score of the noisy marginal.
By Andrej Leban
arXiv:2608. 04327v1 Announce Type: new Abstract: Explicit onshore tsunami inundation forecasting can improve public risk awareness, but deterministically predicted inundation boundaries under highly uncertain conditions, such as near-field tsunamis generated by megathrust earthquakes, may falsely imply safety outside the boundaries.
By Yusuke Oishi, Takashi Furumura, Fumihiko Imamura
arXiv:2608. 05084v1 Announce Type: new Abstract: Diffusion policies are a powerful policy class for continuous control, but their iterative denoising process creates a substantial computational bottleneck.
By Rohit Kumar Salla, Manoj Saravanan, Simon Stepputtis
arXiv:2608. 04448v1 Announce Type: cross Abstract: Diffusion transformers deliver strong image generation, but their training cost grows superlinearly with resolution.
By Seunghyun Ji
arXiv:2608. 04504v1 Announce Type: cross Abstract: Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while underestimating sparse but decision-critical contextual variables.
By Shuo Liu, Huixiang Cai, Weiru Zhang, Xiaoyi Zeng
Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance. To mitigate visual noise and address privacy concerns, recent work has shifted to pose-based VAD, which focuses on motion dynamics rather than raw video data.
Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context, making such content a difficult test case for video-language models.
We introduce the neural echo as a tool for understanding the behavior of neural networks. It generalizes the model-based concepts of impulse responses, diffusion echoes, and filter echoes to learning-based methods.
Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence. Highly realistic synthetic imagery produced for malicious purposes (deepfakes) can have major consequences in the remote sensing domain, where this data is a fundamental source of information for science applications, planning, logistics, and monitoring.
We introduce the Intrinsic Hybrid Latent Diffusion Model (ILDM), a generative framework that integrates probabilistic dimensionality reduction with geometry-aware diffusion on unknown manifolds. While diffusion models (DMs) have achieved state-of-the-art results in high-dimensional data synthesis, they rely on large training datasets and ignore intrinsic geometric structure.
Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods focus primarily on salient regions and therefore have limited sensitivity to the global relationships among the main image components.
Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models frequently misinterpret the metaphorical vehicle and confuse it with the object.
Zero-shot Skeleton Action Recognition (ZSAR) remains ambiguous when unseen actions share similar skeleton joint dynamics but differ in objects or scene context. RGB provides these missing cues, yet existing multimodal methods typically maintain independent skeleton and RGB scoring branches and fuse their outputs.
The style of a painting is not monolithic: color, texture, and structure may come from different sources. Existing reference-guided methods transfer them as one style signal, leaving each attribute's source and strength outside the user's control.
We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, then retrieves instances from other videos or images that depict that state. Unlike action anticipation, which predicts a label, moment retrieval, which localizes an observed event within a video, or video generation, which synthesizes pixels, PSR combines anticipation with cross-instance retrieval across multiple temporal horizons.
arXiv:2608. 03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped.
By Pei Li, Sihan Chen, Delong Ran, Tianshuo Cong
arXiv:2608. 03103v1 Announce Type: cross Abstract: Diffusion policies have shown strong performance in learning complex, multi-modal behaviors for robotic manipulation.
By Rishabh Shukla, Adithya Santhosh, Shaili Gandhi, Samrudh Moode, Satyandra K. Gupta
arXiv:2508. 01597v2 Announce Type: replace Abstract: Score Matching (SM) is a powerful framework for estimating the log-density derivatives of a distribution without calculating its normalizing constants.
By Juyan Zhang, Rhys Newbury, Xinyang Zhang, Tin Tran, Dana Kulic, Michael Burke