arXiv:2605. 00972v2 Announce Type: replace-cross Abstract: Earth system science is producing increasingly large, high-dimensional datasets from both physics-based and AI-driven models.
By Nihanth W. Cherukuru, Matt Rehme, Kirsten J. Mayer, David John Gagne, John Schreck, John Clyne, Charlie Becker
arXiv:2606. 26535v1 Announce Type: cross Abstract: Current VLM evaluations often conflate language priors with genuine spatial reasoning.
By Zhixing Li, Yinan Yu
Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning.
arXiv:2605.04515v2 Announce Type: replace
Abstract: Video Large Language Models (Video-LLMs) excel in general video understanding but often base physical judgments on event expectations rather than o...
By Zicheng Zhao, Chaofan Gan, Shijie Li, Weiyao Lin
HiPhy introduces a hierarchical reinforcement learning framework for video generation that enforces physical laws at both local and global levels. It addresses the challenge of multi-principle interactions—such as buoyancy and fluid dynamics occurring simultaneously—by ensuring each principle’s temporal dynamics and the overall scene’s coherence. The authors also provide a 50K-prompt dataset and the MultiPhyBench benchmark, demonstrating that HiPhy outperforms existing methods, especially in scenes with multiple concurrent physical principles.
By Tahira Kazimi, Shubhankar Borse, Munawar Hayat, Fatih Porikli, Pinar Yanardag
arXiv:2602. 06205v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces.
By Akshit Achara, Tatiana Gaintseva, Mateo Mahaut, Pritish Chakraborty, Viktor Stenby Johansson, Melih Barsbey, Emanuele Rodol\`a, Donato Crisostomi
The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.
By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv:2606. 00384v1 Announce Type: new Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated.
By William Rudman, Abhishek Divekar, Kanishk Jain, Sebastian Joseph, Stella S. R. Offner, Matthew Lease, Kyle Mahowald, Greg Durrett, Junyi Jessy Li
arXiv:2608. 05981v1 Announce Type: new Abstract: High-resolution climate data is crucial for meteorological predictions and for informing decision support across diverse domains.
By Yichen Zhang, Yixiong Xiao, Congxi Xiao, Jingbo Zhou
arXiv:2606. 25128v1 Announce Type: cross Abstract: Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs.
By \"Umit Mert \c{C}a\u{g}lar, Alptekin Temizel
Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs. Synthetic data augmentation can extend existing datasets with realistic images, and the quality of these images is generally assessed through fidelity metrics such as FID, KID, IS, LPIPS and SSIM that measure structural or distributional similarity.
Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize high-level events and lack the visual fidelity required to assess true low-level Newtonian understanding.