arXiv Machine Learning

The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench

arXiv:2605. 24782v2 Announce Type: replace Abstract: While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than underlying structural invariants, making even perception-based out-of-distribution accuracy a poor proxy for scientific utility.

arXiv Computer Vision
2d ago

HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation

HiPhy introduces a hierarchical reinforcement learning framework for video generation that enforces physical laws at both local and global levels. It addresses the challenge of multi-principle interactions—such as buoyancy and fluid dynamics occurring simultaneously—by ensuring each principle’s temporal dynamics and the overall scene’s coherence. The authors also provide a 50K-prompt dataset and the MultiPhyBench benchmark, demonstrating that HiPhy outperforms existing methods, especially in scenes with multiple concurrent physical principles.

By Tahira Kazimi, Shubhankar Borse, Munawar Hayat, Fatih Porikli, Pinar Yanardag
arXiv AI
Jul 7

Multi-Way Representation Alignment

arXiv:2602. 06205v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces.

By Akshit Achara, Tatiana Gaintseva, Mateo Mahaut, Pritish Chakraborty, Viktor Stenby Johansson, Melih Barsbey, Emanuele Rodol\`a, Donato Crisostomi
arXiv AI
6d ago

Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.

By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
Hugging Face Trending Papers
Jun 23

Benchmarking the Alignment of Data-Quality Metrics, Human Judgment and Land-Cover Segmentation Performance for Earth Observation

Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs. Synthetic data augmentation can extend existing datasets with realistic images, and the quality of these images is generally assessed through fidelity metrics such as FID, KID, IS, LPIPS and SSIM that measure structural or distributional similarity.