Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,626 stories · RSS feed

arXiv Machine Learning
Jun 5

Equivariant Neural Belief Propagation

arXiv:2606. 06344v1 Announce Type: new Abstract: Probabilistic inference over spatially embedded variables requires beliefs that respect $SE(3)$ symmetry, yet existing equivariant networks produce only scalars and vectors -- not the rank-2 precision tensors needed for anisotropic uncertainty, and single-component messages collapse multi-modal energy landscapes to physically meaningless averages.

By Zehua Cheng, Wei Dai, Jiahao Sun
arXiv Machine Learning
Jun 5

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?")

arXiv:2606. 05497v1 Announce Type: new Abstract: Given the inherently multimodal nature of human experience, vision-language models (VLMs) hold substantial promise for modeling human cognition as it grows and develops with experience.

By Alvin Wei Ming Tan, David Cardinal, Tania Lorido-Botran, Laura Bravo-Sanchez, Sunny Yu, Michael C. Frank
arXiv Machine Learning
Jun 5

Electricity price forecasting across Norway's five bidding zones in the post-crisis era

arXiv:2604. 26634v2 Announce Type: replace Abstract: Norway's electricity market is heavily dominated by hydropower, but the 2021-2022 energy crisis and stronger integration with Continental Europe have fundamentally altered price formation, reducing the reliability of forecasting models calibrated on historical data.

By My Thi Diem Phan, Trung Tuyen Truong, Hoai Phuong Ha, Dat Thanh Nguyen