arXiv AI By Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang

The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

Read the original on arXiv AI →

The Unwritten Benchmark introduces a novel challenge for multimodal machine learning, focusing on abstract perceptual reasoning through acousto‑kinematic word inference. Models must decode words written only by the audio of pen scratches and the video of hand movements, across three writing styles, without any visible ink trace. Evaluation shows a stark performance gap: humans achieve over 80% ordered letter accuracy, while leading models like GPT‑4o and Gemini 2.5‑Pro fail to exceed 10%, and combining modalities often degrades performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 28

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion

The paper introduces Inverted Asymmetric Fusion (IAF) to address strong-modality collapse in multimodal learning, where dominant modalities are degraded during fusion. IAF preserves the dominant modality by passing it unchanged and letting weaker modalities attend to it, while also strengthening weaker modalities via Modality-Aware Knowledge Distillation. Experiments on MultiHuSE, UR-FUNNY, and MUStARD show that IAF maintains unimodal performance and improves over the best unimodal baseline by up to 8.25%.

By Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat
arXiv Machine Learning
Sep 15

TwinICL: Diagnosing Multimodal In-Context Learning through Paired Counterfactuals

TwinICL is a procedurally generated benchmark that pairs matched text and image versions of tasks to enable controlled comparison of in‑context learning (ICL) across modalities. Experiments on six open‑weight models and 38 tasks show that multimodal ICL consistently underperforms text‑only ICL, with varying gaps by task family. Interventions targeting visual access, task framing, and reasoning can recover strong multimodal performance on a diagnostic subset, yet a modality gap remains even when explicit task instructions are provided, highlighting the dual role of demonstrations as context and evidence.

By Zihan Xue, Po-Yi Lu, Serhii Honcharenko, Zih-Ching Chen, Hsuan-Tien Lin, Nanyun Peng, I-Hung Hsu, Kuan-Hao Huang
arXiv Machine Learning
Jun 10

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning

arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.

By Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos, Paul Pu Liang