arXiv AI
Aug 18

The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

The Unwritten Benchmark introduces a novel challenge for multimodal machine learning, focusing on abstract perceptual reasoning through acousto‑kinematic word inference. Models must decode words written only by the audio of pen scratches and the video of hand movements, across three writing styles, without any visible ink trace. Evaluation shows a stark performance gap: humans achieve over 80% ordered letter accuracy, while leading models like GPT‑4o and Gemini 2.5‑Pro fail to exceed 10%, and combining modalities often degrades performance.

By Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang