The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning
Researchers have introduced a new challenge called The Unwritten Benchmark to test multimodal machine learning models' ability to perform abstract perceptual reasoning. This task involves inferring words being written from audio and video of pen scratches and hand movements without any visible ink trace. Current state-of-the-art models, including GPT-4o and Gemini 2.5-Pro, struggle significantly with this task, achieving less than 10% accuracy compared to human participants w
Researchers have introduced a new challenge called The Unwritten Benchmark to test multimodal machine learning models' ability to perform abstract perceptual reasoning. This task involves inferring words being written from audio and video of pen scratches and hand movements without any visible ink trace. Current state-of-the-art models, including GPT-4o and Gemini 2.5-Pro, struggle significantly with this task, achieving less than 10% accuracy compared to human participants who achieve over 80%. The study also found a paradoxical fusion effect where providing both modalities often degrades performance rather than improving it.
---
Why it matters: This matters because it highlights significant limitations in multimodal machine learning models' ability to perform abstract perceptual reasoning, which is essential for many real-world applications such as human-computer interaction and cognitive assistance. The findings suggest that current models need to improve their cross-modal causal reasoning and understanding of micro-kinematics.
Source: https://arxiv.org/abs/2608.14558
This article was originally published at: https://arxiv.org/abs/2608.14558