The Unwritten Benchmark: New Challenge Exposes Deep Gap Between Humans and AI in Abstract Perception
Introduction
FAQ
What is the Unwritten Benchmark?
It's a new challenge testing models' ability to infer words being written without visible ink, using only pen scratch audio and hand motion video, across three writing styles.
How do models compare to humans on this benchmark?
Humans achieve over 80% ordered letter accuracy, while models like GPT-4o and Gemini 2.5-Pro fail to surpass 10%, showing a massive gap.
What is the paradoxical fusion effect observed?
When models are given both audio and video, their performance degrades compared to using a single modality, indicating a failure to synthesize complementary cues.
Why does this matter for the MENA region?
It highlights limitations of current models in sensitive applications like education and audio analysis, urging investment in local multimodal AI research.
Source: arXiv cs.AI
AI-assisted content, human-reviewed.