NeurIPS 2024 Datasets and Benchmarks Track

Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models

Arshia Hemmat, Adam Davies, Tom A. Lamb, Jianhao Yuan, Philip Torr, Ashkan Khakzar, Francesco Pinto

Figure from Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models

IllusionBench hides letters, faces and animals inside ordinary scenes. Human annotators read the shapes almost perfectly; leading vision-language models do not.

A longer write-up of this work is on its way. Until then the paper and the code are the best place to look.