NeurIPS 2025 Workshop VLMs for Real-World Data, spotlight
From Scenes to Semantics: PersianCLEVR for Bilingual 3D Visual Reasoning
Kianoosh Vadaei, Melika Shirian, Arshia Hemmat, Mohammad Hassan Heydari, Ali Mamanpoosh, Afsaneh Fatemi
A bilingual Persian and English benchmark for compositional 3D visual reasoning, testing whether scene understanding survives a change of language.
A longer write-up of this work is on its way. Until then the paper and the code are the best place to look.