NeurIPS 2025 Workshop VLMs for Real-World Data, spotlight

From Scenes to Semantics: PersianCLEVR for Bilingual 3D Visual Reasoning

Kianoosh Vadaei, Melika Shirian, Arshia Hemmat, Mohammad Hassan Heydari, Ali Mamanpoosh, Afsaneh Fatemi

A bilingual Persian and English benchmark for compositional 3D visual reasoning, testing whether scene understanding survives a change of language.

A longer write-up of this work is on its way. Until then the paper and the code are the best place to look.