MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
Youngwan Lee, Soojin Jang, Yoorhim Cho, et al.
This paper introduces MultihopSpatial, a new benchmark to test how well AI vision-language models can understand complex spatial relationships in images, such as 'the object to the left of the object above the red box.' The benchmark includes a new evaluation metric (Acc@50IoU) that checks both whether the model answers correctly and whether it can precisely pinpoint the location in the image, which is crucial for robots performing real-world tasks. Testing 37 state-of-the-art models reveals that spatial reasoning remains challenging, but training models on their new dataset improves both the models' spatial understanding and their ability to perform physical manipulation tasks.
vision-language modelsspatial reasoningbenchmarkvisual grounding