Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models
Kaito Watanabe, Taisei Yamamoto, Tomoki Doi, Hitomi Yanaka
Why It Matters
What makes this one worth your time
Understanding how VLMs process spatial deictic expressions is crucial for improving their contextual reasoning and multilingual capabilities, which are essential for real-world applications.
This study benchmarks how well vision-language models handle spatial references across languages.
Summary
The paper develops a benchmark to evaluate the multilingual ability of vision-language models in using spatial deictic expressions, revealing differences in their usage compared to humans.
Key contributions
- Development of a benchmark for evaluating spatial deictic expression usage in VLMs.
- Experimental results showing differences in demonstrative selection between models and humans.
Notable insights
- The focus on spatial deictic expressions highlights the complexity of grounding language in visual contexts.
- The multilingual aspect suggests that spatial reasoning may vary significantly across cultures, impacting model training.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2607.07251v1 Announce Type: new Abstract: One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, which are defined as spatial expressions whose referent is determined by their situational context, such as ``this'' and ``that''. To handle spatial deictic expressions, VLMs must jointly reason over language and visual space, grounding context-dependent references in the image's spatial structure. In addition, selecting appropriate spatial deictic expressions across languages requires VLMs to understand the language-specific spatial distinctions encoded by these expressions. In this paper, we develop a benchmark to evaluate the multilingual ability of VLMs to use spatial deictic expressions in four languages. Our experiments using this benchmark reveal that the tested models use demonstratives in a manner different from that of humans, particularly in selecting the appropriate demonstratives based on the distance to the object.