Vision-language model benchmarks have mostly used simple scenes and small human-description samples, leaving model errors on complex social scenes poorly understood. This study introduces a new dataset of images depicting complex social interactions and compares a decade of vision-language models (2017-2025) against human describers, analyzing five error types including hallucination and spatial reasoning. Results show multimodal LLMs now match top human performance and have closed the gap between simple and complex scenes. Applications include informing benchmark design, guiding development priorities for social scene understanding, and identifying remaining weaknesses (spatial dependence) for assistive and robotic vision systems.
Authors: Shravan Murlidaran, Miguel P. Eckstein
Paper: https://arxiv.org/abs/2607.09654v1
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
