
Eye on AI Weekly Research Watch
Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
2 min•15 juli 2026
Om avsnittet
Vision-language model benchmarks have mostly used simple scenes and small human-description samples, leaving model errors on complex social scenes poorly understood. This study introduces a new dataset of images depicting complex social interactions and compares a decade of vision-language models (2017-2025) against human describers, analyzing five error types including hallucination and spatial reasoning. Results show multimodal LLMs now match top human performance and have closed the gap between simple and complex scenes. Applications include informing benchmark design, guiding development priorities for social scene understanding, and identifying remaining weaknesses (spatial dependence) for assistive and robotic vision systems.
Authors: Shravan Murlidaran, Miguel P. Eckstein
Paper: https://arxiv.org/abs/2607.09654v1
Eye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.