Following complex storylines in long-form TV dramas requires accurately identifying who is speaking each line, a task complicated by multiple characters, unclear audio, and short utterances where voice alone is unreliable. This paper introduces DramaSR-532K, a massive benchmark of 532,000 annotated dialogue lines across 900+ characters, and DramaSR-LRM, a reasoning-based model that combines audio, text, and visual cues via tool-use to attribute speech accurately. The approach notably outperforms existing methods on short utterances. Applications include automated subtitle generation, media indexing, content accessibility tools, and improved recommendation or search systems for video streaming platforms.
Authors: Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu, Jiacheng Shao, Pengfei Chen, Jiannan Ge, Kaiwen Duan, Qi Tian
Paper: https://arxiv.org/abs/2607.02504v1
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
