Automatic speech recognition models like Whisper are impressively accurate, but when they fail — or when accountability matters — we rarely know why they made a particular decision. LEAF-X introduces a principled explainability framework that uses entropy patterns in attention heads to identify which audio frames most influenced a transcription. It produces sparser, more faithful attributions than existing methods, with 32% better faithfulness scores. Practical applications include auditable transcription systems for legal or medical settings, debugging ASR failures in edge cases like accented speech or noisy environments, and building regulatory-compliant voice AI where model decisions must be traceable and explainable to non-technical stakeholders.
Authors: Ravi Ranjan, Utkarsh Grover, Xiaomin Lin, Agoritsa Polyzou
Paper: https://arxiv.org/abs/2606.14647v1
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
