Concept-based explainable AI aims to make model reasoning human-readable, but the concepts themselves can be unreliable. This paper introduces an auditing framework that perturbs input regions, measures how concept outputs shift, and fits a surrogate model to assess whether concept explanations are trustworthy - evaluating attribution accuracy, faithfulness, and stability. Tested on retinal fundus images comparing segmentation-based versus vision-language-derived concepts, it reveals that reliability varies by concept type and pathway. Applications include validating explainable AI systems before clinical deployment, particularly in medical imaging, where trustworthy explanations are critical for physician confidence and regulatory approval.
Authors: Mohadeseh Mollapour, Koorosh Aslansefat, Zeinab Dehghani, Bhupesh Kumar Mishra, Tejal Shah, Zhibao Mian
Paper: https://arxiv.org/abs/2607.09649v1
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
