Breast Cancer AI Diagnostic Tool Falls Short of Radiologists' Expectations
A survey of 215 members of the Society of Breast Imaging has revealed a significant gap between expectations for FDA-approved breast cancer detection AI tools and their actual effectiveness. Only 35% of physicians reported experiencing a reduction in recall rates after implementing the tools, falling well short of the 59% who had expected improvements before adoption. This discrepancy was consistently observed across all evaluation metrics measured in the survey.

FDA-approved breast cancer detection AI tools are falling short of expectations in clinical practice, according to a survey by a professional organization. A survey of 215 members of the Society of Breast Imaging found that approximately half of respondents have already integrated these AI tools into their actual clinical practice. While adoption itself is progressing, the benefits achieved are significantly below pre-implementation expectations.
One of the critical metrics in the context of medical AI is the "recall rate." This refers to the proportion of patients deemed to require further evaluation based on mammography interpretation. High recall rates increase unnecessary testing and patient burden, so AI has been expected to reduce this metric by lowering false positives. For radiologists, improving recall rates has been one of the primary goals of implementing AI tools.
However, the survey results reveal a numerical disparity between expectations and reality. Only 35% of physicians reported actually experiencing a reduction in recall rates after implementation. Meanwhile, 59% of physicians had expected this improvement before adoption, creating a gap of nearly 24 percentage points between expectation and reality. Moreover, the report indicates that this discrepancy is not limited to recall rate alone—similar patterns were confirmed across all evaluation metrics measured in the survey.
The implications of this finding become more significant when considering the characteristics of breast cancer diagnosis. Breast cancer screening operates in a delicate balance between the risk of missing cases that could be life-threatening and the risk of over-diagnosis and unnecessary testing, which causes significant physical and psychological stress to patients. The global trend of integrating AI into this domain represents part of this movement, with FDA-approved tools now beginning to be used in clinical settings.
However, it is important to note that this survey does not demonstrate that AI is "unhelpful." Rather, it visualizes the clinical perception that "expected benefits are not being realized." From this survey data alone, it is unclear whether the issue lies with actual tool performance or with overestimated expectations. The possibility that expectations for AI tools were set too high cannot be ruled out.
AI implementation in healthcare is influenced not only by diagnostic accuracy but also by integration with clinical workflows and the human dimension of how physicians interpret and apply AI recommendations. This survey serves as a case study demonstrating that even FDA-approved tools may present different real-world utility than expected, with potential implications for future discussions on the evaluation and implementation of medical AI.
A key point for future observation is how these voices from the field will be reflected in AI tool development and approval criteria. If a gap exists between the evaluation metrics used in approval reviews and effectiveness indicators in actual clinical practice, efforts to close that gap may be needed. As breast cancer detection AI continues to proliferate, awareness appears to be growing among healthcare institutions and policymakers that "FDA approval" and "functioning as expected" are not synonymous.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.