A commercially available artificial intelligence (AI) system can improve upon human-only mammogram screening but not replace radiologist readings, suggested the findings of a new study.
Final results of the Mammography Screening with Artificial Intelligence (MASAI) trial, published in The Lancet, “demonstrated that AI-supported mammography screening outperformed standard double reading by radiologists,” senior study author Kristina Lång, MD, PhD, senior consultant at the Unilabs Mammography Unit at Skåne University Hospital in Malmö, Sweden, told Medscape Medical News.
The trial randomized 105,934 women in Sweden to either AI-supported mammography screening or the European standard of having two radiologists reading each mammogram without AI and compared outcomes of the two approaches. The study participants had digital breast tomosynthesis, also known as pseudo-three-dimensional (3D) mammography. AI-supported screening involved AI triage to single and double reading by radiologists, depending on the examination risk score, and using AI as detection support by highlighting suspicious areas.
“AI support increased cancer detection by 29%, reduced interval cancers by 12%, and did not increase false positive recalls,” Lång said.
She added that previously published interim MASAI trial results reported that AI-supported screening reduced radiologists’ screen-reading workload by 44%. “Together,” Lång added, “these findings suggest that AI can enhance the early detection of clinically relevant breast cancers and may improve outcomes for women participating in screening programs.”
First Study of AI in Mammography
The MASAI trial is the first randomized controlled trial evaluating the use of AI in mammography screening and the first to report on its effect on interval cancers, according to the study authors. Interval cancers are found after a negative screening but before the next scheduled screening.
The trial used one commercially available AI system to interpret examinations in the intervention group. The AI system, trained, validated, and tested with more than 200,000 examinations from multiple institutions in the US, Asia and Europe, analyzed mammograms for suspicious findings, and provided an overall examination risk score ranging from 1 to 10. Scores of 1-9 were assigned to a single reading by a radiologist; scores of 10 were assigned to double reading by radiologists.
During the 2-year follow-up, the study found a rate of 1.55 interval cancers per 1000 women in the AI-supported mammography group vs 1.76 in the standard care group, a 12% reduction in interval cancer diagnosis for the AI group.
“The interval cancer rate is an important measure that reflects both the effectiveness of a screening program and the performance of the screening test itself,” first author Jessie Gommers, a PhD student at Radboud University Medical Centre in Nijmegen, Netherlands, said. “Interval cancers are often more aggressive and are associated with poorer outcomes than cancers detected at screening. For this reason, the interval cancer rate is considered a central indicator of screening efficacy and is commonly used as a surrogate measure for breast cancer mortality.”
Additionally, the study found 16% fewer invasive, 21% fewer large, and 27% fewer aggressive sub-type cancers in the AI group.
In the AI-supported group, 81% of cancer cases were detected at screening vs 74% in the standard reading group. The rate of false positives was similar in both groups: 1.5% and 1.4%, respectively.
“A similar false positive rate indicates that AI-supported screening did not increase unnecessary recalls, which can cause anxiety and additional diagnostic procedures for women,” Gommers said. The recalls prompted by AI were appropriate and were “largely limited to women who were ultimately diagnosed with breast cancer.”
The findings do not obviate a role for radiologists in breast cancer screening, Lång said. “In this study, AI did not replace radiologists.” Instead, she said, it supported them by triaging exams for single or double reading and by highlighting suspicious areas.
“Final recall decisions remained entirely with the radiologist,” Lång added. “The results show that AI detection support helps radiologists avoid overlooking suspicious regions.”
The study’s randomized design in a real-world screening “strengthens the validity of the findings,” Lång said. “While replication in other screening programs will be important, the current results suggest that AI can help radiologists detect relevant cancers early without increasing unnecessary recalls.”
One limitation is that the study was conducted in southwest Sweden using a single AI system, “so additional research in diverse populations and with other AI tools is needed to strengthen the evidence for AI in mammography screening,” Lång said. Also, the trial only covered a single screening round. “Longer-term data across multiple rounds would provide further insight into sustained effects,” she said.
Digital vs 3D Mammography
The study results are not generalizable to breast cancer screening done in the US and other countries, Joann Elmore, MD, MPH, at the David Geffen School of Medicine at UCLA, told Medscape Medical News. “This MASAI trial evaluated the use of AI support tools on digital mammography examinations, but in the US most of the screening exams are now done with 3D tomography,” Elmore said.
She added that the double-reading standard in Europe is not usually done in the US. “Thus, this study needs to be repeated on 3D tomography exams and in health system that do not routinely double read all exams,” Elmore said.
“The reductions shown in the study should be interpreted cautiously until further studies can support the findings across multiple settings, multiple patient populations, and multiple AI algorithms,” said Vignesh Arasu, MD, PhD, radiologist and research scientist at the Kaiser Permanente Division of Research in Northern California.
Arasu noted the study was not powered to show superiority but rather noninferiority. “Even so, the findings suggest there may be a clinical benefit, and that AI could reduce interval cancers,” he said. “I hope we see more studies that will attempt to answer this question.”
The MAISAI results also leaves open the question of whether 3D mammography would reduce the interval cancer rate vs digital breast tomosynthesis, Arasu added.
More studies that look at the ability of AI to reduce interval cancers are needed, both Arasu and Elmore said.
“There is also a need to look at the best ways to introduce AI into radiologist’s workflows,” Arasu said. Two potential pitfalls of AI supported mammography he cautioned about are the risk for overdiagnosis, especially for low-grade or indolent lesions, and a risk for over-reliance on a new technology.
Elmore is a principal investigator of the PRISM trial (Pragmatic Randomized Trial of Artificial Intelligence for Screening Mammography), which launched last year. She said it will include about 400,000 3D tomography exams in the US interpreted with a radiologist assisted by an AI support platform or by radiologists only.
“I would emphasize that the goal of PRISM is not to replace human expertise but to understand how AI might complement it,” she said. “Our expert radiologists will continue to make the final call. AI may be a useful co-pilot, but it’s the radiologist who holds the wheel.”
This study received funding from the Swedish Cancer Society. Lång reported financial relationships with Siemens Healthineers, N23 Health, and AstraZeneca. Gommers, Elmore, and Arasu reported having no relevant financial relationships.
Richard Mark Kirkner is a medical journalist based in Philadelphia.
Admin_Adham