 서울아산병원 융합의학과 김남국 교수와 남유진 연구원_삼성창원병원 영상의학과 홍파 교수.jpg)
▲ (From left) Professor Namkug Kim and Researcher Yoojin Nam of the Department of Convergence Medicine at Asan Medical Center, together with Professor Pa Hong of the Department of Radiology at Samsung Changwon Hospital.
Medical education requires a vast body of knowledge and high quality educational content, but even minor errors can lead to the dissemination of incorrect medical information, making careful review essential.
Recent advances in generative AI have made it possible to produce medical educational materials rapidly. However, a recent study found errors even in specialist training materials that had passed an initial automated AI review, highlighting the need for final expert oversight.
A research team led by Professor Namkug Kim and Researcher Yoojin Nam of the Department of Convergence Medicine at Asan Medical Center, together with Professor Pa Hong of the Department of Radiology at Samsung Changwon Hospital, developed a six step generative AI based system for producing educational content. Using the system, the team created 6,000 flashcards and 833 infographics to support preparation for the radiology board examination.
The research team considered educational materials that had passed its internal quality verification process to be safe for use and set a strict acceptable error threshold of no more than 0.3%, based on error standards used in other medical fields. In other words, fewer than three major errors were permitted per 1,000 educational items.
To evaluate the materials, 20 medical professionals, including 9 radiology residents and 11 board certified radiologists, directly reviewed 1,284 flashcards for factual accuracy, educational appropriateness, and potential errors. The results revealed an error rate of approximately 1.0%, even among materials that had passed the initial AI based verification process, exceeding the safety threshold established by the research team.
This study demonstrates that generative AI has the potential to produce large volumes of specialized medical educational content. At the same time, it showed that rather than fully automating both the creation and distribution of educational materials, a human in the loop system, in which experts perform final verification, is essential for the safe use of AI in medicine.
The study also found that when AI generated feedback was provided alongside the educational materials, reviewers tended to identify more errors and evaluate educational quality more rigorously.
The research team believes this was because reviewers critically reassessed AI generated feedback when it conflicted with their own expertise, rather than simply accepting the AI’s judgment, thereby improving their ability to detect errors. The findings suggest that, with generative AI at its current level of development, a system requiring final expert verification is essential, rather than a fully automated approach to producing medical educational content.
Professor Namkug Kim of the Department of Convergence Medicine at Asan Medical Center said, “Through this large scale study, we demonstrated that expert verification is essential for the safe use of increasingly prevalent generative AI in real world medical education.”
He added, “We need to continue research to systematically manage errors missed by automated verification and establish clear safety standards, thereby creating a generative AI environment that can be used with confidence in medical education.”
Researcher Yoojin Nam of the Department of Convergence Medicine at Asan Medical Center, the first author of the study, said, “This study demonstrated the potential of generative AI to produce large volumes of medical imaging educational materials, while also showing the importance of systematically analyzing and filtering where and what types of errors arise, rather than using AI generated learning materials without further review.”
The findings were recently published in the international journal npj Digital Medicine (impact factor: 18.0).