Can generative AI support risk mapping in hospital radiopharmacy? Comparison of six generative models
30 September 2026
A. Crou, T. Tang, D. Kanan, P. Junior Onzo, S. Blondeel-Gomes, F. El KouariPharmacie, Groupe Hospitalier Grand Paris Nord-Est, GHI Le Raincy Montfermeil, France
Introduction
Risk mapping commonly relies on the GRA (Global Risk Analysis) method, a tool recommended by the Haute Autorité de Santé in France for a priori risk management. However, this process is particularly time-consuming and requires multidisciplinary expertise.
Generative artificial intelligence (AI) models could provide support for risk identification and prioritization, but their ability to produce relevant and usable analysis in pharmacy remains poorly documented. This study aimed to evaluate the performance of several AI models in performing risk mapping applied to the automated preparation process of radiopharmaceutical drug doses.
Materials & methods
A standardized prompt, including regulatory texts, was submitted to 6 generative AIs: Gemini® (Google), Vibe® (Vibe), Perplexity® (Perplexity AI), Claude Opus 4.8® (Anthropic), Copilot® (Microsoft), and ChatGPT 4.5® (OpenAI). Each AI also received an identical corpus of 7 internal procedures and the internal non-conformity register for the process under study. The mapping covered 8 steps of the dose preparation process using the Unidose® dispenser (Trasis) with a target of at least 25 failure modes. Severity and likelihood scoring were left to each AI. The responses were scored a posteriori by 3 radiopharmacists (RPh) on the following criteria: comprehensiveness of identified risks, relevance of the scenarios and their associated scoring, methodological compliance with GRA, specificity of radiopharmaceutical terminology, quality and feasibility of corrective actions, format of results, use of the provided documents and institutional data, and identification of original and relevant scenarios. Each criterion was scored from 1 (insufficient) to 4 (excellent), and the resulting scores were weighted according to the importance of each criterion to obtain a final score out of 20.
Results
Some tools generated preliminary clarification requests outside the scope of the standardized prompt, undermining protocol reproducibility. The three RPhs reached concordant rankings. Claude Opus 4.8 clearly outperformed the other models, achieving the maximum score (20/20) and providing a ready-to-use spreadsheet. The other models showed more heterogenous performance: Perplexity (16/20), ChatGPT (15/20), Copilot (11/20), Gemini (10/20), and Vibe (9/20).
Discussion & conclusion
This study highlights substantial heterogeneity among AI-generated risk mappings. Nevertheless, certain models may provide relevant assistance in their development, provided that structured prompts and suitable models are used, together with systematic validation by experts.
Keywords: Radiopharmacy, Risk mapping, Artificial intelligence