TY - JOUR
T1 - Leveraging large language models for heuristic usability assessment of medical software
T2 - Insights with the Radiation Planning Assistant
AU - Court, Laurence E.
AU - Smit, Jacobus
AU - Strauss, Lourens
AU - Shaw, William
AU - Marais, Andrea
AU - Trauernicht, Christoph
AU - Joubert, Nanette
AU - Smith, Elaine
AU - Badre, Shona
AU - Lazarus, Graeme L.
AU - Khotle, Thekiso
AU - Netherton, Lauren
AU - van Heerden, Wanda
AU - Cardenas, Carlos
AU - Serban, Monica
AU - Seuntjens, Jan
AU - Chung, Christine V.
AU - Govyadinov, Pavel
AU - Khan, Meena
AU - Nair, Saurabh
AU - Netherton, Tucker
AU - Zhang, Lifei
N1 - Publisher Copyright:
© 2026 The Author(s). Journal of Applied Clinical Medical Physics published by Wiley Periodicals LLC on behalf of American Association of Physicists in Medicine.
PY - 2026/2
Y1 - 2026/2
N2 - Background: Usability engineering is essential for ensuring the safety and effectiveness of medical software, as design-related issues are a leading cause of use errors in clinical settings. Heuristic evaluation provides a practical approach to identifying usability problems, but its outcomes depend heavily on expert interpretation. Large Language Models (LLMs), such as ChatGPT, offer a potential means to augment heuristic evaluation by generating structured, context-aware usability feedback. This study explored the use of ChatGPT to support heuristic assessment of the Radiation Planning Assistant (RPA), a web-based radiotherapy planning tool designed to support clinical teams in low- and middle-income countries. Methods: ChatGPT was provided with the RPA user and technical guides, training videos for each functional dashboard, and Zhang et al.’s 14 usability heuristics. The model was instructed to score each dashboard according to these heuristics, using Zhang's 0–4 severity scale, and to propose concrete interface improvements. The resulting feedback was reviewed and scored independently by the RPA developer team and by 13 users during a dedicated User Meeting. Comparative analysis was performed between ChatGPT, developer, and user ratings. Results: ChatGPT identified 26 potential usability issues across six heuristic domains. The developer team considered nine of these actionable, though all were classified as minor (severity ≤ 2). User ratings showed wide variability, with nine suggestions achieving mean scores ≥ 1.5. Qualitative agreement between users and developers was limited, underscoring the importance of diverse perspectives in heuristic evaluation. Three suggestions—enhanced upload logs, reversible actions (“reopen request”), and stronger error prevention—were rated as potentially high priority by a minority of users. ChatGPT's ratings were consistent across dashboards. Conclusions: While ChatGPT did not reveal any critical usability failures, its heuristic assessment proved valuable in prompting discussion, identifying minor refinements, and enriching both developer and user engagement with the RPA's interface design. This study demonstrates that LLMs can serve as an effective, low-cost complement to conventional heuristic evaluation, supporting early-stage usability review and stakeholder dialogue in the development of medical software.
AB - Background: Usability engineering is essential for ensuring the safety and effectiveness of medical software, as design-related issues are a leading cause of use errors in clinical settings. Heuristic evaluation provides a practical approach to identifying usability problems, but its outcomes depend heavily on expert interpretation. Large Language Models (LLMs), such as ChatGPT, offer a potential means to augment heuristic evaluation by generating structured, context-aware usability feedback. This study explored the use of ChatGPT to support heuristic assessment of the Radiation Planning Assistant (RPA), a web-based radiotherapy planning tool designed to support clinical teams in low- and middle-income countries. Methods: ChatGPT was provided with the RPA user and technical guides, training videos for each functional dashboard, and Zhang et al.’s 14 usability heuristics. The model was instructed to score each dashboard according to these heuristics, using Zhang's 0–4 severity scale, and to propose concrete interface improvements. The resulting feedback was reviewed and scored independently by the RPA developer team and by 13 users during a dedicated User Meeting. Comparative analysis was performed between ChatGPT, developer, and user ratings. Results: ChatGPT identified 26 potential usability issues across six heuristic domains. The developer team considered nine of these actionable, though all were classified as minor (severity ≤ 2). User ratings showed wide variability, with nine suggestions achieving mean scores ≥ 1.5. Qualitative agreement between users and developers was limited, underscoring the importance of diverse perspectives in heuristic evaluation. Three suggestions—enhanced upload logs, reversible actions (“reopen request”), and stronger error prevention—were rated as potentially high priority by a minority of users. ChatGPT's ratings were consistent across dashboards. Conclusions: While ChatGPT did not reveal any critical usability failures, its heuristic assessment proved valuable in prompting discussion, identifying minor refinements, and enriching both developer and user engagement with the RPA's interface design. This study demonstrates that LLMs can serve as an effective, low-cost complement to conventional heuristic evaluation, supporting early-stage usability review and stakeholder dialogue in the development of medical software.
KW - Heuristic evaluation
KW - Large language models
KW - Usability
KW - User interface design
UR - https://www.scopus.com/pages/publications/105030492803
UR - https://www.scopus.com/pages/publications/105030492803#tab=citedBy
U2 - 10.1002/acm2.70495
DO - 10.1002/acm2.70495
M3 - Article
C2 - 41708070
AN - SCOPUS:105030492803
SN - 1526-9914
VL - 27
JO - Journal of applied clinical medical physics
JF - Journal of applied clinical medical physics
IS - 2
M1 - e70495
ER -