This cross-sectional comparative study evaluated AI-generated PILs for five oral anticoagulants (warfarin, apixaban, dabigatran, rivaroxaban, and edoxaban) against FDA-referenced patient information materials and was reported in accordance with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guideline [22] and informed by the MedinAI guidelines for artificial intelligence studies in medicines, pharmacotherapy, and pharmaceutical services [23, 24].
PIL developmentAI model and comparator materialsThree publicly accessible LLMs were evaluated: ChatGPT (OpenAI, GPT-5.4), Gemini (Google DeepMind, 2.5 Pro), and DeepSeek (DeepSeek AI, V3). All models were accessed through their default web interfaces on 5 April 2026, with memory features and prior chat histories cleared before each session to minimise personalisation bias. FDA-approved PILs for Coumadin (warfarin), Eliquis (apixaban), Pradaxa (dabigatran), Xarelto (rivaroxaban), and Savaysa (edoxaban) served as comparator materials. Brand names were replaced with generic equivalents before evaluation to preserve blinding. The FDA-approved Medication Guides used as comparator materials are documented in Supplementary Table S8 (Supplementary Material S2).
Prompt design and formattingThe structure and content of the PILs were based on U.S. FDA patient labelling frameworks, specifically the Patient Information and Medication Guide templates [25,26,27]; the template used as the structural reference is provided in Supplementary Material S1, Table S2. All nine question domains outlined in these templates were adapted for this study, covering indications, contraindications, administration, precautions, adverse effects, storage, and general safety information (Supplementary Material S1, Table S1). The prompt was developed in accordance with established prompt engineering principles [28,29,30,31] and included clear task instructions, clinical context, alignment with FDA-style patient information, and safeguards against inappropriate or unsafe outputs. Preliminary versions were refined iteratively until all nine domains were consistently addressed without further prompting. A zero-shot approach was used, in which each LLM received a single instruction without example responses, to standardise inputs across models, minimise prompt-related variation, and ensure fair comparison between information sources while avoiding stylistic bias introduced by example-based prompting [28]. The final prompt is provided in the Supplementary Material S1, Supplementary Text S1. AI-generated and FDA-referenced PILs were then formatted using a uniform template aligned with health communication recommendations from the European Medicines Agency (EMA), Centres for Disease Control and Prevention (CDC), and U.S. Food and Drug Administration (FDA) [32,33,34]. This standardisation did not alter the underlying content of any material. All twenty evaluated patient information leaflets, comprising the AI-generated and FDA-referenced materials for each of the five anticoagulants, are provided in Supplementary Material S2.
Validation processExpert panel and blindingThree clinical pharmacists with at least five years of experience in anticoagulation management formed the validation panel. All reviewers were independent of the study design and the PIL generation process. Materials were anonymised using a computer-generated randomisation sequence and assigned unique alphanumeric identifiers, and reviewers remained blinded to both the content source and the model identity throughout the evaluation. Following a structured training session, reviewers independently rated all PILs using the PEMAT-P and DISCERN instruments. Disagreements were resolved through discussion until full consensus was reached.
Reproducibility assessmentReproducibility was assessed by examining whether each model produced consistent outputs when the same prompt was repeated at different times. Two dimensions were examined: intra-day reproducibility, referring to consistency between three outputs generated within the same day, and temporal reproducibility, referring to consistency between outputs generated at Day 1, 14, and 28. Each output was then evaluated across five domains: accuracy, clarity, completeness, internal consistency, and temporal stability, each scored using a five-point Likert scale as shown in Supplementary Material S1, Table S3. The five-domain rubric used to rate reproducibility was developed specifically for this study. It has not undergone formal psychometric validation, and the reproducibility findings should therefore be read as exploratory.
Inter-rater reliabilityInter-rater reliability was calculated prior to consensus resolution. Fleiss’ kappa [35] was used for PEMAT-P scores, with agreement interpreted using the criteria of Landis and Koch [36], and intraclass correlation coefficients (ICC [2, k]) [37] together with Cronbach's alpha [38] were used for DISCERN. For PEMAT-P, agreement ranged from substantial to almost perfect across all sources (Fleiss' kappa = 0.76–0.82), with higher agreement for actionability (kappa = 0.81–1.00) than for understandability (kappa = 0.55–0.75). For DISCERN, reliability was good to excellent (ICC [2, k] = 0.93–0.96; Cronbach's alpha = 0.93–0.96). Full reliability statistics are provided in the Supplementary Material S1, Tables S6 and S7.
Outcome measuresInformational quality and usability assessmentPEMAT-P evaluates usability through two subscales: understandability (PEMAT-U), the extent to which patients can process the information presented, and actionability (PEMAT-A), the extent to which patients can identify and carry out the actions described [14]. DISCERN evaluates informational quality, including whether materials present balanced, evidence-informed content that supports informed treatment decision making [15]. Both instruments were selected because oral anticoagulant PILs must inform patients clearly, enable safety–critical actions, acknowledge uncertainty, and support patients as active participants in their own care.
For PEMAT-P, a score of 70% or higher on each subscale was considered acceptable. For DISCERN, items 4, 5, and 7, which assess source citations, were excluded because regulatory Medication Guides are intentionally designed without references in accordance with FDA patient labelling guidance; scoring FDA materials on these items would penalise a structural regulatory convention rather than reflect a genuine difference in informational quality. This modification is consistent with previous evaluations of LLM-generated health information [16, 39, 40]. DISCERN results were reported as overall scores, with higher scores indicating better informational quality and a maximum possible score of 65. The PEMAT-P and DISCERN instruments are provided in the Supplementary Material S1, Tables S4 and S5.
Readability assessmentReadability was assessed using seven validated indices: Flesch Reading Ease [41], Flesch–Kincaid Grade Level [42], Gunning Fog Index [43], Coleman–Liau Index [44], SMOG Index [45], Automated Readability Index [46], and FORCAST [47]. These indices estimate reading difficulty from different text features. Flesch Reading Ease and Flesch–Kincaid Grade Level are based on average sentence length and average syllables per word, the former reported as a 0–100 ease score and the latter as a US grade level. The Gunning Fog Index and the SMOG Index are based on sentence length and the proportion of multisyllabic words. The Coleman–Liau Index and the Automated Readability Index are character-based, using characters per word rather than syllable counts. FORCAST is based on the proportion of single-syllable words and does not use sentence length, making it suitable for materials that are not written in continuous prose, such as the partly list-formatted leaflets evaluated here. Applying these complementary indices together provides a more robust assessment of readability than any single measure. Results were interpreted against the recommended sixth- to eighth-grade reading level for patient education materials. All readability analyses were performed using Readable.com (Readable Software Ltd., UK).
Statistical analysisPEMAT-P and DISCERN outcomes were summarised as medians with interquartile ranges, given their ordinal nature and small group sizes, while readability indices were summarised as means with standard deviations. Between-source differences for PEMAT-P, DISCERN, and readability outcomes were compared using the Kruskal–Wallis H test, with epsilon-squared reported as the effect size. Where significant differences were detected, pairwise comparisons were performed using the Mann–Whitney U test with Bonferroni correction. Reproducibility was analysed using repeated-measures ANOVA, with Greenhouse–Geisser corrections applied where sphericity was violated. Output stability was assessed from the observed mean differences across conditions and the presence or absence of statistically significant within-model variation, rather than against a formal equivalence margin. Given the small group sizes, a post-hoc sensitivity analysis was conducted to estimate the minimum detectable effect for the primary between-source comparisons (DISCERN, PEMAT-P, and readability) at five observations per group, using a significance level of 0.05 and 80% power. The analysis indicated that the study was powered to detect only large between-group effects (approximately η2 = 0.41); smaller differences may not have been detected. All analyses were performed using IBM SPSS Statistics version 31.0.2.0 and RStudio Desktop version 2026.04.0.
Ethics approvalThis study did not involve patients, clinical data, or identifiable personal information and was therefore exempt from formal ethics review in accordance with the Indian Council of Medical Research (ICMR) National Ethical Guidelines for Biomedical and Health Research Involving Human Participants (2017) and the Declaration of Helsinki. Written informed consent was obtained from all expert evaluators prior to participation.
Comments (0)