Image-based diagnostic accuracy in molar incisor hypomineralisation and its differential diagnosis following standardised training: a comparison between undergraduate students and paediatric dentistry specialists

This study is reported following the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines (von Elm et al. 2008) and was approved by the Human Research Ethics Committee of CES University (project code: 1361; approval record:  264). All participants provided written informed consent.

Study design and setting.

This was a cross-sectional analytical study. It included undergraduate dental students and paediatric dentistry specialists. The study was conducted at the  Faculty of Dentistry, CES University in Medellin,  Colombia, with participant recruitment and data collection occurring from July 2025 to December 2025.

Participants.

Two groups of participants were included: individuals without previous clinical experience (undergraduate students) and individuals with clinical experience (paediatric dentistry specialists).

The group without clinical experience included second-year undergraduate dental students at  CES University enrolled in the course ‘’ Diagnosis of Hard Tissue Lesions of the Tooth during the July–December 2025 semester (n = 40). Census sampling included all students officially enrolled in the course, with no additional inclusion or exclusion criteria. At the time of the study, students had not received theoretical, preclinical, or clinical training in diagnosing DDEs, as this topic was being introduced to their curriculum for the first time. The course, which is part of the professional development area, is offered face-to-face, has a workload of 48 h, and is organised into two modules: cariology and developmental defects of enamel. The latter adopts a theoretical and practical preclinical approach focused on studying DDEs, including MIH, dental fluorosis, hypoplasia, amelogenesis imperfecta, and defects related to trauma, infection, and medications. For each condition, clinical features, aetiology, classification criteria, and differential diagnosis are covered. By the end of the course, students are expected to be able to classify the clinical features and extent of MIH and distinguish it from other DDEs and dental caries lesions.

The group with clinical experience comprised paediatric dentistry specialists with at least 1 year of active professional practice after completion of specialist training in private, public, academic, or combined settings. Purposive sampling with predefined criteria was used; participants were required to be in active clinical practice and to report caring for at least 20 paediatric patients per week. Eligible specialists were identified from the membership directory of the national/regional paediatric dentistry society (National Paediatric Society Name, Country) and the postgraduate alumni registry of  CES University. Invitations were delivered individually by email, including an explanation of the study aims and eligibility criteria, and a link to the informed-consent form; non-responders received up to 2 reminder emails. Forty-eight specialists who met the eligibility criteria were invited to participate, of whom 30 agreed to take part (response rate: 62.5%). Practice setting (private, public, academic, or combined) was recorded for each specialist; given the limited subgroup sizes, this variable was used to describe the sample rather than as a stratification factor in the diagnostic-accuracy analyses.

Training.

Two distinct instructional components are relevant to this study. First, as part of the undergraduate curriculum, students completed the developmental defects of enamel module of the course ‘’Diagnosis of Hard Tissue Lesions of the Tooth a face-to-face course with a total workload of 48 h, organised into two modules (cariology and developmental defects of enamel). The developmental defects of the enamel module followed a theoretical and practical preclinical approach, covering MIH, dental fluorosis, hypoplasia, amelogenesis imperfecta, and defects related to trauma, infection, and medication, and addressing, for each condition, its clinical features, aetiology, classification criteria, and differential diagnosis. Specialists did not take part in this undergraduate module. Second, and central to the present study, all participants in both groups attended the same standardised MIH training session one week before the photographic assessment. This session combined theory and practice and lasted 8 h: a 4-h lecture on MIH covering its definition, clinical features, differential diagnosis, and the criteria of the long-form MIH Index (Ghanim et al. 2015), followed by 4 h of calibration exercises and discussions based on Module V of the MIH Index training manual (Ghanim et al. 2017), delivered on the Genially platform to reinforce the standardised application of the diagnostic criteria. The session was led by an instructor with expertise in DDEs and more than 10 years of teaching and research experience. Because this training was applied identically to both groups and no pre-training assessment was performed, the study compares post-training diagnostic performance rather than evaluating the effect of the training itself.

Variables.

Study variables were classified as outcome, exposure, predictor, and potential effect-modifier variables.

The primary outcome variable was diagnostic accuracy in identifying MIH, defined as the proportion of correctly classified clinical images relative to a previously validated reference standard. Accuracy was operationalised as a dichotomous variable (correct/incorrect) for each case assessed and as the overall proportion of correct responses for each participant. The correct/incorrect classification at the level of each individual image was the raw response from which diagnostic accuracy was computed. Because the reference standard comprised more than two diagnostic categories (a non-binary gold standard), these image-level responses were aggregated using the Obuchowski area under the curve (AUC) method rather than conventional sensitivity/specificity, which requires a binary outcome. In this framework, the AUC summarises a participant’s ability to correctly rank or distinguish images belonging to different diagnostic categories; the dichotomous correct/incorrect scoring and the AUC are, therefore, two linked representations of the same underlying classification task rather than competing outcomes (see Study size and statistical methods).

The main exposure variable was professional training level, categorised as an undergraduate dental student (no clinical experience) or a paediatric dentistry specialist (with clinical experience).

Predictor variables included individual participant characteristics, specifically age (years), sex (male/female), and, for the specialist group, years of clinical experience after obtaining a specialist qualification. Self-reported frequency of exposure to MIH cases in clinical practice and previous familiarity with MIH diagnosis were also recorded.

Additional training in MIH diagnosis (yes/no) and self-reported confidence in identifying hard dental tissue alterations, measured on an ordinal scale, was examined as potential effect modifiers. Two related, but conceptually distinct experience indicators were recorded. “Familiarity with MIH diagnosis” captured the participant’s cumulative diagnostic history (a four-level scale ranging from having seen MIH cases without ever diagnosing them, through having diagnosed some or several cases, to having extensive diagnostic experience). “Frequency of exposure to MIH cases” captured how often MIH cases were currently encountered in routine practice (rarely vs sometimes). The former reflects accumulated diagnostic experience over a career, whereas the latter reflects current case mix and the recency of exposure; the two need not coincide, as a clinician may have diagnosed many cases historically yet now encounters them infrequently, or vice versa. We acknowledge that, with the present sample size, these constructs are only partially separable and that their subgroup analyses (Tables 2 and 4) should be interpreted with this overlap in mind.

Data sources and measurement.

Diagnostic accuracy was determined by classifying participants’ evaluations of 115 clinical photographs using the long-form MIH Index (Ghanim et al. 2015), a proposed and validated instrument (Ghanim et al. 2019). For this study, the following 11 categories of the long-form MIH Index were used (Ghanim et al. 2015): no visible enamel defect; diffuse opacities; hypoplasia; amelogenesis imperfecta; hypomineralisation defect not attributable to MIH/HSPM; white or creamy demarcated opacities; yellow or brown demarcated opacities; post-eruptive enamel breakdown; atypical restoration; atypical caries; and missing due to MIH/HSPM. Of these, ten correspond to developmental defects of enamel and one (no visible enamel defect) corresponds to sound teeth. The image set comprised 115 photographs: 100 depicting developmental defects of enamel—ten photographs for each of the 10 defect categories—together with 15 photographs of teeth with no visible enamel defect (sound teeth). Images were drawn from an institutional repository of intraoral photographs taken under standardised clinical and photographic conditions and accompanied by the corresponding clinical diagnosis. For each defect category, candidate images were screened for adequate quality and unambiguous representation of the defect, and purposively selected to span the range of severity and presentation encountered clinically (from early/mild to more advanced or atypical forms), including both prototypical and more challenging examples; these classifications were based on the expert judgement of the two reference examiners rather than on predefined quantitative criteria. The reference (gold-standard) diagnosis for every image was established by two examiners. The first was a paediatric dentist with more than 10 years of clinical and research experience in DDEs who had previously been calibrated against the long-form MIH Index using the criteria and reference images of the official MIH Index training manual (Ghanim et al. 2017). The second was an independent examiner with expertise in DDEs who classified the images, blinded to the first examiner’s scores. Intra-examiner agreement (κ = 0.89) and inter-examiner agreement (κ = 0.80) were calculated, and any discrepancies were resolved by consensus to determine the final reference standard against which all participant responses were evaluated.

Photographs were presented in random order on the Moodle platform under standardised and identical conditions for both groups, and participants did not have access to the reference standard during the assessment. For each photograph, participants assigned the diagnostic clinical category of the long-form MIH Index. The eruption and extension criteria of the Index, which require information not reliably obtainable from a single clinical photograph (such as the tooth's eruption stage and the proportion of the surface affected), were not assessed; the analysis was, therefore, based solely on the diagnostic classification of each image. Diagnostic accuracy was scored at the level of this single classification per image (correct vs incorrect relative to the reference standard) rather than as a composite of multiple Index criteria, and these image-level classifications were the input for the Obuchowski area under the curve (AUC) analysis.

The exposure variable (level of clinical experience) was derived from institutional academic records in the student group and from professional self-reports in the specialist group. Data on self-assessment of professional experience, diagnostic confidence, and additional training in MIH were collected through a structured questionnaire administered prior to the evaluation. Age, sex, and years of clinical experience were obtained via self-report.

Bias.

Methodological measures were implemented to minimise potential selection and information bias. To address selection bias, the student group was included through census sampling (the entire enrolled cohort), while the specialist group was recruited via purposive sampling with predefined eligibility criteria; in the latter group, the potential for volunteer bias is acknowledged. To reduce information (classification) bias, both groups evaluated the same set of 115 standardised clinical photographs using the same diagnostic tool (the MIH Index) and under comparable conditions. Images were randomly ordered on Moodle, with a maximum completion time of 1.5 h. All participants used the same type of computer and monitor in a dimly lit computer room to minimise variation in image visualisation. Additionally, participants did not have access to the reference standard during assessment. The reference standard was derived from a previously validated image repository and established by consensus between examiners, thereby decreasing the risk of systematic error in reference classification. Nevertheless, because only 10 images were used per defect category, the possibility of spectrum bias cannot be excluded: if the selected images predominantly represented typical or clear presentations, diagnostic accuracy may have been overestimated relative to the full clinical spectrum. To mitigate this, images were purposely selected to span different levels of severity and diagnostic difficulty, although the limited number per category remains a constraint addressed in the Limitations section.

To minimise differences caused by varying familiarity with the diagnostic tool, both groups completed the same standardised 8-h training session one week before the assessment, as detailed in the Training subsection.

Study size and statistical methods.

Participants’ demographic, professional, and experience characteristics were described using means and standard deviations for quantitative variables, and frequencies and percentages for categorical variables. The primary outcome measure throughout the study was the AUC estimated with the Obuchowski method for a non-binary gold standard. Conceptually, the AUC summarises a participant’s ability to correctly discriminate between diagnostic categories: higher values indicate greater discriminative ability, with values close to 1.0 representing near-perfect diagnostic accuracy and values near 0.5 indicating discrimination no better than chance. The analysis followed a hierarchical structure that is reported in the same order throughout the Results: first, the overall (global) diagnostic accuracy of each group across all categories; second, pairwise diagnostic discrimination, that is, the AUC for distinguishing each specific pair of diagnostic categories; and third, subgroup-stratified AUC analyses within the specialist group according to experience-related characteristics. The 70 participants assessed 115 images, yielding 8,050 ratings. This sample size provided precisions of ± 5.6% and ± 4.5% for the estimation of the AUC in undergraduate students and paediatric dentistry specialists, respectively. With respect to statistical power for comparing pairwise AUCs between students and specialists, the effective sample size varied with the number of photographs per diagnostic category; nevertheless, power ≥ 0.80 was achieved to detect differences of ≥ 0.05. These precision and power figures were derived under the assumptions of the Obuchowski approach for a non-binary gold standard (Obuchowski 2005), namely a two-sided significance level of α = 0.05, the observed group AUCs (approximately 0.89 for students and 0.95 for specialists), and the number of participants and images per category contributing to each estimate; the variance of the AUC was estimated using the Obuchowski formulation, which accounts for the multiple-reader, multiple-case structure of the data (Obuchowski 2005). The 8,050 ratings were not statistically independent because each participant evaluated multiple images, and each image was evaluated by multiple participants. This clustering at both the participant and image levels is intrinsic to the multi-reader, multi-case design and was accommodated by the Obuchowski method (Obuchowski 2005), which models the correlation between repeated ratings from the same reader and the same case when estimating AUCs and their confidence intervals, rather than treating each rating as an independent observation. The unit of analysis was, therefore, the discriminative performance across readers and cases, not the individual rating. For the pairwise comparisons, the large number of category pairs raises the risk of type I error inflation; the AUC differences reported should, therefore, be interpreted as exploratory, and we focus interpretation on the pattern of consistently lower AUCs for clinically overlapping pairs rather than on the statistical significance of any single comparison. The potential influence of this non-independence (pseudo-replication) on precision estimates and apparent statistical significance is acknowledged and revisited in the Limitations section.

Overall diagnostic accuracy and differential diagnostic accuracy were estimated using the Obuchowski area under the curve (AUC) method for a non-binary gold standard (Obuchowski 2005). Estimates are presented with 95% confidence intervals.

To assess the impact of experience on diagnostic accuracy, two analyses were conducted: (1) comparing the AUCs of professionals and undergraduate students, and (2) comparing AUCs based on self-reported professional experience. The significance of differences in AUC was evaluated using the test proposed by Obuchowski (2005). Additionally, the most difficult differential diagnoses among professionals were analysed by identifying pairs of categories with AUC values significantly lower than the overall diagnostic accuracy.

Analyses were performed in R using the nonbinROC package (Nguyen 2007).

Comments (0)

No login
gif