Validation of a digital, partly automated three-dimensional cast analysis for evaluation of orthodontic treatment assessment

In this study, we investigated the accuracy, reliability and efficiency of digital model analysis based on virtual models digitised with the model scanner orthoXscan (Dentaurum GmbH & Co. KG, Ispringen, Germany) and analysed with the software OnyxCeph3TM (Image Instruments GmbH, Chemnitz, Germany), in comparison to conventional manual model analysis on plaster models and a digital calliper. Our findings indicate that partly automated digital cast analysis is a timesaving and therefore highly efficient process and represents an accurate and reliable alternative to traditional manual cast analysis, which are of particular interest to orthodontists.

A comparison of the measurement of tooth width between digital and plaster models revealed that the partially automated digital method delivered by an amount of 0.023 mm resulted in increased values according to Kardach et al. [17]. Studies have shown varying results in measuring tooth width, with some finding no statistical difference but a slight decrease in digital values up to 0.023mm [18]. Regardless of the statistical significance observed in the discrepancy between analogue and digital methods in measuring tooth width, the results obtained with the Bolton analysis demonstrated no significant differences in terms of OR [18]. In contrast to our results, Lo Giudice et al. found no significant differences in OR [19]. Conversely, the differences in AR were not significant, which is not consistent with our results and those of Kim et al. [20]. The lack of discrepancies in the overall Bolton analysis can be attributed to the assumption of a consistent bias in tooth width. After accounting for this assumption, the proportional value of OR effectively compensates for the differences in single tooth width [11, 17, 18]. In some cases, especially in the measurement of crown height, the digital measurement process proved to be challenging due to the occurrence of boundary blurring during the reconstruction of attached gingiva in the segmentation process. This resulted in difficulties in accurately defining the cervical reference point. Consequently, this can result in a higher absolute mean difference of 0.241 mm regarding to the crown height compared to Keating et al.´s (0.1 mm) and Camardella et al.´s (0.17 mm) findings [21, 22]. In this case, other measuring tool as the “calliper measurement’’ could be more accurate to measure the crown height of these teeth. Furthermore, is also important to note that the measurement ranges for tooth width and height are similar, but that for crown height, there is a nearly four times higher deviation. Although the discrepancy is below the 0.3 mm threshold [5], it is not clinically relevant. In our study the arch dimensions showed measurements with a significantly decreased difference of 0.297 mm. In disagreement to the literature, increased values for arch widths and arch perimeter were found [23]. This result can be attributed to potential misinterpretations during the selection of reference points (centre of the cusp), particularly in cases where teeth exhibit attrition on the reference cusp. The estimation of the arch perimeter in the context of space analysis represents a key aspect of the diagnostic and planning processes employed in orthodontics. Arch perimeter estimation, such as intermolar and intercanine widths, displayed significant differences between plaster and digitised models. The clinical relevance of the detected results was questioned by the authors [23]. This type of measurement revealed in our study the greatest absolute difference between analogue and digital measurements. It is important to note that the results should be interpreted in the context of the large distances that were analysed. Therefore the percentage difference is the smallest of all the measurements that were conducted, and it is evident that the results are not clinically significant when compared to the literature [23]. In contrast to other studies, the measurement of overjet was found to be statistically different in our study [11, 18]. Similarly, the measurement of the overbite differed with statistical significance in other studies, amounting to 0.3 mm or 0.49 mm, respectively [11, 18]. In contrast, our determined mean difference was much smaller, with 0.236 mm for the overbite. In the aforementioned studies [11, 18], the discrepancy was deemed to be of no clinical significance despite statistical significance. This may be indicative of the similarly insignificant discrepancy observed in our study.

In conclusion, this prompts a fundamental debate about the clinical significance of statistical significance. A perusal of Table 2 reveals that all the observed differences in measurements are statistically significant. However, it is not immediately evident whether these differences are automatically clinically relevant for orthodontists. To address this fundamental question, a review of the relevant literature was conducted. It was recommended by Bishara et al. that a maximum difference of 0.2 mm should be allowed for repeated measurements carried out by a single examiner. The discrepancy in measurements for tooth width and overjet between the digital and plaster methods in our study was below the threshold levels required for repeated measurements with the same method, as previously stipulated [24, 25]. Another author has stated that for a digital model in orthodontics, 0.3 mm is set as the threshold of required accuracy [5]. Deviation for non-calculated measurements of up to 0.5 mm is considered to be clinically irrelevant in the literature [11, 26]. All the differences in the measurements presented in Table 2 are below this value, thus satisfying the requirements. When analysing the measurement results, it can be assumed that the statistically significant but at the same time clinically irrelevant measured values should not significantly influence the treatment decision. It may be necessary to make an exception for the crown height. As previously stated, there are discrepancies between the findings of different studies, which raises the question of what factors are responsible for the observed discrepancies in study results.

When analysing the Bland-Altman plot, it should be noted that both limits of agreement increase when measuring individual distances (with the exception of arch width), and even more so when cumulative measurements are made. Whether this variability is clinically acceptable cannot be easily answered. The tolerance limit of ± 0.5 mm postulated in the literature [11] suggests that the non calculated measurements of tooth and arch widths are clinically acceptable. At the same time, it should be noted that more complex measurements (arch length or girth) and cumulative measurements were not included in the above study and therefore the tolerance limit of ± 0.5 mm may not be applicable. Our results suggest a complexity-dependent proportionality. The more complex and extensive the measured structure (from individual teeth to dental arches to commutative groups of teeth), the greater the percentage deviations. This suggests a proportional relationship in which measurement uncertainty increases with the size and complexity of the measured structure. In general, it can be said that the mean of difference/distortion is low and the relatively small range of the limits of agreement in relation to the complexity of the measurements indicates a low variability of the differences.

As described above, there are inconsistencies between different investigations, which leads to the question what factors are causing differences between study results.

One of the factors that can affect the accuracy of a measurement is the level of experience of the observer. In the context of our study, the rater was experienced, familiar with the software being used, and completed all measurements in the allotted time of two months. The research group has many years of experience in digital measurement with the OnyxCeph software, due to regular internal training by the developers of the software program itself, previous digital studies and prior research results, the daily use of digital analysis of patient cases in clinical orthodontic practice and to support, continuously optimise and further develop the software tool. It can be observed that observers with more experience tend to produce more consistent readings when undertaking repeated measurements, although this may not always be the case [27]. However, the results of repeated measurements can also vary, even if performed by the same observer. This is particularly true when the interval between repeated measurements is extended [28]. Furthermore, it is important to acknowledge that the process of marking reference points in a three-dimensional model on a two-dimensional computer screen may present certain challenges [29]. This is due to the fact that points may be evaluated differently depending on the adjusted point of view, which can be particularly problematic when working with complex models. The challenge lies in identifying an identical reference point in both analogue and digital forms [6, 30].

In the present study, teeth exhibiting close approximate contact with neighbour teeth and the presence of crowding demonstrated an elevated discrepancy. Regarding the issue of point positioning, callipers are unable to attain the maximum mesiodistal diameter of teeth, especially in the presence of crowding. Moreover, impression materials are unable to accurately reproduce the space between crowded teeth in plaster casts [31]. In contrast, software solutions for digital model analysis offer a range of functions, including zooming and view rotation particularly at proximal contacts, even in the presence of crowding. Conversely, one source for bias in this context can be the poor resolution of the approximate area between teeth in the plaster model, in which the software is making a proposal for the missing information delineating the merging teeth surfaces of adjacent teeth. This approach may diverge from that of the examiner when the calliper is applied. A review of the literature reveals that there are discrepancies between analog and digital measurements, which are attributed to varying degrees of crowding [19, 32, 33]. The authors also propose that these discrepancies have clinical implications [33]. Furthermore, it would be beneficial to discuss whether the selection of traditional plaster models can be considered the gold standard and using an intraoral scanner by a trained user may avoid this bias.

In the analysis of the time required to perform the measurements, a significant reduction was observed using the digital method, which is consistent with other studies [8,9,10]. Looking at the significant results of timekeeping that have been studied for more than 10 years, the technology has improved considerably. Today, the software places the reference points automatically and accurately, and only needs to be checked and adjusted by the user, which can explain another time saving in our study. The first step in the process chain, which includes the traditional impression or direct digital intraoral scanning, was not studied here. Recent studies on chairside time shows a differentiated picture of the time efficiency of digital workflows, which is determined by complex interactions between technology maturity, user experience, and clinical application context, which reveal situation-dependent delay potentials [34,35,36]. Again, the literature is clear on the benefits and efficiency of digital intraoral scanning, which should be kept in mind [37,38,39].

The comparison of measurement methods on plaster models represents a limitation in terms of bias, as the use of intraoral scans only would allow a more accurate analysis. For users with non-specialised digital equipment, the results may be of limited value.

Comments (0)

No login
gif