In this study, we evaluated the segmentation performance of four neural networks for delineating ablation regions following thermal ablation therapies. The liver and spine datasets cover a variety of clinical and technical aspects, making them suitable benchmarks for assessing the current capabilities of deep learning models in medical image segmentation. Across both datasets, state-of-the-art DSC were achieved with nnUNet: 83.3 ± 13.2% for the spine dataset, surpassing the performance of Steffen et al. [13] with a DSC of 77.2 ± 15.6% and 82.2 ± 12.4% for the liver dataset, where no comparable MRI-based segmentation accuracy has been reported, establishing our result as a baseline. In contrast, previous studies investigating liver ablation zone delineation from smaller CT imaging datasets have reported lower segmentation accuracies, i.e., 77 ± 4.7% DSC across twelve clinical cases [16], and 76 ± 3% DSC in a preclinical swine model without tumors [18]. Although these results show lower variability, the absence of tumors in [18] simplifies the delineation task, as there is no need to differentiate between residual tumor tissue and ablation areas. Passera et al. [17] reported a higher true positive rate (TPR) of 90%, but a lower true negative rate (TNR) of 96%, indicating increased oversegmentation. In comparison, our results exhibit a more balanced performance with a TPR of 83.7% and a TNR of 99.7%, suggesting minimal oversegmentation and greater precision in necrotic tissue delineation.
Notably, the performance on our spine dataset is comparable to the IRV of 82.4 ± 5.9%, as reported in [15], which can be considered an estimate of the achievable performance ceiling for learning-based segmentation approaches. For liver MRI, no direct IRV values for ablation zone segmentation are available in the literature. However, a reference DSC of 88.8 ± 3.3% has been reported in CT imaging [16]. This value currently surpasses the performance of the proposed model, which may be attributed to several factors: the presence of outliers, the inclusion of untreated lesions that introduced ambiguity during both training and inference, and the limited availability of complementary imaging information due to the use of a single MRI sequence (ce\(\text _1\)). Incorporating additional sequences in future studies could mitigate these limitations. However, the ce\(\text _1\) MRI sequence plays a critical role in the segmentation of ablation zones and lesions, particularly in spinal and brain imaging [14, 24] most likely also transferring to liver imaging as well.
Notably, the HD95 values observed in both scenarios (4.63 mm for the spine and 7.35 mm for the liver) are of a similar magnitude to the 5 mm safety margin applied around the tumor. Consequently, larger localized boundary deviations may have a relevant impact on the resulting margin. However, because HD95 quantifies the 95th percentile of surface distance errors rather than the overall average segmentation accuracy, this metric is particularly sensitive to more pronounced local discrepancies, even when the overall segmentation quality remains high, as indicated by the high DSC values.
Limitations may arise from the use of a standardized patches, particularly in cases where multiple adjacent lesions are present within the same image. This can complicate the model’s ability to distinguish between distinct pathological zones. Therefore, the network was specifically trained to segment only the actively treated and most centrally located ablation zone. This strategy reflects clinical practice, which prioritizes evaluating the currently treated lesion rather than all lesions simultaneously. Another limitation is the use of separate networks for segmenting spinal and liver ablation zones, rather than employing a unified model. Although a combined network may appear more generalizable, its performance would likely decline because the spine and liver differ in anatomical context and tissue characteristics. Moreover, extending a unified model to cover additional ablation sites would require complete retraining. In contrast, separate networks allow for greater modularity and flexibility, enabling the integration of new anatomical regions into the tool without affecting existing models especially in cases like ours, where the dataset size is limited and each additional data point could increase the robustness.
In contrast to the organ-specific models, we didn’t separate models segmenting ablation with or without inserted needles as represented by the liver dataset. Additional analysis showed, that cases with a visible needle (n=17) achieved a higher average DSC (84.78 ± 6.90 %) compared to cases without a needle (n=11) (78.18 ± 17.64 %). However, this comparison should be interpreted with caution, as the no-needle group predominantly consists of more complex cases with multiple tumors per patient. In these patients, only the currently treated lesion contains a needle, while other lesions do not, resulting in an uneven distribution of case complexity between groups. Although the results may suggest that the needle provides additional spatial information that helps the network identify the ablation location more accurately, training separate networks for needle and no-needle cases would not be practical due to the clinical workflow of sequentially treating multiple tumors and evaluating treatment success only at the end with an MRI. Future work employing explainable AI approaches and advanced visualization techniques could provide further insights into how the network utilizes needle-related information.
A potential medical limitation is associated with ablation induced tissue shrinkage, which may affect both manual and automated delineations by reducing the apparent pathological volume of the ablated region, resulting in a systematic underestimation of the true ablation extent relative to pre-operative imaging. However, this bias may be less critical in clinical decision-making, as it is generally more relevant to ensure complete ablation coverage than to risk overestimation of the treated area.
Future work will focus on developing a tool for integration into clinical routine. Here, a critical step is the incorporation of both pre-interventional and post-interventional imaging, including image registration and initial tumor segmentation. This enables direct assessment of tumor coverage by the ablation zone, as well as 3D distance visualization to ensure a consistent 5 mm safety margin in all directions. Compared to conventional 2D slice-by-slice evaluation, this method offers greater spatial accuracy while the final decision making and manual segmentation inspection will be subject to human experts. An additional advantage lies in its potential intraoperative application: real-time assessment using a needle-based reference could help identify residual tumor tissue, allowing immediate re-ablation and potentially avoiding the need for a second intervention and LTP, particularly important in liver ablation for curative intent.
Comments (0)