In this study, we sought to evaluate the repeatability of SUV, ADC, and SUV/ADC via voxel-wise analysis, compared to conventional whole-tumor summary metrics, and found that whole-tumor summary metrics generally had better repeatability compared to their voxel-wise equivalents (though especially for ADC). Furthermore, the type of deformation target (i.e., PET1 vs. ADC1) played a significant role in the repeatability of voxel-wise metrics. In light of their observed lower test-retest repeatability compared to whole-tumor summary metrics, voxel-wise metrics appear to be susceptible to a variety of sources of variation beyond those affecting whole-tumor summary metrics, potentially limiting their use in clinical settings.
Our findings should be evaluated in the context of several prior studies assessing the repeatability of SUV and ADC on [18F]FDG-PET/CT and/or [18F]FDG-PET/MRI. For example, in a meta-analysis of [18F]FDG-PET/CT test-retest repeatability studies performed in patients with various solid tumor types, Lodge reported a mean within-subject coefficient of variation of 10.96% (equating to an RC-rel of 30.4%) for SUVmax [20]. In a test-retest repeatability study of [18F]FDG-PET/MRI among patients with pelvic tumors, Fraum et al. found even better repeatability of whole-tumor summary statistics, with an RC-rel of 18.3% for SUVmax and 9.6% for ADCmedian [3]. In our whole-tumor summary metric analysis, we observed RC-rel values (with ranges reflecting different deformation targets) of 29.7–36.0% for SUVmean and 8.5–9.7% for ADCmean. Our voxel-wise repeatability was worse, with median RC-rel values (with ranges again reflecting different deformation targets) of 44.4–50.2%, 30.9–36.6%, and 56.7–58.2% for SUV, ADC, and SUV/ADC, respectively.
Notably, the apparent superior repeatability of ADC (voxel-wise) and ADCmean (whole-tumor) seen for RC-rel was seemingly discordant with the ICCs, which instead indicated more robust results for SUV and SUV/ADC compared with ADC in the voxel-wise analysis and similar results for ADCmean and SUVmean in the whole-tumor summary metric analysis. This finding, which reflects the nuances of RC versus ICC described in the Statistical Analysis section, is likely attributable to differences in the dynamic range of ADC (values almost never below 500 × 10− 6 mm2/s in tumors) compared with SUV (values can be near-zero in some tumor areas). Inspection of the RC-rel equation reveals that the high ‘minimum’ for ADC values effectively biases ADC-based metrics to higher RC-rel compared with SUV-based metrics, as the test-retest mean will generally be quite high relative to typically observed test-retest differences. The lower ICCs for ADC compared with SUV could potentially be due to susceptibility effects from adjacent structures (e.g., rectal gas), reducing the quantitative accuracy of our single-shot DWI sequence. Although we purposefully acquired DWI twice in each session, selecting the better of two image sets for analysis, persistent susceptibility effects could have nonetheless diminished the voxel-wise repeatability of ADC. Newer PET/MRI and MRI scanners employing multi-shot DWI sequences would likely produce ADC results with higher voxel-wise repeatability.
An unexpected finding during our study was the significantly better voxel-wise repeatability (in terms of RC-rel) for SUV with the ADC1 deformation target and, conversely, for ADC with the PET1 deformation target. For instance, one might imagine that SUV repeatability would be better when PET2 is deformed to PET1 (i.e., only one deformation step) than when both PET1 and PET2 are deformed to a third target, such as ADC1 (i.e., two deformation steps) One possible explanation for this finding is that the deformation process may mitigate the effects of outlier voxels that tend to limit repeatability, effectively smoothing the image and enhancing repeatability. In the example above, when PET2 is deformed to PET1, voxels with outlier SUVs in the PET1 dataset are not averaged with other voxels, as the PET1 data are not resampled to fit the contours and spatial resolution of a separate deformation target. Importantly, the better voxel-wise repeatability for SUV with the ADC1 deformation target might be related to the lower spatial resolution (larger voxel size) of the ADC reconstructions compared with the SUV reconstructions. Arguing against this possibility though is our observation that ADC had higher RC-rel when using the higher spatial resolution PET1 deformation target than the lower spatial resolution ADC1 deformation target. In our analysis of voxel-wise standard deviations for session 1 SUV and ADC values, we found that deformable registration decreased image noise, even when that process involved a change from larger to smaller voxels (which should generally increase image noise). These findings support our supposition that the transformation process effectively smooths the voxel-wise data, reducing the effects of outliers and potentially improving repeatability. Future studies could explore these possibilities via resampling the PET1 and ADC1 targets to the same spatial resolution before deformation or performing deformation to a third “neutral” target distinct from the PET and ADC images.
To our knowledge, our study is the first to assess the voxel-wise repeatability of SUV and ADC and provides an important complement to existing PET/MRI-based biomarker data and voxel-wise analysis studies. For example, Chenevert et al. reported that voxel-wise analysis of ADC maps from pre-treatment and mid-treatment brain MRIs predicted response to therapy among patients with high-grade glioma treated with chemoradiation [21]. In that study, significant treatment differences in ADC were defined by stratifying voxel-wise changes according to outcomes data rather than by a test-retest assessment of voxel-wise variability at the pre-treatment time point, as in our study. Applying thresholds based on intrinsic statistical variations rather than outcomes might increase the sensitivity of voxel-wise analysis for more subtle therapeutic responses. Several studies have evaluated the ability of quantitative metrics derived from PET/MRI to predict tumor features in cervical cancer [15, 16]. For example, Surov et al. reported that the SUVmax/ADCmean ratio correlated with the tumor proliferation index Ki-67, a marker of tumor aggressiveness [16]. The ability of PET/MRI to predict such features might be enhanced by voxel-wise analysis, especially in heterogeneous tumors.
Our study had some important limitations. First, the cohort size was relatively small with only ten subjects, precluding analysis of cervical cancer histologic subtypes and resulting in greater potential for noise/outlier-related errors. The analysis was conducted solely on cervical cancer; as such, it is unclear if repeatability varies across tumor types. All imaging was performed at a single center on a single PET/MRI scanner and using a single deformation approach; other PET/MRI scanner models or deformation software programs might produce different repeatability results. Our PET/MRI acquisitions were substantially longer than the typical 60-minute uptake time; the repeatability of voxel-wise metrics could be different (either better or worse) at standard uptake times. Lastly, because our patients did not undergo post-therapy PET/MRI, we were unable to determine if there was a correlation between treatment-related changes in voxel-wise metrics and outcomes. As a result, our study does not offer any thresholds for defining a clinically significant voxel-wise change.
In conclusion, voxel-wise SUV and ADC metrics derived from [18F]FDG-PET/MRI in cervical cancer patients had reasonable repeatability. Although whole-tumor metrics appear to exhibit less intrinsic variability between imaging sessions, voxel-wise analysis holds promise for various clinical applications. By capturing nuanced details and characteristics of tumor subregions that are neglected by whole-tumor summary statistics, voxel-wise analysis has the potential to be used in clinical settings to improve understanding of tumor heterogeneity (at baseline and during treatment), to craft personalized treatment plans, and, ultimately, to improve patient outcomes. Through future PET/MRI-based voxel-wise studies targeting other cancer types, increasing study cohort numbers, and correlating metrics with clinical outcomes, a robust understanding of how to incorporate voxel-wise disease assessments into clinical oncology may be achieved. Importantly, the results of our study should prompt future investigations to consider using a “neutral” deformation target for both SUV and ADC data to mitigate outliers and improve repeatability.
Comments (0)