On the Importance of a Clear Definition of Time Horizon for Time-to-Event Dynamic Predictions: A Systematic Review and a Concrete Illustration in Kidney Transplantation

Lucas Chabeau,1,2 Vincent Bonnemains,1 Pierre Rinder,2 Magali Giral,3,4 Solène Desmée,5 Etienne Dantan1

1Nantes Université, Univ Tours, INSERM UMR 1246, MethodS in Patients-Centered Outcomes and HEalth Research, SPHERE, Nantes, France; 2SEMEIA, Paris, Toulouse, France; 3Nantes Université, CHU Nantes, INSERM, Center for Research in Transplantation and Translational Immunology, UMR 1064, ITUN, F-44000, Nantes, France; 4Centre d’Investigation Clinique en Biothérapie, Nantes, France; 5Univ Tours, Nantes Université, INSERM UMR 1246, MethodS in Patients-centered outcomes and HEalth Research, SPHERE, Tours, France

Correspondence: Etienne Dantan, Nantes Université, Univ Tours, INSERM UMR 1246, MethodS in Patients-centered outcomes and HEalth Research, SPHERE, IRS2, 22 Boulevard Bénoni Goullin, Nantes, 44200, France, Tel +33 2 53 00 91 28, Email [email protected]

Introduction: Dynamic predictions estimate survival probability until a given horizon, conditionally on being event-free at landmark times and additional information on predictive variables available at these times. The horizon may be defined as a final time horizon or the end of a sliding horizon window.
Methods: Following PRISMA, CHARMS, and TRIPOD recommendations, a systematic review of dynamic predictions, querying Medline in May 2025 with no date restriction and including only English articles, assessed heterogeneity in time horizon reporting. Moreover, from 2,523 kidney recipients, prognostic capacities of the Dynamic predictions of Patient and kidney Graft survival have been studied considering either a final time horizon or a sliding horizon window.
Results: From 171 articles included in the systematic review, 97 articles (57%) used a sliding horizon window, 36 (21%) a final time horizon and 38 articles (22%) had no clear time horizon definition. From the kidney transplant recipients’ sample, discrimination and calibration of dynamic predictions are not comparable between the two time horizon definitions due to differing event incidence depending on the nature of the prediction window that can be either reduced or sliding. For a 5-year sliding horizon window, discrimination slightly increased with landmark times, and calibration appeared reasonable, especially at earliest landmark times. For an 11-year final time horizon, discrimination was high for earliest landmark times and increased over time, while calibration revealed underestimated predictions for earliest landmark times and overestimated for later ones.
Discussion: Our systematic review highlights the need for clearer time horizon reporting due to heterogeneous definitions. Our renal transplantation illustration revealed that the prognostic performances of Dynamic predictions of Patient and kidney Graft survival differ given the time horizon definition. More broadly, this study highlights the need for caution when interpreting prognostic performances, as it depends on whether the horizon window is sliding or reduced.

Keywords: time-to-event dynamic predictions, landmark times, horizon window, time horizon, discrimination, calibration

Introduction

In the personalized medicine era,1 time-to-event dynamic predictions are becoming more widespread. They are defined as the probability of being event-free until a defined time (time horizon) given being event-free at the time of making a prediction (landmark time) and given available predictive variables at such prediction times.2–5 Dynamic predictions take into account the valuable information consisting of the entire marker trajectory known at landmark time and have shown their importance in improving time-fixed predictions that only consider baseline variables.5–7 The dynamic predictions are updated predictions whenever additional longitudinal data becomes available during the patient follow-up. For example, Teramukai et al focused on dynamically predicting the cardiovascular endpoints of hypertensive patients until 3 years after inclusion using repeated on-treatment blood pressure measurements.8 Ben-Hassen et al proposed dynamically predicting dementia based on longitudinal neurocognitive tests over the next 5 years following the time of making prediction.9

In such dynamic prediction context, the prediction window (horizon window) corresponds to the delay between the landmark time and the time horizon. The horizon window requires precisions in reporting since it can be defined in several ways.6,10 One objective could be to predict the survival probability until a final time horizon, as, for instance, in Teramukai et al.8 An alternative objective could be to predict the survival probability until the end of a sliding horizon window6,11,12 as, for instance, in Ben-Hassen et al.9 Whatever the objective and the chosen definition of the horizon window, prognostic scores require good prognostic performances to be useful.13,14 They can be studied through global performances using time-dependent Brier Score or R2-curve,11,15 discrimination property using time-dependent AUC of ROC curves,15,16 and calibration property using time-dependent calibration plots.17

The literature on dynamic predictions is growing; however, to the best of our knowledge, a clear definition of the prediction window corresponding to the two aforementioned objectives does not yet exist, and a thorough comparison between the two approaches has not been reported. Such topics are of crucial importance when assessing the prognostic performances, as the same metrics have different interpretations depending on the objective.

The objective was to illustrate that the prognostic performances and their interpretations differ given the two time horizon definitions. We presented the mathematical framework to obtain dynamic predictions according to the two definitions (Section 2). We then conducted a systematic review of articles concerning dynamic predictions to objectively assess how the horizon window was reported in the literature (Section 3). By reanalysing kidney recipient data previously used to develop and internally validate the Dynamic predictions of Patient and kidney Graft survival (DynPG),17 we illustrated that the prognostic performances of dynamic predictions differ between time horizon definitions (Section 4). Finally, Section 5 offers a discussion and conclusions.

Methods Notations

Let us assume a learning sample of independent and identically distributed patients with the observed data . Here, corresponds to the set of longitudinal marker values measured at the corresponding times in individual . Let us consider with the true time-to-event, the censoring time for subject , and the event indicator function taking 1 if the event is not censored and 0 otherwise. We note the baseline variables. From a learning sample of patients, we can estimate a prediction model of parameters . Among those, landmarking and joint modeling of longitudinal and survival data are popular approaches.2,18,19

Dynamic Predictions Definitions

The dynamic prediction for a new patient is defined as the probability of being event-free until a time horizon given being event-free at the landmark time , given baseline variables and the longitudinal marker history available at time (ie ):

(1)

We report two different uses of this generic definition of dynamic predictions,2,3,5,10,20 that depend on the definition of time horizon and that would not have the same clinical objective.

Dynamic Predictions Given a Final Time Horizon

The objective would be to dynamically predict the survival probability for a patient until a fixed final time horizon (Figure 1). It corresponds to the probability to not suffer the event on a horizon window defined as the delay between the landmark time and the final time horizon :

Displayed are two plots illustrating the change in serum creatinine levels and patient-graft survival likelihood during the observation period.

Figure 1 Scheme of dynamic predictions of patient and kidney graft survival until a final time horizon (red line) given a landmark time Part (A) and given a landmark time Part  (B), with longitudinal measures (blue crosses) and predicted marker evolution available before the landmark time (blue line).

Using the estimated parameters of the prediction model, the dynamic predictions considering a final time horizon can be estimated as the ratio between the survival probability until the final time horizon and the survival probability until the time of making prediction :

Following this, the horizon window is progressively reduced when landmark time increases (Figure 1). We note that for a fixed , the smaller the difference between and , the bigger will be. Indeed, is bounded with:

Dynamic Predictions Given a Sliding Horizon Window

Some authors have proposed an alternative definition of the dynamic predictions.6,11,12 The target of inference of a patient is to predict the survival probability until the end of a horizon window of length :

Following this definition, the horizon window is thus sliding when updating the prediction, ie. the landmark time increases (Figure 2). Using the estimated parameters of the prediction model, the dynamic predictions given a sliding horizon window can be decomposed as the ratio between the survival probability until the end of the horizon window and the survival probability until the time of making prediction :

Two graphs showing serum creatinine values and patient-graft survival probability over follow-up time.

Figure 2 Scheme of dynamic predictions of patient and kidney graft survival for a fixed horizon window (red line) given landmark time Part (A) and given landmark time Part (B), with longitudinal measures (blue crosses) and predicted marker evolution available before the landmark time (blue line).

Accuracy Measures for Dynamic Predictions

The global prognostic performance of dynamic predictions can be assessed by measuring the prediction error,21 with the popular Brier Score metric, for instance. The time-dependent expected Brier Score is defined as follows:15,18

The Brier Score is a mean square error term. The closer the prediction is to the observation, the closer the Brier Score is to 0. When considering the final time horizon, it means that the survival probability until a final time horizon is expected to be close to the observed event indicator after the final time horizon . When considering the sliding horizon window, it means that the survival probability after an horizon window is expected to be close to the observed event indicator after this horizon window . In a dynamic context, one limit of the Brier Score is that it is sensitive to the marginal event probability that could change as the landmark time increases. Van Houwelingen et al proposed studying the relative error reduction,21 also named R2-curve:11

where is the Brier Score obtained using the model of interest and is the Brier Score of a reference model (Kaplan–Meier estimator for instance) allowing to assess the global performance of dynamic predictions whatever the evolution of the marginal event probability along the landmark times. Since the Brier Score computation would be impacted by the consideration of the final time horizon or the sliding horizon window, the R2-curve would also not be the same between the two approaches.

Discrimination is the ability of a prognostic tool to order the risk between subjects. To assess discrimination, the well-known Area Under the Receiver Operating Characteristics Curve (AUC) can be explicitly defined in a dynamic context as:15,16

When considering the final time horizon, the AUC corresponds to the probability that a subject at-risk at landmark time and who would not suffer the event before the final time horizon would have a higher survival prediction than a subject at-risk at landmark time and who would suffer the event before the final time horizon U . When considering the sliding horizon window, the AUC corresponds to the probability that a subject at-risk at landmark time and who would not suffer the event before the end of the horizon window would have a higher survival prediction than a subject at-risk at landmark time and who would suffer the event between and .

The calibration property assesses the ability of a prognostic tool to provide a prediction close to the observed outcome. Usually assessed through calibration plots, the calibration is described by comparing the predicted values within subgroups (defined from quantiles of predictions) to the observed survival probabilities (computed using the Kaplan–Meier estimator for instance). When considering the final time horizon, a good calibration means that, for any given value, we expect that among all subjects who are predicted to survive with a % probability, out of 100 will not experience the event before the final time horizon. When considering the sliding horizon window, a good calibration means that, for any given value, we expect that among all subjects who are predicted to survive with a % probability, out of 100 will not experience the event before the end of the horizon window.

Systematic Review of Dynamic Predictions Search Strategy

We conducted a systematic review following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses statement - PRISMA (Table S1).22 We searched the Medline database on 5th May 2025, with no date restriction and including only English articles. The search equation used is reported in Appendix 1.

Study Selection

As this systematic review questioned the definition and the choice of time horizon, we included all articles focusing on the development or the validation of individual dynamic prediction scoring systems as well as methodological papers concerning dynamic predictions. Two reviewers (LC and ED) independently screened references by title and abstract, and their concordance was assessed based on the proportion of observed agreements and Cohen’s kappa. Our exclusion criteria were reviews or meta-analyses, full texts not found or conference abstracts and editorials.

Data Extraction

Each article was randomly allocated to two of the five reviewers (LC – VB – PR – SD – ED). Following the CHecklist for critical Appraisal and data extraction for systematic Reviews of prediction Modelling Studies (CHARMS),23 we predefined a standardized form for data extraction. Whether it was a methodological article or an article concerning the development or the validation of dynamic predictions, for each eligible article, we collected descriptive information as the author’s name, year and journal, the data source. If we finally decided to not include it, the reason for was specified. All eligible texts were independently assessed by the two reviewers for inclusion/exclusion, and for time horizon classification when included. The inter-rater agreement between reviewers was measured based on the proportion of observed agreement and Cohen’s kappa. For each included article, the time horizon was qualified as either a final time horizon or the end of a sliding horizon window based on an explicit equation, explicit texts, and/or a clear graphic representing individual dynamic prediction, or qualified as unclear information when such explicit information was unavailable or contradictory information was present. When necessary, any discrepancies were resolved through discussion with another reviewer to reach a consensus.

We also collected the study design, the sample size, the population characteristics, the predicted outcomes, the prediction times (ie. landmark times corresponding to the times when one calculates the prediction), the time horizon (ie. the end of the prediction time window), the predictive tools and the development details (such as the methodology used, the variables of the scoring systems). The TRIPOD statement recommends to develop a prognostic score from a learning sample and to validate the prognostic performances from an independent internal and/or external validation sample to avoid reporting the prognostic performances on the learning sample only, which would lead to overestimating the performances.24,25 We thus extracted information on internal and external validations as well.

Results

The search identified 425 unique articles. Screening of titles and abstracts identified 275 papers eligible for full-text review. As detailed on the flow diagram (Figure 3), 104 articles were excluded. Details of the inter-rater reliability are provided in Table S2. Finally, we included a total of 171 articles (Appendix 2), among which 66 (39%) were methodological articles and 105 (61%) were applied articles.

PRISMA flow diagram showing study selection process from 425 articles to 171 included studies.

Figure 3 PRISMA flow diagram, selection of included studies in the systematic review.

A general overview of the included articles is presented. Dynamic predictions appear to be a relatively recent research topic in biostatistics with 169 (99%) articles published after 2010. Sixty-six (39%) of the retained articles are methodological studies that proposed new modeling approach. This systematic review confirmed the dominance of landmarking (n=69, 40%) and joint modeling for longitudinal and survival data (n=66, 39%) to develop dynamic predictions. To a lesser extent, we identified the emergence of machine learning approaches (n=33, 19%), with, for instance, random survival forest,26–28 or neural networks.29,30 All articles referring to machine learning approaches were published after 2018. This will probably increase in the future with the fast development of such modeling approaches. Among the applied articles (n=105, 61%), we also observed that 22 studies (21%) only presented a development modeling of dynamic prediction without validation of their predictive performance on an independent dataset (Table 1). This does not respect the TRIPOD recommendations about the right process to develop and validate prognostic performances. Nevertheless, 58 articles included internal validation, 8 articles included an external validation and 15 both types of validation, leading to 66% of the applied articles in agreement with the TRIPOD.

Table 1 Description of the Articles Included in the Systematic Review

According to the main objective of our systematic review, the different approaches used to define prediction horizons were examined. Among the included articles, we identified 36 (21%) articles that defined dynamic predictions using a final time horizon, 97 (57%) articles using sliding horizon windows, and 38 (22%) articles in which the definition of time horizon was unclear (Table 1). We observed that this distribution differed between methodological and applied articles. Methodological articles mainly used sliding horizon windows for dynamic predictions (n=52, 79%), while only 7 (11%) articles used a final time horizon, and 7 (11%) articles did not clearly define the approach adopted. Regarding applied articles, the definitions used were much more heterogeneous, as 45 (43%) used sliding horizon windows, 29 (28%) used a final time horizon, and 31 (30%) were unclear regarding the approach used.

Beyond the description of the time horizon considered, our systematic review allows to appraise the reporting quality on dynamic predictions. As clearly recommended in the literature,13 assessing prognostic performances is of major importance for prediction tools to propose useful scores. Discrimination capacity is well reported (n=148, 87%) with AUC as the principal performance indicator. Global performances (n=76, 44%) are less often provided, with Brier Score the most frequent indicator. Calibration (n=51, 30%) is also less reported, with calibration plot the most used approach. Details of such prognostic performances are also not homogeneous between methodological and applied articles (Table 1).

Application Context

In kidney transplantation, patients are particularly interested in their kidney graft survival, before any risk of adverse outcomes or infections.31 In this context, we developed and internally and externally validated “Dynamic predictions of Patient and kidney Graft survival” (DynPG) for kidney recipients alive with a functioning graft at 1-year post-transplantation.17,32 The main outcome was the delay from 1-year post-transplantation to patient and kidney graft failure defined as the first event between return-to-dialysis, pre-emptive retransplantation and death with a functioning graft. Such dynamic predictions can be obtained from six baseline variables: recipient age, graft rank, cardiovascular histories, pretransplantation anti-HLA class I immunization, serum creatinine at 3-months post-transplantation, occurrence of acute rejection in the first year post-transplantation, and also the complete longitudinal serum creatinine trajectory available at the time of prediction. Our aim was to provide patients and physicians with information regarding mid-term prognosis by considering a sliding horizon window of 5 years. In line with this clinical objective, the DynPG was defined as the probability of being kidney graft failure-free over the 5 years following the landmark time, for each landmark time from 1 to 6 years post-transplantation. We retained 6 years post-transplantation as the maximum landmark time because there was still a reasonable number of at-risk patients at the end of the prediction window (178 patients still at risk of patient and kidney graft failure at 11 = 6 + 5 years post-transplantation in the validation sample).

We aimed to illustrate that the prognostic performances would not be the same under the assumptions of a sliding horizon window or a final time horizon. While a sliding horizon window may be relevant to monitor the patient risk along his/her follow-up with always the same horizon window of 5 years, the clinical objective is different under the prism of a final time horizon. Regarding the final time horizon, we aim to provide long-term dynamic predictions with good confidence up to the final time horizon. Therefore, we chose the final time horizon as the maximum time considered in the sliding horizon window (ie. 11 years). Benefiting from the previously estimated joint model and the same internal validation sample, we compared the prognostic performances of the DynPG under the assumption of a 5-year sliding horizon window and under the assumption of an 11-year final time horizon.

Study Population

Data were extracted from the French multicentric observational and prospective DIVAT cohort (Données Informatisées et VAlidées en Transplantation; (www.divat.fr). The study was conducted in accordance with the Declaration of Helsinki and approved by the institutional review board of the DIVAT network. Data confidentiality was ensured based on the recommendations of the French Commission for Data Protection (Commission Nationale Informatique et Liberté, CNIL no. 914184, ClinicalTrials.gov recording NCT02900040). All participants provided written informed consent.

The inclusion criteria were adult recipients who received a first or second renal graft, transplanted between January 2000 and October 2016, from a living or heart-beating deceased donor, who were alive with a functioning graft at 1 year post-transplantation and maintained under tacrolimus and mycophenolate. The extracted DIVAT cohort data consisted of a learning set of 2,749 patients, initially used to estimate the shared-random joint model, and an internal validation set of 2,589 patients. Due to missing data for variables required to calculate the predictions, 66 patients (less than 3% of the eligible patients) were excluded from the analysis of the validation set (n = 2,523). However, the excluded patients had characteristics comparable to those of the included patients, suggesting that the missing data process was random; more details on this validation sample are reported in Fournier et al.17

Results

Under the 5-year sliding horizon window setting, the results were as previously published.17 The global prognostic performance of the DynPG appeared relatively stable along the landmark times with R2 values ranging from 14% (95% CI 7–21%) to 15% (95% CI −2% to 33%) at 1 and 6 years post-transplantation, respectively (Figure 4A). The discrimination slightly increased along the prediction times with the AUC values ranging from 0.72 (95% CI 0.67–0.78) to 0.76 (95% CI 0.68–0.85) at 1 and 6 years post-transplantation (Figure 4B). The calibration properties appeared reasonable (Figure S1). Note that, because the discrimination performances increased and the global prognostic performances remained stable, the calibration properties decreased for the late landmark times. As the landmark times increased, the calibration slope progressively deteriorated, with the estimated intercept and slope moving further away from 0 and 1, respectively (Figure S1).

Two graphs showing estimations of R squared and AUC over post-transplantation time from 1 to 6 years.

Figure 4 Prognostic capacities of the dynamic predictions obtained from the DIVAT internal validation sample (n=2,523, 66 observations deleted due to missing data concerning covariates) for landmark times varying from 1 to 6 years post-transplantation for a given 5-year horizon window; R2 evaluated global performance (A) and the AUC appraised the discrimination accuracy (B). Estimations are drawn as solid lines and the corresponding 95% CIs are drawn as dashed lines.

Under the 11 years final time horizon assumption, the global prognostic performances of the DynPG were at a higher level for the earliest landmark times, but decreased along the landmark times. We estimated R2 values of 22% (95% CI 12–32%), 15% (95% CI −2% to 31%) and −2% (95% CI −41% to 37%) at 1, 6 and 10 years post-transplantation, respectively (Figure 5A). The discrimination was high for the earliest landmark times and increased over times with AUC values of 0.76 (95% CI 0.70–0.82), 0.76 (95% CI 0.67–0.84) and 0.88 (95% CI 0.79–0.98) at 1, 6 and 10 years post-transplantation, respectively (Figure 5B). In contrast, we observed a very poor calibration plot, where the estimated intercept and slope of the calibration slope were far from 0 and 1, respectively, since the earliest landmark times (Figure S2). The predictions appeared underestimated for the earliest landmark times and overestimated for the subsequent ones.

Two graphs showing R squared and AUC estimations over 10 years post-transplantation.

Figure 5 Prognostic capacities of the dynamic predictions obtained from the DIVAT internal validation sample (n=2,523, 66 observations deleted due to missing data concerning covariates) for landmark times varying from 1 to 10 years post-transplantation for a final time horizon of 11 years post-transplantation; R2 evaluated global performance (A) and the AUC appraised the discrimination accuracy (B). Estimations are drawn as solid lines and the corresponding 95% CIs are drawn as dashed lines.

Following these two definitions, the horizon window is either reduced or sliding as the landmark time increases. Therefore, the at-risk population at landmark times are identical given these two definitions. However, the number of observed events and the number of censored subjects in each window differ between the two approaches due to different horizon windows, resulting in a number of at-risk patients at the end of the horizon window that are not comparable (Figure S3, S4). Therefore, the incidence rates in the prediction window are not the same between the two approaches. Consequently, dynamic predictions cannot be compared, and they tell different stories. For instance, under the sliding horizon window assumption, at 1 year post-transplantation, we estimated a 72% probability that the predicted survival of a patient who actually experienced a graft failure within the 5 years was lower than that of a patient who did not (Figure 4 – part B). Under the final time horizon assumption, at 1 year post-transplantation, we estimated a 76% probability that the predicted survival of a patient who actually experienced a graft failure before 11 years post-transplantation was lower than that of a patient who did not (Figure 5 – part B). In terms of calibration properties, we may reasonably accept that DynPG was sufficiently well calibrated for the earliest landmark times under the 5-year sliding horizon window assumption (Figure S1). In contrast, under the 11-year final time horizon assumption, the calibration is quite poor (Figure S2).

Discussion

In this study, we distinguished two types of time horizon – final time horizon or end of a sliding horizon window – for dynamic predictions. We conducted a systematic review that stated the heterogeneity of the used time prediction horizons in the literature about dynamic predictions. While the two definitions are similar, a specific definition is of major importance since the prognostic performances obtained are different given the nature of the prediction window. This is also well illustrated by our concrete application in kidney transplantation.

The concept of P4-medicine (predictive, preventive, personalized and participatory) is now largely developed in the literature, but still difficult to apply in clinical practice.1 Van Calster et al recently highlighted several challenges that make difficult the development, the validation and the implementation of prediction models for a clinical use33 Dynamic predictive tools can be beneficial for such a health policy and promote shared medical decision making, provided that the prognostic performances are sufficiently good. While dynamic predictions are updated predictions whenever additional information is available during the patient follow-up,6,10 their associated prognostic performances depend on the at-risk population at the time of making a prediction and can evolve as the landmark time increases. In this work, we insist on the fact that prognostic performances are also related to the prediction window. We showed that the incidence rates of the event would not be identical between the two time horizon definitions due to the nature of the window that can be either reduced or sliding, resulting in prognostic performances that are not comparable. The corresponding interpretations of prognostic performances should therefore be formulated with caution since they do not tell the same story in the two contexts.

In our concrete application to kidney transplantation, the dynamic predictions obtained with the sliding horizon window framework provide good calibration properties from 1-year until 6-year landmark times. In our opinion, these properties make the DynPG suitable for following and monitoring patient health evolution across landmark times and make it a promising option for individualized and personalized medicine.17,32 The international SONG-Tx initiative highlights the importance of core outcome domains in kidney transplantation, as graft health and mortality remain among the most important criteria in the eyes of patients.34 Therefore, the mid-term prognostic information provided by the DynPG could constitute individual levers for action. This effect could increase patient adherence to their treatment and increase patient empowerment, as patients can take an active role in their chronic disease management.35,36 This could also help patients to better manage their feelings of uncertainty about the survival of their transplant in the not-too-distant future. Furthermore, we may envisage guiding patients through the care organization based on stratified survival probabilities.

When considering the final time horizon (ie. reduced window), the comparison of prognostic performances across the landmark times do not make sense since the lengths of the prediction window are not the same. Despite the bounded character of dynamic predictions for a final time horizon, it is important to note that such predictions are not always monotonic along the landmark times because of the actualization of the longitudinal marker. The amelioration or deterioration of the prognosis between two landmark times is not necessarily due to a difference in health state characterized by the new marker measurement. This is noised by the mathematical artefact brought by the reduction of the prediction window. Since it is difficult to know to what extent the prognosis evolution is due to the reduction of the window rather than to the marker evolution, a comparison cannot be done between predictions realized at two different landmark times, contrary to the sliding horizon window framework. For this reason, the final time horizon approach does not seem suited to follow the patient health evolution through survival probabilities and cannot be envisaged to support patient personalized and individualized care. In our kidney transplantation application, the poor calibration property across the landmark times for a final time horizon of 11 years post-transplantation do not allow to consider that individual survival probabilities are correctly predicted. The more the prediction time approaches 11 years, the more we predicted patient and kidney survival probabilities close to 1, which is not confirmed by the observed event repartition. This inevitably results in poor calibration. Based on the predicted patient and kidney graft survival probabilities, such an approach cannot be used to individualized the kidney recipient taking care along the follow-up. Nevertheless, in a context where the objective is not the individualisation of patient care, but rather the stratification of the studied population, considering a final time horizon may be of interest. In kidney transplantation, the TELEGRAFT study aimed to assess interest in telemedicine depending on the strata of patient graft failure risk calculated at baseline.37 By extension, the satisfying levels of discrimination performance of the Dynamic predictions of Patient and kidney Graft survival can support stratified medicine. For example, healthcare visit schedules could be optimized by reducing monitoring or considering telemedicine consultations for low-risk patients while reinforcing follow-up for patients at higher risk.37,38 From a public health perspective, such an efficient allocation of resources could also provide economic benefits. In the context of a chronic disease such as kidney transplantation, the definition of a clinically relevant final time horizon may be difficult from a patient-centred perspective. Other clinical contexts may be of interest when a final time horizon is already known. For example, manufacturers of medical devices are required to specify a maximum duration of use. The aortic valve duration or hip replacement duration could be predicted for such final time horizons announced by suppliers39 For pregnancy outcomes, the prediction of adverse neonatal outcomes could be provided for a final time horizon of nine months.40,41

Considering the methodological differences between the two time horizon definitions, the studies’ clinical objectives should be clearly anticipated to correctly assess the prognostic performances. Indirectly, the heterogeneity that we observed in the reporting of time horizon windows in our systematic review, which was more substantial in applied articles, reveals a lack of clarity regarding the clinical objectives justifying dynamic prediction development. Van Calster clearly stated poor or ambiguous reporting as one of the important limits for the use of prediction models.33 They highly recommend to follow adequate reporting guidelines as TRIPOD,24 or the recent TRIPOD+AI for which specific items concern the AI predictive algorithms.42 Our systematic review also informed about the methodological quality of dynamic predictions. A non negligible proportion of the literature did not sufficiently respect the TRIPOD recommendation,24 with a lack of validation studies for instance. We also noted that discrimination properties were mainly reported and referred to the stratification of the studied population, but calibration metrics that are essential for personalized care were not reported in numerous studies.43 Such findings are consistent with previous literature demonstrating the lack of validation studies and the poor reporting of calibration performances.44–46

Our study has several limitations. First, we conducted a systematic review that suffered from several drawbacks. We performed it from a unique database (Medline). We did not use a formal risk-of-bias tool such as the Prediction model Risk Of Bias Assessment Tool (PROBAST), which is based on TRIPOD recommendations.47,48 Instead, in line with TRIPOD, we developed a standardized data extraction form capturing information similar to the PROBAST items. We do not believe this choice affected our conclusions, as the main objective of our systematic review was not to perform a meta-analysis, but rather to provide an overview of the types of horizon windows for dynamic prediction tools. Therefore, the main limitation of our systematic review is related to a possible publication bias issue, to which all systematic reviews are subject. Indeed, we cannot exclude the underrepresentation of external validation studies with non-confirmatory prognostic performances.49 Second, the results of our illustration depend on the chosen prediction horizons. Fournier et al previously chose a 5-year sliding horizon window to provide mid-term prognosis for patients and physicians.17 In this work, we adopted a similar choice for the sliding horizon window framework, while choosing an 11-year horizon for the final time horizon framework. We fully recognise that other choices could have been considered in accordance with other clinical interests. For example, a shorter horizon w

Comments (0)

No login
gif