Longitudinal Non-interventional Changes of the FORTA Score are Associated with Changes of Cognitive and Physical Function Tests in Community-Dwelling Older People

2.1 Study Design and Study Population

This work was based on the Study on Ageing, Cognition and Dementia (AgeCoDe), a multi-center (Bonn, Düsseldorf, Hamburg, Leipzig, Mannheim and Munich), population-based, longitudinal cohort study in primary care patients commencing in 2003/2004 [6].

Primary care patients aged 75 years and above without the diagnosis of dementia at baseline were recruited through general practitioners’ (GP) offices. Follow-up (FU) assessments were performed at 1.5-year intervals on average [6, 18] until FU 7; from FU 7–8 the interval was only 10 months. The 8th follow-up finished in 2016 [6, 19].

Data from FU 7 onwards were collected from AgeQualiDe (a study on needs, health service use, costs and health-related quality of life in a large sample of oldest-old primary care patients [85+]). AgeQualiDe represents an extension of the AgeCoDe study, continuing largely comparable data collection procedures, diagnostic criteria and core outcome assessments to enable longitudinal analyses across study waves. However, some modifications were introduced, including shorter follow-up intervals as indicated above and additional measures (e.g., quality of life and health care utilization), meaning the methodologies were expanded.

Attrition throughout the study up to FU 6 is described elsewhere [18]. At FU 3, there were 1161 participants, 69 died and 145 dropped out, at FU 4 these numbers were 977, 76, 82; at FU 5 they were 822, 74, 66; at FU 6 they were 652, 62, 68. Similarly, for the following FUs in AgeQualiDe it is stated that attrition was mainly due to death and dropout [20, 21]. In detail, at FU 7, 136 participants died and 46 refused participation; at FU 8, 78 died and 17 refused; and at FU 9, 92 died and 18 refused [20, 21].

The main outcomes of this retrospective analysis of longitudinal cohort data were to find multivariate and univariate associations between the FORTA score and clinical variables, and statistical comparisons of these variables for patients showing an increase and those showing a decrease of the FORTA score; thus, the cutoff for this dichotomization of the pre/post FORTA score difference (delta FORTA score) was set to zero.

2.2 Data Collection

Relevant data collected at the 6th through 8th FU (FU 6–8) “included drug use (ATC codes), age, gender, GP diagnoses and blood pressure. Data were electronically entered into the database” [6]. At FU 6–8 of the AgeCoDe/AgeQualiDe study, patients (by patient ID) were included if data sufficient for comparison were present for FU 7–6/8–7/8–6 [6].

2.3 FORTA (Fit fOR The Aged) Diagnoses

Data were not available for all FORTA diagnoses. The alignments of non-FORTA diagnoses have been detailed before [16]; in brief, any of the following diagnoses were assumed to reflect the FORTA diagnosis ‘gastrointestinal disease’: gastritis, reflux gastritis, reflux, esophageal carcinoma and gastrointestinal bleeding. Similarly, stroke, cerebral infarction, stenosis of the afferent cerebral arteries and transient ischemic attacks were attributed to the FORTA diagnosis ‘stroke’ [6].

Further details of these studies are provided elsewhere [18, 19, 22,23,24,25,26,27,28,29,30] and in the earlier papers on the FORTA analysis of those data [6, 15].

2.4 Ethics

Approval from the ethics committees of all the participating centers [23] had been obtained for the AgeCoDe and AgeQualiDe studies, which were conducted in compliance with the Code of Ethics of the World Medical Association [6, 31]. Furthermore, ethics approvals that were previously received (reference number: 2007-253E-MA) and recently renewed by the ethics committee in Mannheim, University of Heidelberg (reference number: TEMP558252-AF 11) cover the work in this paper [6].

2.5 Determination of the FORTA Score

In brief, “the FORTA-list assigns 4 FORTA classes to drugs that are defined as follows” [6]:

“Class A (A-bsolutely) = indispensable drug, clear-cut benefit in terms of efficacy/safety ratio proven in elderly patients for a given indication

Class B (B-eneficial) = drugs with proven or obvious efficacy in the elderly, but limited extent of effect or safety concerns

Class C (C-areful) = drugs with questionable efficacy/safety profiles in the elderly, to be avoided or omitted in the presence of too many drugs, lack of benefits or emerging side effects; review/find alternatives

Class D (D-on’t) = avoid in the elderly, omit first, review/find alternatives” [6].

“The related FORTA-score is the sum of medication errors classified as over- and/or under-treatment errors in an individual patient as checked against these labels. An error was counted if an indication was not appropriately treated though beneficial options (FORTA A or B exist – under-treatment) or if a prescription was suboptimal regarding the FORTA categories (e.g., FORTA C though A or B drugs exist) or not indicated (over-treatment)” [6].

“Proton-pump-inhibitors are often not indicated and would trigger an overtreatment error; oral anticoagulation is strictly indicated in atrial fibrillation, and the absence of a positively labeled oral anticoagulant (e.g., apixaban) would be considered as undertreatment error. To apply FORTA, a demand analysis for drug treatment is required and the determination of the FORTA score relies on it.” [6].

A more detailed description of the FORTA score can be found elsewhere [14, 32] and in the first paper on FORTA in AgeCode [6].

For FU 6, the FORTA scores have already been published by Pazan et al. [6, 16]. The score was determined at FU 7 and 8 in the same way.

2.6 Statistical Analysis and Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) Statement

A preceding sample size analysis [16] showed that 310 patients from this cohort would be sufficient to detect a minimal clinically important difference (MCID, one score point) assuming a standard deviation of 2.7 (power = 0.9). The MCID for FORTA changes was oriented on the difference of FORTA scores in the interventional VALFORTA trial [15], which was 1.7, and resulted in the clinical benefits; a clear dose–response relationship could not be established from these data, thus, the assumption of 1 as MCID is consensual, but within the frame of clinical success.

Associations between changes in FORTA score (delta FORTA) and corresponding changes in ADL, IADL and MMSE were analyzed for three follow-up pairings (FU7 vs FU6, FU8 vs FU7 and FU8 vs FU6) using Spearman’s correlation analysis, with all variables treated as continuous variables. Besides, to minimize the risk of type I error, the Bonferroni correction was applied reflecting the total number of tests. Changes in FORTA score (delta FORTA) were analyzed as a continuous variable in order to capture the full range and magnitude of change and to avoid loss of information associated with dichotomization. Associations between FORTA score changes and clinical outcomes (MMSE, ADL and IADL) were assessed using multivariable linear regression models, adjusting for potential confounders including age, sex, number of medications and number of comorbidities. Clinically relevant decline was defined a priori as a decrease of ≥ 1 point in MMSE and IADL, and ≥ 5 points in ADL, reflecting differences in scale granularity and established clinical interpretability. Receiver operating characteristic (ROC) curve analyses were conducted to evaluate the ability of continuous FORTA score changes to discriminate between participants with and without clinically relevant decline. The area under the curve (AUC) was used as a measure of discriminative performance. Optimal cut-offs were estimated using the Youden index and are reported for exploratory purposes only, given the limited discriminatory ability observed.

To facilitate comparability with previous approaches, an additional sensitivity analysis was performed in which participants were categorized according to the direction of FORTA score change (delta FORTA > 0 vs < 0). Group differences were assessed using the Wilcoxon rank-sum test; however, this dichotomization was considered secondary due to its inherent limitations. The strength of correlation was classified according to [33]; all correlations reported here are ‘weak’ (e.g., < ± 0.3).

Statistical significance was assumed at p < 0.05. Statistical analyses were performed using SAS Version 9.4 software for Windows (SAS Institute Inc., Cary, NC, USA) and R version 4.5.1 (R Foundation for Statistical Computing, Vienna, Austria). Moreover, the STROBE checklist (Electronic Supplementary Material 1 [ESM 1] [34]) was used to ensure the inclusion of relevant data.

Comments (0)

No login
gif