The shaky voice of aging localized to the larynx: dissociation of frequency and amplitude tremor

Subjects

A total of 291 native speakers (126 men, 165 women) were recruited from August 2023 to September 2025. The overall study design, including the flow of subjects through clinical examination, is illustrated in Fig. 1. The enrollment was conducted to balance the age distribution for both sexes in the age range from 18 to 94 (mean 46 ± 20 SD) years for males and from 20 to 94 (mean 48 ± 20 SD) years for females. All participants had completed at least 8 years of elementary education. None of the participants were professional voice users, such as singers or actors. Each participant was examined by a trained speech-language pathologist (author I.B.) to exclude voice, speech, language, and communication impairment.

Fig. 1Fig. 1

The flowchart illustrates the experimental design, including data gathering, flow of subjects, and data analysis. R1 = first repetition of perceptual rating, R2 = second repetition of perceptual rating

We established the following inclusion criteria to ensure that the observed effects are not biased by voice-affecting conditions: no history of vocal pathologies or aerodigestive malignancies, no history of communication disorders including voice, speech (e.g., stuttering), and language and communication disorders; no history of hearing impairment or use of hearing aids; no history of psychiatric, cognitive, and neurological disorders; no history of thyroidectomy or any other voice-affecting surgery; no current use of medication with known effects on speech, such as antidepressants and antipsychotics; and no current abuse of recreational drugs such as marijuana and cocaine. The voice-affecting conditions and habits commonly expected in the population—including gastrointestinal reflux, respiratory diseases (asthma, allergies), endocrine diseases, local head and neck infections, history of smoking (all tobacco products, water pipes, and vaping), and alcohol abuse—were not exclusion criteria but were treated as factors in the statistical analyses.

To determine a clear effect of age, we defined a normative control group as subjects without voice-affecting medical conditions and with no history of smoking or alcohol abuse (n = 179). For studying the effect of voice-affecting medical conditions, only subjects affected by a single factor were analyzed: normative controls, smoking (n = 46), respiratory disease (n = 31), endocrine disease (n = 17), gastroesophageal reflux disease (n = 4), a combination of respiratory and endocrine disease (n = 2), and alcohol abuse (n = 2). Only normative controls, smoking, respiratory disease, and endocrine disease were compared, given the low numbers of subjects in the other groups. The rationale for studying smoking and alcohol abuse separately was that they were sparsely distributed across disease groups; statistical analysis of interaction would be misleading when clear categories can instead be defined.

The characteristics of each sex in the control groups under 40 years and above 40 years are summarized in Table 1. The division at 40 years was based on natural menopause for women and the expected decline of testosterone levels in males [24,25,26].

Table 1 Overview of cohort characteristics. Age, height, and weight are described by mean ± standard deviation (minimum–maximum). Other characteristics represent the total number of positive subjects within the category/total number of subjects in the category (percentual ratio of positive subjects per total number of subjects in the category)Clinical examination

All participants underwent standard otolaryngological evaluation, including indirect laryngoscopy without stroboscopy and perceptual voice assessment, to exclude subjects with overt vocal fold pathology and clinically apparent voice disorders. Each participant then underwent the following battery of questionnaires and tests: demographic questionnaire (age, gender, handedness, height, and weight), anamnestic interview (medical history of diseases and comorbidities with medication history), lifestyle habits (smoking, alcohol consumption), and acoustic voice examination (voice recording and acoustic analysis). The acoustic examination was held in a quiet room with low reverberation. The speech was recorded with a head-mounted omnidirectional condenser microphone with linear frequency response (DR4 made by Vocalgebra, Prague, Czech Republic) positioned approximately 4 cm from the speaker’s mouth. The speech signals were sampled at 44.1kHz with 16-bit resolution. Each participant was instructed to complete sustained phonation of the vowel/a/for as long and as steadily as possible, within one breath for at least 10s. Each subject performed phonation of the vowel/a/three times, individually.

Perceptual ratings

A total of 434 acoustic recordings of sustained vowel/a/from 227 subjects were blindly rated using the CAPE-V, where a score of 0% indicates a perfectly stable voice and 100% indicates a tremor strong enough to interrupt voicing or introduce pausing [27]. These rated recordings represent a subset of the complete database, as all recordings were initially rated and the database was subsequently extended. Examples of mild, moderate, and severe tremor were rated and selected based on the consensus of experienced speech and language pathologists to define the rating scale; these example recordings were available to raters throughout the perceptual rating as a reference. The strongest perceived tremor was rated preferentially. The continuity was scored as intermittent when the rated tremor occurred in intervals shorter than 5s.

The rating was performed using a simple graphical user interface (GUI) implemented in MATLAB (MathWorks, Natick, Massachusetts, USA), allowing one to score the perceived shakiness of the voice on a visual scale via a slider with 5% resolution. According to the original CAPE-V, mild, moderate, and severe tremors were classified as 10%, 35%, and 75%, respectively [27]. Recordings were presented in random order by the GUI, without revealing the identities of the subjects or the recording names, and were played via headphones connected to the computer. The raters (ENT specialist B.B. and speech and language pathologist I.B.—the authors of the study) were thoroughly trained before the experiment and scored each recording twice for a baseline and retest comparison. Baseline and retest values were averaged to describe each speaker with a single representative value.

Acoustic analysis

The recording system was measured in an anechoic chamber to verify its capacity to capture the subtle acoustic fluctuations associated with vocal tremor. Acoustic signals were first manually examined for quality using oscillograms and spectrograms. Tremor was then analyzed automatically using the Dysarthria Analyzer (www.dysan.cz) [20, 28]. The method extracts the signal amplitude envelope and modal fundamental frequency contour, both decimated to 50 Hz, and analyzes them within voiced intervals using a 2-s sliding window with 100-ms steps. Within each window, the signal is detrended, Gaussian-weighted, and analyzed spectrally via the Fast Fourier Transform over the 1–20 Hz frequency range. A tracking algorithm links spectral peak candidates across successive windows into continuous tremor tracks, identifying the dominant tremor as the track with the highest cumulative modulation depth spanning at least 3 s or 50% of phonation duration.

This approach distinguishes clinically significant tremor from physiological oscillations and random movements and has been validated on synthetic data with known tremor parameters, demonstrating accuracy substantially superior to conventional long-term averaged spectrum methods [28]. Tracking was performed on signal amplitude (loudness variation) and fundamental frequency (melody changes). Tremor strength was quantified as the prominence of amplitude tremor (PAT) and the prominence of frequency tremor (PF0T), representing modulation depth in percent amplitude and semitones, respectively.

Figure 2 illustrates modulograms and corresponding PF0T values for various subjects and conditions. All outlying measured values were inspected using modulograms, oscillograms, and spectrograms to ensure the validity of the results.

Fig. 2Fig. 2

Examples of PF0T modulograms with the detected dominant tremor track are highlighted as black scatter. Individual modulograms demonstrate various subjects including female (A) and male (D) subjects (< 40 years old) without medical conditions, no alcohol abuse, and no history of smoking; a female subject > 40 years old without a medical condition but with alcohol abuse and history of smoking (B); a female subject > 40 years old with gastrointestinal reflux disease, with no alcohol abuse and no history of smoking (C); a male subject > 40 years old without a medical condition, no alcohol abuse, and with history of smoking (E); and a male subject > 40 years old with cardiovascular disease, no alcohol abuse, and no history of smoking (F). X-axes represent time in seconds, while Y-axes represent the frequency of the measured tremor tracks. Each subplot is accompanied by a colormap legend scaled for the range of measured prominence/depth of tremor oscillations, where the darker color represents deeper oscillations and the white color indicates no oscillations. PF0T = prominence of frequency tremor

Statistical analysis

The minimal sample size for testing a medium effect of correlation (R = 0.4) between age and clinical characteristics, given a probability of rejecting the null hypothesis of 0.05 and a power of 0.8, was estimated to be 46 [29]. The normality of distributions was tested using the Lilliefors test on the normative control group. We compared the effect of sexual dimorphism on the normative control group via t-test for normally distributed data and Wilcoxon’s rank-sum test for nonnormally distributed variables. The relationships between acoustic and clinical variables were tested on the normative control group via Pearson’s and Spearman’s correlation coefficient for normally and nonnormally distributed data, respectively. Interpretation of correlation coefficients followed the guide by Chan [30]. Correlation tests were conducted separately for each sex if sexual dimorphism was observed for the acoustic feature. The threshold of significant probability was set to p < 0.05. The inflation of Type I errors in correlation tests was controlled via Bonferroni’s correction.

Analysis of interrater correlations was conducted on the average rating of each rater. Relationships of perceptual ratings to other variables were tested using an average of all perceptual ratings across both raters.

The normative characteristics were determined by fitting a regression line of age and acoustic biomarkers in the normative control group subjects for acoustic features with a significant correlation to age. The residuals of the fit were inspected to test the validity of the regression model. Scaling to the natural logarithm was used for log-normally distributed data. Acoustic features uncorrelated with age were modeled with mean and standard deviation for normally distributed data or median with median absolute deviation (MAD) scaled to standard deviation for nonnormally distributed data. The model and its residuals were used to determine z-scores of acoustic features for all males and females in the dataset.

Finally, we compared the normative control group with the groups having a history of smoking, respiratory disease, and endocrine disease, to determine their effect on PF0T and PAT z-scores using the Kruskal–Wallis omnibus test followed by a many-to-one nonparametric comparison [31] with normative controls as the reference. This design, which was preferable given the unbalanced data with low presence of smoking subjects in the respiratory and endocrine disease groups, breaks the assumptions for two-way statistical comparison. Statistical analysis was performed in R (version 4.5.2).

Comments (0)

No login
gif