Evidence of cross-domain phase entrainment effects between nonspeech tones and speech sounds

Human listeners use event duration to accurately identify complex sounds such as speech and music. One hypothesized, supramodal neural mechanism underlying duration estimation and other temporal processing is entrainment (originally proposed by Jones, 1976), in which brain oscillations synchronize with external rhythms to scaffold the processing of temporally predictable sensory cues such as musical beats or word stress (Goswami, 2012, Goswami, 2019, Haegens and Zion Golumbic, 2018, Harding et al., 2025, Lakatos et al., 2019, Poeppel and Assaneo, 2020, Rimmele et al., 2018). Researchers have claimed that language processing is shaped by entrainment (Dilley and McAuley, 2008, Doelling et al., 2014, Giraud and Poeppel, 2012, Heffner et al., 2017, Steffman, 2021, Zion Golumbic et al., 2013). Many findings further highlight a strong link between musical and language processing (Fiveash et al., 2021, Fujii and Wan, 2014, Goswami, 2012, Ladányi et al., 2020), hinting at a shared, domain-general mechanism for temporal processing (Bauer et al., 2020). Despite this intense interest, little research addresses how entrainment is transferred among different acoustic domains. Here, we test phase entrainment effects both within domains (nonspeech tone precursors, tone targets; speech precursors, speech targets) and across domains (tone precursors, speech targets; speech precursors, tone targets).

Existing research suggests that the phase of a target stimulus relative to previous events affects its processing. Numerous researchers have found better sensory detection of in-phase events when brain oscillations entrain to match stimulus rhythm (de Graaf et al., 2013, Henry et al., 2014, Henry et al., 2016, Hickok et al., 2015, Kizuk and Mathewson, 2017, Mathewson et al., 2010, Mathewson et al., 2012, Spaak et al., 2014; see Haegens and Zion Golumbic, 2018 Section 2 for a thorough review). Behaviorally, when a sequence of tones evenly spaced in time precedes a target tone, in-phase tone targets are processed more accurately than out-of-phase (slightly early or late) targets (Barnes and Jones, 2000, Hickok et al., 2015, Jones et al., 2002, McAuley and Jones, 2003, McAuley and Kidd, 1998; see Saberi & Hickok, 2023 for a comprehensive discussion). Additionally, multiple researchers (McAuley and Jones, 2003, McAuley and Kidd, 1998) including us (Cheng & Creel, 2020) have shown that perception of the target’s duration is distorted systematically, with early targets reported as shorter than on-time or late targets of equal duration, and late targets perceived as slightly longer (Fig. 1). This entrainment distortion provides a useful behavioral probe of presence or absence of entrainment.

An exciting emerging trend is the extension of the entrainment findings from nonspeech sounds (sine tones, light flashes) to speech sounds (Ten Oever & Martin, 2021, Ten Oever and Sack, 2015, Ten Oever et al., 2024), in which duration information contributes to listeners’ recognition of sounds/words. For example, coda consonant voicing in English is differentiated by (among numerous other features) the duration of the preceding vowel (Raphael, 1972), such as lap (shorter vowel) vs. lab (longer vowel; see Section 2.1.2 in Methods for more details).

Perception of speech stimuli is affected by the temporal properties of both speech and nonspeech precursors (Diehl et al., 1980, Port, 1979, Summerfield, 1981). A few recent studies have investigated entrainment-like effects on duration perception in spoken language by manipulating rate and rhythmic properties of precursor speech or nonspeech. Precursor speech rate (Heffner et al., 2017) and precursor speech rhythmic patterns (Steffman, 2021) affect perception of coda voicing of the target word (see related work by Kidd, 1989; and work on segmentation by Dilley and McAuley, 2008). Especially relevant to the current work, a few studies have reported context rate effects of nonspeech precursors on speech categorization: faster tones bias listeners to perceive target speech sounds as longer, while slower tones bias listeners to perceive target speech sounds as shorter, thus changing apparent word identity (Dutch /as/ vs. the longer-duration /a:s/ in Bosker, 2017; English [b] vs. the longer-duration [w] in Wade & Holt, 2005, but also see Pitt et al., 2016, who found that only precursors perceived as speech can generate context rate effects on speech perception). These studies are predominantly consistent with a general auditory timing mechanism underlying speech perception, counter to modular accounts of language processing (Liberman et al., 1961, Mattingly et al., 1971).

Yet only one study so far (Bosker and Kösem, 2017) has examined phase effects from nonspeech tone precursors on speech perception, and it found no phase influences. This is crucial for claims that speech undergoes entrainment, in that phase is the element in the originally proposed entrainment model that distinguishes interval-based timing and entrainment-based timing (see Jones, 1976, McAuley and Jones, 2003, Repp, 2002, Cheng and Creel, 2020 for detailed description) and appears critical to findings of attentional benefit for in-phase stimuli (de Graaf et al., 2013, Henry et al., 2014, Henry et al., 2016, Hickok et al., 2015, Kizuk and Mathewson, 2017, Mathewson et al., 2010, Mathewson et al., 2012, Spaak et al., 2014; Saberi & Hickok, 2023). In brief, it is possible that an interval model tracking the prevalence of particular durations could account for precursor rate effects. However, because interval models do not track phase, they cannot account for phase effects. If entrainment is occurring, duration distortion effects based on the phase relationship between the context and the target should also occur.

To date, it is unclear to what extent precursor entrainment, especially phase, spreads from one acoustic domain in the precursor (e.g., tones) to another domain for the target (e.g., speech), and it’s also unclear whether speech is subject to phase entrainment at all, even from other speech. This leaves uncertain whether entrainment is truly domain-general, cross-cutting speech and nonspeech, or domain-specific, perhaps limited to highly simplified perceptual materials or musical contexts. If entrainment is occurring during speech perception, then duration distortion effects based on the phase relationship between the context and the target should also occur. This raises the question of whether speech is truly impervious to phase entrainment, which would suggest that entrainment is domain-specific, or, instead, whether controlling for additional factors might reveal phase entrainment effects in speech, which would suggest that entrainment is a domain-general process.

Here we asked whether there are cross-domain phase entrainment effects between nonspeech tones and speech sounds, and whether speech sounds are subject to entrainment generally. An answer to this question contributes to the long-standing debate between domain-specific and domain-general accounts of speech processing as well as a fundamental understanding of entrainment mechanisms in speech and nonspeech tone perception. We measured entrainment distortion (Fig. 1) as a function of the phase of the target onset (early, on-time, late) as in our previous study (Cheng & Creel, 2020) across different auditory targets (Speech vs. speech-shaped Tone). We had two goals. Our primary goal was to ask whether phase entrainment is domain-specific vs. domain-general. If entrainment is domain-general, entrainment distortion should occur for both speech-shaped tone and speech target regardless of the tone/speech precursors. If it is domain-specific in the sense that only nonspeech sounds are susceptible to phase entrainment, then speech targets may not entrain at all. These predictions were addressed in three experiments, with different sets of speech sound targets and corresponding tone targets (Experiment 1: lab/lap; Experiments 2 and 3: add/at). Note that Experiment 1′s outcome was foreseen in two earlier, separately-run pilot experiments (Supplementary Section 1), but here we explored a different set of speech sounds/matched tones and ran each of the current studies as a unitary experiment with random assignment to conditions to verify that previous outcome. Also note that, unlike the sine-wave context tones, all tone targets in this study are speech-shaped tones, modulated by the speech envelope to enhance their similarity to speech. For convenience, speech-shaped tone targets will henceforth be referred to simply as “tone targets.”

Our secondary goal was, if no entrainment is observed for speech (as implied by our earlier experiments), to make an initial probe as to potential reasons why. One possibility is that auditory dissimilarity between context tones and target words blocks phase effects of entrainment, which we term the auditory grouping hypothesis. To this end, we included in Experiment 1 an additional Tones-as-Speech condition which asked subjects to classify Tone targets as speech stimuli, and across Experiments 2–3, we varied the nature of the entrainers themselves (tones in Experiment 2 vs. speech sounds in Experiment 3). On the auditory grouping hypothesis, Tones-as-Speech targets, which do auditorily group with context tones while subjects use their internal representations to categorize the words, should show entrainment. Further, entrainment distortion should occur for tone targets but not for speech targets after tone precursors (Experiments 1–2), and it should occur for speech targets but not for tone targets after speech precursors (Experiment 3). Another possible reason why speech might not show entrainment is that subjectively perceived event onset time between tone and speech, known as the perceptual center (“p-center”) is misaligned. Various researchers (Cooper et al., 1986, Marcus, 1981, Morton et al., 1976) have found p-centers to be influenced by many factors such as the type of consonants prior to the vowel, amplitude envelope, and more. We attempted to control for this in Experiment 1 by matching amplitude envelopes between speech and tone stimuli, but to prefigure our findings somewhat, a control study revealed p-center differences between Experiment 1’s tone and speech stimuli. This and a second control study also identified a set of speech stimuli with p-centers similar to tones, which we used in Experiment 2 and 3.

Comments (0)

No login
gif