Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2016 Dec 20:7:13619.
doi: 10.1038/ncomms13619.

Perceptual restoration of masked speech in human cortex

Affiliations

Perceptual restoration of masked speech in human cortex

Matthew K Leonard et al. Nat Commun. .

Abstract

Humans are adept at understanding speech despite the fact that our natural listening environment is often filled with interference. An example of this capacity is phoneme restoration, in which part of a word is completely replaced by noise, yet listeners report hearing the whole word. The neurological basis for this unconscious fill-in phenomenon is unknown, despite being a fundamental characteristic of human hearing. Here, using direct cortical recordings in humans, we demonstrate that missing speech is restored at the acoustic-phonetic level in bilateral auditory cortex, in real-time. This restoration is preceded by specific neural activity patterns in a separate language area, left frontal cortex, which predicts the word that participants later report hearing. These results demonstrate that during speech perception, missing acoustic content is synthesized online from the integration of incoming sensory cues and the internal neural dynamics that bias word-level expectation and prediction.

PubMed Disclaimer

Figures

Figure 1
Figure 1. Stimuli and single electrode online phoneme restoration effects.
(a,b) Participants listened to pairs of spoken words (/fæstr/ (a) versus/fæktr/ (b)) that were acoustically identical except for a critical phoneme that differentiated their meaning (vertical solid and second dashed lines; first dashed line is word onset). (c) The critical phoneme was also replaced by broadband noise (/fæ#tr/), and on each trial, participants reported which word they heard. (d) Behavioural results show bistable perception on noise trials. (e) Location of representative posterior STG electrode in f. (f) STG electrode shows selectivity for /s/ compared to /k/ (solid blue line stronger response than solid red line immediately after critical phoneme, unshaded region). Trials were sorted depending on which word participants perceived. Responses to noise stimuli were similar to the original version of the perceived phoneme (dotted lines; *signifies 99% CIs only overlapping for same coloured curves; shaded error±s.e.m. across trials). (g) RI describes the magnitude of neural restoration as the relative distances between each noise and original pair in f. When the dotted line is in the region shaded with the same colour, the electrode's activity reflects the participant's percept. (h) Across all participants, word pairs and electrodes, the magnitude of the difference between RI values illustrates that when these neural populations differentiate original stimuli, they also differentiate noise trials, beginning at the onset of the critical phoneme (red bar, one-way t tests, P<0.05, Bonferroni corrected). Shaded error±s.e.m. across word pairs.
Figure 2
Figure 2. Stimuli and single electrode online phoneme restoration effects for a representative word pair where the participant did not show bistable perception.
(a,b) Subjects listened to pairs of spoken words (/wƅkǝrz/ (a) versus/ wƅtǝrz/ (b)) that were acoustically identical except for a critical phoneme that differentiated their meaning (vertical solid and second dashed lines; first dashed line is word onset). (c) The critical phoneme was also replaced by broadband noise (/wƅ#ǝrz/), and on each trial subjects reported which word they heard. (d) Behavioural results showed that the noise was always perceived as /t/. (e) Location of representative STG electrode in f. Data are from the same subject as in Fig. 1. (f) Single representative left hemisphere STG electrode shows selectivity for /k/ compared with /t/ (solid blue line stronger response than solid red line immediately after critical phoneme, unshaded region). Responses to noise stimuli were similar to the original version of /t/ (dotted red line; *signifies 99% CIs only overlapping for red curves; shaded error±s.e.m. across trials). (g) RI describes the magnitude of neural restoration as the relative distances between each noise and original pair in f. When the dotted line is in the region shaded with the same colour, the electrode's activity reflects the subject's percept. (h) Across all word pairs that did not exhibit bistable perception, the average timecourse of the RI metric for all electrodes shows neural restoration effects beginning ∼150 ms after critical phoneme onset (red bar: one-way t test against baseline, P<0.05, false discovery rate corrected for time points). Shaded error±s.e.m. across word pairs.
Figure 3
Figure 3. Stimulus spectrogram reconstruction reveals warping of noise to perceived phoneme.
(a,b) Acoustic spectrograms for a representative word pair (/fæstr/, (a), versus /fæktr/, (b)) differ primarily in the presence of a high-frequency component during the critical phoneme in a (green arrow). (c,d) Spectrograms from (a,b) reconstructed from electrode population activity show that the high-frequency component is present in / fæstr/ (c, green arrow) and absent in /fæktr/ (d). (e,f) Spectrogram reconstruction of noise trials was divided according to which word the participant heard on each trial. During the critical phoneme, a high-frequency component is visible only for trials perceived as /fæstr/ (e, green arrow) and not for /fæktr/ (f). (g) Power spectra of the critical phoneme for c–f show close correspondence between noise and original phonemes, particularly in mid-high frequencies.
Figure 4
Figure 4. Timecourse of stimulus classification shows pre-stimulus frontal lobe bias for restored phonemes.
(a) Trials were classified using population neural activity and compared with reported perception. Original word classification accuracy peaked ∼200 ms after critical phoneme onset (black line, blue arrow). Noise trial classification accuracy was similar, and showed above-chance classification before critical phoneme onset (green line; orange arrow). Shaded error±s.e.m. across word pairs. (b–e) Classification weights for all subjects mapped onto a common cortical surface (MNI). During the pre-critical phoneme period, classification performance was driven by bilateral superior temporal cortex for original (b) and noise (c) trials. Noise trials also showed significantly greater weights in left inferior frontal cortex compared with original trials (orange box). During the post-critical phoneme period, classification performance was driven by bilateral superior temporal cortex for original (d) and noise (e) trials, with greater weights in left superior temporal cortex for original trials (blue box).

References

    1. Guediche S., Blumstein S. E., Fiez J. A. & Holt L. L. Speech perception under adverse conditions: insights from behavioral, computational, and neuroscience research. Front. Syst. Neurosci. 7, 126 (2013). - PMC - PubMed
    1. Samuel A. G. Phonemic restoration: insights from a new methodology. J. Exp. Psychol. Gen. 110, 474 (1981). - PubMed
    1. Warren R. M. Perceptual restoration of missing speech sounds. Science 167, 392–393 (1970). - PubMed
    1. Bregman A. S. Auditory Scene Analysis: The Perceptual Organization Of Sound MIT press (1990).
    1. Miller G. A. & Licklider J. The intelligibility of interrupted speech. J. Acoust. Soc. Am. 22, 167–173 (1950).

Publication types