IP Library › Granted Patent US 12,444,429
Granted Patent B2
US 12,444,429 · App. 18/085,705 · Granted Oct 14, 2025

Speech identification and extraction from noise using extended high frequency information

Inventors: Brian B Monson (Champaign, IL); Rohit M Ananthanarayana (Champaign, IL)
Assignee: The Board of Regents of the University of Illinois
G10L21/0232G10L21/0224G10L21/0272G10L21/0308G10L25/09G10L25/18G10L25/21G10L25/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,429
App. No.
18/085,705
Granted
Oct 14, 2025
Kind
B2
Abstract

Improved systems and methods are provided herein for extracting target speech from audio signals that can contain masking speech or other unwanted noise content. These systems and methods include detection of target speech in an input signal by detecting elevated frequency content in the signal above a threshold frequency. Portions of the signal determined to contain such elevated high frequency content are then used to generate audio filters to extract target speech from subsequently-obtained audio signals. This can include performing non-negative matrix factorization to determine a set of basis vectors to represent noise content in the spectral domain and then using the set of basis vectors to decompose subsequently-obtained audio signals into noise signals that can then be removed from the audio signals.

Claims (74)

1. A non-transitory computer readable medium comprising program instructions executable by at least one processor to cause the at least one processor to perform a method comprising:

obtaining a first audio sample;

determining that a first portion of the first audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds a threshold energy level;

responsive to determining that the first portion contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level, determining a first audio filter based on the first portion of the first audio sample by:

determining a first spectrogram for the first portion; and

performing non-negative matrix factorization to generate a first matrix and a second matrix whose product corresponds to a low-frequency portion of the first spectrogram that is below a threshold frequency, wherein the first matrix is composed of a set of column vectors that span along a frequency dimension of the first spectrogram, and wherein the second matrix is composed of a set of row vectors that span along a time dimension of the first spectrogram;

subsequent to obtaining the first audio sample, obtaining a second audio sample; and

applying the first audio filter to the second audio sample to generate a first audio output by:

determining a second spectrogram for the second audio sample;

applying the first matrix to a low-frequency portion of the second spectrogram that is below the threshold frequency to generate a third spectrogram that represents noise content of the second audio sample; and

using the third spectrogram to remove the noise content from the second audio sample, thereby generating the first audio output.

2. The non-transitory computer readable medium of claim 1 , wherein the method further comprises:

determining a plurality of zero-crossing rates across time for the first audio sample; and

determining a plurality of signal energy levels across time for the first audio sample, wherein determining that the first portion contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level comprises determining (i) that a zero-crossing rate, of the plurality of zero-crossing rates, that corresponds to the first portion exceeds a threshold zero-crossing rate and (ii) that a signal energy level, of the plurality of signal energy levels, that corresponds to the first portion exceeds a threshold signal energy level.

3. The non-transitory computer readable medium of claim 1 , wherein the first audio sample is divided into a plurality of non-overlapping frames, and wherein determining that the first portion contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level comprises:

determining that a contiguous subset of the plurality of non-overlapping frames of the first audio sample all contain frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level, wherein the first portion consists of the contiguous subset of frames of the first audio sample.

4. The non-transitory computer readable medium of claim 3 , wherein each frame of the plurality of non-overlapping frames of the first audio sample has a duration between 15 milliseconds and 50 milliseconds.

5. The non-transitory computer readable medium of claim 1 , wherein the method further comprises:

prior to obtaining the first audio sample, obtaining a third audio sample;

determining that a second portion of the third audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level; and

responsive to determining that the second portion of the third audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level, determining a second audio filter based on the second portion of the third audio sample by:

determining a fourth spectrogram for the second portion; and

performing non-negative matrix factorization to generate a third matrix and a fourth matrix whose product corresponds to a portion of the fourth spectrogram below the threshold frequency, wherein the third matrix is composed of a further set of column vectors that span along a frequency dimension of the fourth spectrogram, and wherein the fourth matrix is composed of a further set of row vectors that span along a time dimension of the fourth spectrogram,

wherein performing non-negative matrix factorization to generate the first matrix and the second matrix comprises using, as an initial estimate of the first matrix, the third matrix.

6. The non-transitory computer readable medium of claim 1 , wherein determining that the first portion contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level comprises:

determining a spectrogram for the first portion; and

determining that a total energy in the spectrogram above 5.6 kilohertz exceeds the threshold energy level.

7. The non-transitory computer readable medium of claim 1 , wherein using the third spectrogram to remove the noise content from the second audio sample comprises:

performing an inverse transform on the third spectrogram to generated a time-domain noise signal; and

subtracting the time-domain noise signal from the second audio sample to generate the first audio output.

8. The non-transitory computer readable medium of claim 1 , wherein the method further comprises:

prior to obtaining the first audio sample, obtaining a third audio sample;

determining that a second portion of the third audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level; and

responsive to determining that the second portion of the third audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level, determining a second audio filter based on the second portion of the third audio sample,

wherein determining the first audio filter based on the first portion comprises:

determining a third audio filter based on the first portion; and

determining the first audio filter as a weighted combination of the second audio filter and the third audio filter.

9. A method comprising:

obtaining a first audio sample;

determining that a first portion of the first audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds a threshold energy level;

responsive to determining that the first portion contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level, determining a first audio filter based on the first portion of the first audio sample by:

determining a first spectrogram for the first portion; and

performing non-negative matrix factorization to generate a first matrix and a second matrix whose product corresponds to a low-frequency portion of the first spectrogram that is below a threshold frequency, wherein the first matrix is composed of a set of column vectors that span along a frequency dimension of the first spectrogram, and wherein the second matrix is composed of a set of row vectors that span along a time dimension of the first spectrogram;

subsequent to obtaining the first audio sample, obtaining a second audio sample; and

applying the first audio filter to the second audio sample to generate a first audio output by:

determining a second spectrogram for the second audio sample;

applying the first matrix to a low-frequency portion of the second spectrogram that is below the threshold frequency to generate a third spectrogram that represents noise content of the second audio sample; and

using the third spectrogram to remove the noise content from the second audio sample, thereby generating the first audio output.

10. The method of claim 9 , further comprising:

determining a plurality of zero-crossing rates across time for the first audio sample; and

determining a plurality of signal energy levels across time for the first audio sample, wherein determining that the first portion contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level comprises determining (i) that a zero-crossing rate, of the plurality of zero-crossing rates, that corresponds to the first portion exceeds a threshold zero-crossing rate and (ii) that a signal energy level, of the plurality of signal energy levels, that corresponds to the first portion exceeds a threshold signal energy level.

11. The method of claim 9 , wherein the first audio sample is divided into a plurality of non-overlapping frames, and wherein determining that the first portion contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level comprises:

determining that a contiguous subset of the plurality of non-overlapping frames of the first audio sample all contain frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level, wherein the first portion consists of the contiguous subset of frames of the first audio sample.

12. The method of claim 11 , wherein each frame of the plurality of non-overlapping frames of the first audio sample has a duration between 15 milliseconds and 50 milliseconds.

13. The method of claim 9 , further comprising:

prior to obtaining the first audio sample, obtaining a third audio sample;

determining that a second portion of the third audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level; and

responsive to determining that the second portion of the third audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level, determining a second audio filter based on the second portion of the third audio sample by:

determining a fourth spectrogram for the second portion; and

performing non-negative matrix factorization to generate a third matrix and a fourth matrix whose product corresponds to a portion of the fourth spectrogram below the threshold frequency, wherein the third matrix is composed of a further set of column vectors that span along a frequency dimension of the fourth spectrogram, and wherein the fourth matrix is composed of a further set of row vectors that span along a time dimension of the fourth spectrogram,

wherein performing non-negative matrix factorization to generate the first matrix and the second matrix comprises using, as an initial estimate of the first matrix, the third matrix.

14. The method of claim 9 , wherein determining that the first portion contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level comprises:

determining a spectrogram for the first portion; and

determining that a total energy in the spectrogram above 5.6 kilohertz exceeds the threshold energy level.

15. The method of claim 9 , wherein using the third spectrogram to remove the noise content from the second audio sample comprises:

performing an inverse transform on the third spectrogram to generated a time-domain noise signal; and

subtracting the time-domain noise signal from the second audio sample to generate the first audio output.

16. The method of claim 9 , further comprising:

prior to obtaining the first audio sample, obtaining a third audio sample;

determining that a second portion of the third audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level; and

responsive to determining that the second portion of the third audio sample contains frequency content at frequencies higher than 5.6 kilohertz that exceeds the threshold energy level, determining a second audio filter based on the second portion of the third audio sample,

wherein determining the first audio filter based on the first portion comprises:

determining a third audio filter based on the first portion; and

determining the first audio filter as a weighted combination of the second audio filter and the third audio filter.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2023
From: MONSON, BRIAN; ANANTHANARAYANA, ROHIT M.
To: THE BOARD OF TRUSTEES OF THE UNIVERSITY OF ILLINOIS
Reel/Frame 064424/0028 →
Continuity (2)
Provisional Application 63292307 · Dec 21, 2021
Related Publication 20230197099A1 · Jun 22, 2023
References Cited (132)
US 9165567B2 · Visser · 2015 [cited by examiner]
US 9467569B2 · Femal · 2016 [cited by examiner]
US 9552825B2 · Venkatesha · 2017 [cited by examiner]
US 10923111B1 · Fan · 2021 [cited by examiner]
US 11479184B2 · Kobayashi · 2022 [cited by examiner]
US 11482225B2 · Pinder · 2022 [cited by examiner]
US 11631404B2 · Pereira · 2023 [cited by examiner]
US 12217761B2 · Chen · 2025 [cited by examiner]
US 20170318374A1 · Dolenc · 2017 [cited by examiner]
US 20240144952A1 · Ikeshita · 2024 [cited by examiner]
Trine, Allison, and Brian B. Monson, “Extended High Frequencies Provide Both Spectral and Temporal Information to Improve Speech-in-Speech Recognition”, Dec. 21, 2020, Trends in Hearing, vol. 24, pp. 1-8. (Year: 2020). [cited by examiner]
Levy, Suzanne Carr, Daniel J. Freed, Michael Nilsson, Brian C. J. Moore, and Sunil Puria, “Extended high-frequency bandwidth improves reception of speech in spatially separated masking speech”, Sep. 2015, Ear and Hearin… [cited by examiner]
Gaikwad, Snehal S., Pallavi P. Ingale, and S. L. Nalbalwar, “Separation of singing voice from background musical noise using modified NMF and Filtering”, Mar. 2016, 2016 International Conference on Electrical, Electroni… [cited by examiner]
Nathwani, Karan, Anurag Kumar, and Rajesh M. Hegde, “Monaural Speaker Segregation using Group Delay Spectral Matrix Factorization”, May 2014, 2014 Twentieth National Conference on Communications (NCC), pp. 1-6. (Year: 2… [cited by examiner]
Nikunen, Joonas, and Tuomas Virtanen, “Source Separation and Reconstruction of Spatial Audio Using Spectrogram Factorization” , Oct. 2017, Parametric Time-Frequency Domain Spatial Audio, John Wiley & Sons, pp. 215-250. … [cited by examiner]
Feng, Yuxiao, and Christian Ritz, “Single-channel speech separation by including spectral structure information within non-negative matrix factorization”, Jul. 2015, 2015 IEEE China Summit and International Conference o… [cited by examiner]
Kim, Seon Man, Ji Hun Park, Hong Kook Kim, Sung Joo Lee, and Yun Keun Lee, “Non-negative Matrix Factorization Based Noise Reduction for Noise Robust Automatic Speech Recognition”, Mar. 2012, 10th International Conferenc… [cited by examiner]
Mohammadiha, Nasser, Timo Gerkmann, and Arne Leijon, “A New Approach for Speech Enhancement Based on a Constrained Nonnegative Matrix Factorization”, Dec. 2011, International Symposium on Intelligent Signal Processing a… [cited by examiner]
Virtanen, Tuomas, “Monaural Sound Source Separation by Nonnegative Matrix Factorization with Temporal Continuity and Sparseness Criteria”, Mar. 2007, IEEE Transactions on Audio, Speech, and Language Processing, vol. 15,… [cited by examiner]
Dos Santos Moura, Mateus, Alexandre Miccheleti Lucena, Kenji Nose Filho, and Ricardo Suyama, “Source Extraction based on Binary Masking and Machine Learning”, Nov. 2021, 6th Workshop on Communication Networks and Power … [cited by examiner]
Shoba, S., and R. Rajavel, “Adaptive energy threshold for monaural speech separation”, Apr. 2017, 2017 International Conference on Communication and Signal Processing (ICCSP), pp. 0905-0908. (Year: 2017). [cited by examiner]
Monson et al., “Horizontal directivity of low- and high-frequency energy in speech and singing”, J. Acoust. Soc. Am. 132 (1), Jul. 2012, pp. 433-441. [cited by applicant]
Middlebrooks, “Sound localization”, Handbook of Clinical Neurology, vol. 129 (3rd series), Chapter 6, 2015, pp. 99-116. [cited by applicant]
McDermott et al., “Sound Texture Perception via Statistics of the Auditory Periphery: Evidence from Sound Synthesis”, Neuron Article, 71, Sep. 8, 2011, pp. 926-940. [cited by applicant]
Masterton et al., “The Evolution of Human Hearing”, The Journal of the Acoustical Society of America, vol. 45, No. 4, (1969) pp. 966-985. [cited by applicant]
Martin et al., “Spatial release from speech-on-speech masking in the median sagittal plane”, J. Acoust. Soc. Am. 131 (1), Jan. 2012, pp. 378-385. [cited by applicant]
Geoffrey A. Manley, “Comparative Auditory Neuroscience: Understanding the Evolution and Function of Ears”, Journal of the Association for Research in Otolaryngology, 18: 1-24, 2017. [cited by applicant]
Richard P. Lippmann, “Accurate Consonant Perception Without Mid-Frequency Speech Energy”, IEEE Transactions on Speech and Audio Processing, vol. 4, No. 1, Jan. 1996, pp. 66-69. [cited by applicant]
Liberman et al., “The Discrimination of Speech Sounds Within and Across Phoneme Boundaries”, Journal of Experimental Psychology, vol. 54, No. 5, 1957, pp. 358-368. [cited by applicant]
Levy et al., “Extended high-frequency bandwidth improves reception of speech in spatially separated masking speech”, Ear Hear. 2015; 36(5), pp. 1-27. [cited by applicant]
H. Levitt, “Transformed Up-Down Methods in Psychoacoustics”, The Journal of the Acoustical Society of America, vol. 49, No. 2, 1971, pp. 467-477. [cited by applicant]
Kuhl et al., “Linguistic Experience Alters Phonetic Perception in Infants by 6 Months of Age”, Science, vol. 255, Jan. 31, 1992, pp. 606-608. [cited by applicant]
Kocon et al., “Horizontal directivity patterns differ between vowels extracted from running speech”, J. Acoust. Soc. Am. 144 (1), Jul. 2018, pp. EL7-EL12. [cited by applicant]
Sergei Kochkin, “MarkeTrak VIII: Consumer satisfaction with hearing aids is slowly increasing”, The Hearing Journal, Jan. 2010, vol. 63, No. 1, pp. 19-32. [cited by applicant]
King et al., “The Impack of Signal Bandwidth on Auditory Localization: Implications for the Design of Three-Dimensional Audio Displays”, Human Factors, 1997, 39(2), 287-295. [cited by applicant]
Kato et al., “Spatial acoustic cues for the auditory perception of speaker's facing direction”, Proceedings of 20th International Congress on Acoustics, ICA 2010, pp. 1-8. [cited by applicant]
Imbery et al., “Auditory Facing Angle Perception: The Effect of Different Source Positions in a Real and an Anechoic Environment”, ACTA Acustica United with Acustica, vol. 105, 2019, pp. 492-505. [cited by applicant]
Hoy et al., “The Evolution of Hearing in Insects as an Adaptation to Predation from Bats”, The Evolutionary Biology of Hearing, 1992, pp. 115-129. [cited by applicant]
Heffner et al., “High-Frequency Hearing”, 2008, pp. 55-60. [cited by applicant]
Heffner et al., “Hearing Ranges of Laboratory Animals”, Journal of the American Association for Laboratory Animal Science, vol. 46, No. 1, Jan. 2007, pp. 20-22. [cited by applicant]
Rickye S. Heffner, “Primate Hearing From a Mammalian Perspective”, The Anatomical Record Part A, 281A:1111-1122, 2004. [cited by applicant]
“The Evolution of Communication: Historical Overview”, University of Illinois, Urbana Champaign, Chapter 2, Mar. 2023, pp. 17-70. [cited by applicant]
Halkosaari et al., “Directivity of Artificial and Human Speech”, J. Audio Eng. Socl, vol. 53, No. 7/8, Jul./Aug. 2005, pp. 620-631. [cited by applicant]
Green et al., “High-frequency audiometric assessment of a young adult population”, J. Acoust. Soc. Am. 81 (2), Feb. 1987, pp. 485-494. [cited by applicant]
Glasberg et al., “Derivation of auditory filter shapes from notched-noise data”, Hearing Research, 47 (1990), pp. 103-138. [cited by applicant]
Fletcher et al., “Articulation Testing Methods”, Downloaded on Mar. 8, 2023, IEEE, pp. 806-854. [cited by applicant]
Fletcher et al., “The Perception of Speech and Its Relation to Telephony”, The Journal of the Acoustical Society of America, vol. 22, No. 2, 1950, pp. 89-151. [cited by applicant]
Richard R. Fay, “Comparative psychoacoustics”, Hearing Research, 34, 1988, pp. 295-306. [cited by applicant]
Edlund et al., “On the effect of the acoustic environment on the accuracy of perception of speaker orientation from auditory cues alone”, 2012, 5 pages. [cited by applicant]
Derey et al., “Localization of complex sounds is modulated by behavioral relevance and sound category”, J. Acoust. Soc. Am. 142 (4), Oct. 2017, pp. 1757-1773. [cited by applicant]
Crouzet et al., “On the Various Influences of Envelope Information on the Perception of Speech in Adverse Conditions: An Analysis of Between-Channel Envelope Correlation”, 2001, 4 pages. [cited by applicant]
Chu et al., “Detailed Directivity of Sound Fields Around Human Talkers”, National Research Council Canada, Dec. 2022. [cited by applicant]
Cherry, “Some Experiments on the Recognition of Speech, with One and with Two Ears”, The Journal of the Acoustical Society of America, vol. 25, No. 5, Sep. 1953, pp. 975-979. [cited by applicant]
Carlile et al., “the localisation of spectrally restricted sounds by human listeners”, Hearing Research 128 (1999), pp. 175-189. [cited by applicant]
Brungart et al., “Effects of bandwidth on auditory localization with a noise masker”, J. Acoust. Soc. Am. 126 (6), Dec. 2009, pp. 3199-3208. [cited by applicant]
Best et al., “The role of high frequencies in speech localization”, J. Acoust. Soc. Am. 118 (1), Jul. 2005, pp. 353-363. [cited by applicant]
Berlin et al., “Superior Ultra-Audiometric Hearing: A New Type of Hearing Loss Which Correlates Highly With Unusually Good Speech in the Profoundly Deaf”, vol. 86, Jan.-Feb. 1978, pp. ORL-111 to ORL-116. [cited by applicant]
Bench et al., “The BKB (Banford-Kowal-Bench) Sentence Lists for Partially-Hearing Children”, British Journal of Audiology, 1979, 13, pp. 108-112. [cited by applicant]
Badri et al., “Auditory filter shapes and high-frequency hearing in adults who have impaired speech in noise performance despite clinically normal audiograms”, J. Acoust. Soc. Am. 129 (2), Feb. 2011, pp. 852-863. [cited by applicant]
Arbogast et al., “Achieved Gain and Subjective Outcomes for a Wide-Bandwidth Contact Hearing Aid Fitted Using CAM2” Ear & Hearing, vol. 40, No. 3, 2018, pp. 741-756. [cited by applicant]
Wilson et al., “Perception of head orientation”, Vision Research, 40, 2000, pp. 459-472. [cited by applicant]
Werker et al., “Cross-Language Speech Perception: Evidence for Perceptual Reorganization During the First Year of Life”, Infant Behavior and Development 7, 1984, pp. 49-63. [cited by applicant]
Vitela et al., “Phoneme categorization relying solely on high-frequency energy”, J. Acoust. Soc. Am. 137 (1), Jan. 2015, pp. EL65-EL70. [cited by applicant]
Theunissen et al., “Neural processing of natural sounds”, Nature Reviews, vol. 15, Jun. 2014, pp. 355-366. [cited by applicant]
Strelcyk et al., “Effects of interferer facing orientation on speech perception by normal-hearing and hearing-impaired listeners”, J. Acoust. Soc. Am. vol. 135, No. 3, Mar. 2014, pp. 1419-1432. [cited by applicant]
Stelmachowicz et al., “The effect of stimulus bandwidth on auditory skills in normal-hearing and hearing-impaired children”, Ear Hear, Aug. 2007, 28(4): 483-494. [cited by applicant]
Stelmachowicz et al., “High-frequency audiometry: Test reliability and procedural considerations”, J. Acoust. Soc. Am., vol. 85, No. 2, Feb. 1989, pp. 879-887. [cited by applicant]
Shamma et al., “Temporal coherence and attention in auditory scene analysis”, Trends in Neurosciences, Mar. 2011, vol. 34, No. 3, pp. 114-123. [cited by applicant]
Lord Rayleigh, O.M., Pres. R.S., (1908) “XVIII. Acoustical Notes.-VIII”, The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, Apr. 21, 2009, 16:92, pp. 235-246. [cited by applicant]
Lippmann, “Accurate Consonant Perception Without Mid-Frequency Speech Energy”, IEEE Transactions on Speech and Audio Processing, vol. 4, No. 1, Jan. 1996, pp. 66-69. [cited by applicant]
Jongman et al., “Acoustic characteristics of English fricatives”, J. Acoust. Soc. Am., vol. 108, No. 3, Pt. 1, Sep. 2000, pp. 1252-1263. [cited by applicant]
Halkosaari et al., “Directivity of Artificial and Human Speech”, J. Audio Eng. Soc., vol. 53, No. 7/8, Jul./Aug. 2005, pp. 620-631. [cited by applicant]
Fullgrabe et al., “Preliminary evaluation of a method for fitting hearing aids with extended bandwidth”, International Journal of Audiology, vol. 49, No. 10, Jul. 30, 2010, pp. 741-753. [cited by applicant]
Flanagan, “Analog Measurements of Sound Radiation from the Mouth”, The Journal of the Acoustical Society of America, vol. 32, No. 12, Dec. 1960, pp. 1613-1620. [cited by applicant]
Dunn et al., “Exploration of Pressure Field Around the Human Head During Speech”, The Journal of the Acoustical Society of America, vol. 10, Jan. 1939, pp. 184-199. [cited by applicant]
Chu et al., “Detailed directivity of sound fields around human talkers”, NRC Publications Archive, IRC-RR-104, Dec. 2002, 49 pages. [cited by applicant]
Cabrera et al., “Long-Term Horizontal Vocal Directivity of Opera Singers: Effects of Singing Projection and Acoustic Environment”, Journal of Voice, vol. 25, No. 6, 2011, pp. e291-e303. [cited by applicant]
Bozzoli et al., “Balloons of Directivity of Real and Artificial Mouth Used in Determining Speech Transmission Index”, AES 118th Convention, Barcelona, Spain, May 28-31, 2005, pp. 1-5. [cited by applicant]
Best et al., “The role of high frequencies in speech localization”, J. Acoust. Soc. Am. vol. 118, No. 1, Jul. 2005, pp. 353-363. [cited by applicant]
Badri et al., “Auditory filter shapes and high-frequency hearing in adults who have impaired speech in noise performance despite clinically normal audiograms”, J. Acoust. Soc. Am., vol. 129, No. 2, Feb. 2011, pp. 852-86… [cited by applicant]
Stelmachowicz et al., “Effect of stimulus bandwidth on the perception of /s/ in normal- and hearing-impaired children and adults”, J. Acoust. Soc. Am., vol. 110, No. 4, Oct. 2001, pp. 2183-2190. [cited by applicant]
Spitzer et al., “Acoustic cues to lexical segmentation: A study of resynthesized speech”, J. Acoust. Soc. Am., vol. 122, No. 6, Dec. 2007, pp. 3678-3687. [cited by applicant]
Lord Rayleigh O.M. Pres.R.S., (1908) XVIII. Acoustical notes—VIII, the London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 16:92, 235-246. [cited by applicant]
Pittman, “Short-Term Word-Learning Rate in Children With Normal Hearing and Children With Hearing Loss in Limited and Extended High-Frequency Bandwidths”, Journal of Speech, Language, and Hearing Research, vol. 51, Jun.… [cited by applicant]
Moreno et al., “Human head directivity and speech emission: a new approach”, vol. 1, 1978, pp. 78-84. [cited by applicant]
Moore et al., “Effect of spatial separation, extended bandwidth, and compression speed on intelligibility in a competing-speech task”, J. Acoust. Soc. Am., vol. 128, No. 1, Jul. 2010, pp. 360-371. [cited by applicant]
Monson et al., “Detection of high-frequency energy changes in sustained vowels produced by singers”, J. Acoust. Soc. Am., vol. 129, No. 4, Apr. 2011, pp. 2263-2268. [cited by applicant]
Monson, “High-Frequency Energy in Singing and Speech”, 2011,161 pages. [cited by applicant]
McKendree, “Directivity indices of Human Talkers in English Speech”, Immission: Effects of Noise, Jul. 1986, 6 pages. [cited by applicant]
Marshall et al., “The Directivity and Auditory Impressions of Singers”, Acustica, vol. 58, 1985, pp. 130-140. [cited by applicant]
Maniwa et al., “Acoustic characteristics of clearly spoken English fricatives”, J. Acoust. Soc. Am., vol. 125, No. 6, Jun. 2009, pp. 3962-3973. [cited by applicant]
Monson et al., “The perceptual significance of high-frequency energy in the human voice”, frontiers in Psychology, Jun. 2014, vol. 5, Article 587, pp. 1 to 11. [cited by applicant]
Monson et al., “Analysis of high-frequency energy in ling-term avarage spectra of singing, speech, and voiceless fricatives”, J. Acoust. Soc. Am. 132(3), Sep. 2012, pp. 1754-1764. [cited by applicant]
Monson et al., “Horizontal directivity of low-and high-frequency energy in speech and singing”, Special Issue: Fish Bioacoustics: Hearing and Sound Communication, vol. 132, No. 1, Jul. 2012, pp. 433-441. [cited by applicant]
Monson et al., “The maximum audible low-pass cutoff frequency for speech”, J. Acoust. Soc. Am. 146(6), Dec. 2019, pp. EL496-EL501. [cited by applicant]
Liberman et al., “Toward a Differential Diagnosis of Hidden Hearing Loss in Humans”, PLOS One, Sep. 12, 2016, pp. 1-15. [cited by applicant]
Hunter et al., “Extended high frequency hearing and speech perception implications in adults and children”, Hearing Research 397 (2020), pp. 1-20. [cited by applicant]
Gamer et al., “Are You Looking at Me? Measuring the Cone of Gaze”, Journal of Experimental Psychology: Human Perception and Performance, 2007, vol. 33, No. 3, pp. 705-715. [cited by applicant]
Frischen et al., “Gaze Cueing of Attention: Visual Attention, Social Cognition, and Individual Differences”, Psychological Bulletin, 2007, vol. 133, No. 4, pp. 694-724. [cited by applicant]
American Auditory Society Scientific and Technology Meeting, Feb. 28-Mar. 2, 2019, 85 pages. [cited by applicant]
Chu et al., “Detailed directivity of sound fields around human talkers”, NRC Publications Archive, Dec. 2002, IRC-RR-104, 49 pages. [cited by applicant]
Ching et al., “Speech recognition of hearing-impaired listeners: Predictions from audibility and the limited role of high-frequency amplification”, The Journal of the Acoustical Society of America, 103 (2), Feb. 1998, p… [cited by applicant]
Bench et al., “The Bkb (Bamford-Kowal-Bench) Sentence Lists for Partially-Hearing Children”, British Journal of Audiology, 1979, 13, pp. 108-112. [cited by applicant]
Barbee et al., “Effectiveness of Auditory Measures for Detecting Hidden Hearing Loss and/or Cochlear Synaptopathy: A Systematic Review”, Seminars in Hearing, vol. 39, No. 2, 2018, pp. 172-209. [cited by applicant]
Apoux et al., “Relative importance of temporal information in various frequency regions for consonant identification in quiet and in noise”, J. Acoust. Soc. Am. 116 (3), Sep. 2004, pp. 1671-1680. [cited by applicant]
Alexander et al., “Acoustic and perceptual effects of amplitude and frequency compression on high-frequency speech”, J. Acoust. Soc. Am. 142 (2), Aug. 2017, pp. 908-923. [cited by applicant]
Yeend et al., “Working Memory and Extended High-Frequency Hearing in Adults: Diagnostic Predictors of Speech-in-Noise Perception”, Ear & Hearing, vol. 40, No. 3, 2018, pp. 458-467. [cited by applicant]
Van Eeckhoutte et al., “Speech recognition, loudness, and preference with extended bandwidth hearing aids for adjult hearing aid users”, International Journal of Audiology, vol. 59, No. 10, 2020, pp. 780-791. [cited by applicant]
Strelcyk et al., “Effects of interferer facing orientation on speech perception by normal-hearing and hearing-impaired listeners”, J. Acoust. Soc. Am. 135 (3), Mar. 2014, pp. 1419-1432. [cited by applicant]
Smith et al., “Investigating peripheral sources of speech-in-noise variability in listeners with normal audiograms”, Hearing Research 371, 2019, pp. 66-74. [cited by applicant]
Seeto et al., “Investigation of Extended Bandwidth Hearing Aid Amplification on Speech Intelligibility and Sound Quality in Adults with Mild-to-Moderate Hearing Loss”, Journal of the American Academy of Audiology, vol. … [cited by applicant]
Prendergast et al., “Effects of Age and Noise Exposure on Proxy Measures of Cochlear Synaptopathy”, Trends in Hearing, vol. 23, 2019, pp. 1-16. [cited by applicant]
Zadeh et al., “Extended high-frequency hearing enhances speech perception in noise”, PNAS, Nov. 19, 2019, vol. 116, No. 47, pp. 23753-23759. [cited by applicant]
Moore et al., “Perceived naturalness of spectrally distorted speech and music”, J. Acoust. Soc. Am. 114 (1), Jul. 2003, pp. 408-419. [cited by applicant]
Moore et al., “Frequency difference limens at high frequencies: Evidence for a transition from a temporal to a place code”, J. Acoust. Soc. Am. 132 (3), Sep. 2012, pp. 1542-1547. [cited by applicant]
Monson et al., “Ecological cocktail party listening reveals the utility of extended high-frequency hearing”, Hearing Research 381, 2019, pp. 1-7. [cited by applicant]
Monson et al., “Detection of high-frequency energy level changes in speech and singing”, J. Acoust. Soc. Am. 135 (1), Jan. 2014, pp. 400-406. [cited by applicant]
Prendergast et al., “Effects of noise exposure on young adults with normal audiograms II: Behavioral measures”, Hearing Research 356, 2017, pp. 74-86. [cited by applicant]
Plomp et al., “Effect of the Orientation of the Speaker's Head and the Azimuth of a Noise Source on the Speech-Reception Threshold for Sentences”, Acustica, vol. 48, 1981, pp. 325-328. [cited by applicant]
Neuhoff, “Twist and Shout: Audible Facing Angles and Dynamic Rotation”, Ecological Psychology, 2003, 15(4), 335-351. [cited by applicant]
Moore et al., “Benefits of Extended High-Frequency Audiometry for Everyone”, The Hearing Journal, Mar. 2017, pp. 50-55. [cited by applicant]
Moore et al., “Effect of spatial separation, extended bandwidth, and compression speed on intelligibility in a competing-speech task”, J. Acoust. Soc. Am. 128 (1), Jul. 2010, pp. 360-371. [cited by applicant]
Moore et al., “Spectro-Temporal Characteristics of Speech at High Frequencies, and the Potential for Restoration of Audibility to People with Mild-to-Moderate hearing Loss”, Ear & Hearing, 2008, vol. 29, No. 6, pp. 907-… [cited by applicant]
Moore et al., “Preliminary comparison of bone-anchored hearing instruments and a dental device as treatments for unilateral hearing loss”, International Journal of Audiology, Jul. 17, 2013, pp. 678-686. [cited by applicant]
Monson et al., “The perceptual significance of high-frequency energy in the human voice”, Frontiers in Psychology, Jun. 2014, vol. 5, Article 587, 11 pages. [cited by applicant]
Monson et al., “Analysis of high-frequency energy in long-term average spectra of singing, speech, and voiceless fricatives”, J. Acoust. Soc. Am. 132 (3), Sep. 2012, pp. 1754-1764. [cited by applicant]
McCreery et al., “The Influence of Audibility on Speech Recognition With Nonlinear Frequency Compression for Children and Adults With Hearing Loss”, Ear & Hearing, vol. 35, No. 4, 2014, pp. 440-447. [cited by applicant]
McCreery et al., “The Effects of Limited Bandwidth and Noise on Verbal Processing Time and Word Recall in Normal-Hearing Children”, Ear and Hearing, vol. 34, No. 5, 2013, pp. 585-591. [cited by applicant]
Crouzet et al., “On the Various Influences of Envelope Information on the Perception of Speech in Adverse Conditions: An Analysis of Between-Channel Envelope Correlation”, Human and Machine Perception Research Centre, M… [cited by applicant]
Arbogast et al., “Achieved Gain and Subjective Ourcomes for a Wide-Bandwidth Contact Hearing Aid Fitted Using CAM2”, Ear & Hearing, vol. 40, No. 3, 2019, pp. 741-756. [cited by applicant]
Jers, “Directivity Measurements of Adjacent Singers in a Choir”, 19th International Congress on Acoustics—ICA2007Madrid, 2007, 5 pages. [cited by applicant]
Crandall et al., “Analysis of the Energy Distribution in Speech”, Bell System Technical Journal, vol. 19, No. 3, 1922, pp. 221-232. [cited by applicant]