IP Library Granted Patent US 12,354,617
Granted Patent B2
US 12,354,617 · App. 17/669,615 · Granted Jul 8, 2025

Context-aware voice intelligibility enhancement

Inventors: Daekyoung Noh (Calabasas, CA); Pavel Chubarev (Calabasas, CA); Xiaoyu Guo (Calabasas, CA)
Assignee: DTS, Inc.
G10L21/0232G10L15/18G10L15/22G10L21/038G10L2021/02082G10L2021/02163
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,617
App. No.
17/669,615
Granted
Jul 8, 2025
Kind
B2
Abstract

A method comprises: detecting noise in an environment with a microphone to produce a noise signal; receiving a voice signal to be played into the environment through a loudspeaker; performing multiband correction of the noise signal based on a microphone transfer function of the microphone, to produce a corrected noise signal; performing multiband correction of the voice signal based on a loudspeaker transfer function of the loudspeaker to produce a corrected voice signal; and computing multiband voice intelligibility results based on the corrected noise signal and the corrected voice signal.

Claims (70)

1. A method comprising:

detecting noise in an environment with a microphone to produce a noise signal;

receiving a voice signal to be played into the environment through a loudspeaker;

determining a frequency analysis region for a multiband voice intelligibility computation based on a relationship between a microphone transfer function of the microphone and a loudspeaker transfer function of the loudspeaker;

performing multiband correction of the noise signal based on the microphone transfer function, to produce a corrected noise signal;

performing multiband correction of the voice signal based on the loudspeaker transfer function, to produce a corrected voice signal;

computing a global speech-to-noise ratio of (i) voice power based on the voice signal across a voice analysis bands limited to an overlap passband to (ii) noise power based on the noise signal across a microphone passband; and

computing multiband voice intelligibility results over the frequency analysis region based on the corrected noise signal and the corrected voice signal, wherein the multiband voice intelligibility results include long segments analyzed by a long-term voice and noise profiling obtained based on an accumulation of short-term voice intelligibility results over time with a sliding window.

2. The method of claim 1 , further comprising:

enhancing intelligibility of the voice signal using the multiband voice intelligibility results.

3. The method of claim 1 , wherein:

determining includes determining, as the frequency analysis region, the overlap passband over which the microphone passband of the microphone transfer function and a loudspeaker passband of the loudspeaker transfer function overlap; and

computing includes computing per-band voice intelligibility values across voice analysis bands limited to the overlap passband.

4. The method of claim 3 , wherein determining includes:

identifying start frequencies and stop frequencies that define the microphone passband and the loudspeaker passband, respectively; and

computing the overlap passband as a passband that extends from a maximum of the start frequencies to a minimum of the stop frequencies.

5. The method of claim 1 , wherein:

computing the multiband voice intelligibility results includes computing per-band voice intelligibility values and a global voice-to-noise ratio.

6. The method of claim 1 , wherein computing the multiband voice intelligibility results includes:

performing multiband voice intelligibility analysis based on short/medium length segments of the voice signal and the noise signal, to produce short term voice intelligibility results; and

performing multiband voice intelligibility analysis based on long segments of the voice signal and the noise signal that are longer than the short/medium length segments of the voice signal and the noise signal, to produce long term voice intelligibility results.

7. The method of claim 1 , further comprising:

prior to performing the multiband correction of the noise signal, performing digital-to-acoustic level conversion of the noise signal based on a sensitivity of the microphone; and

prior to performing the multiband correction of the voice signal, performing digital-to-acoustic level conversion of the voice signal based on a sensitivity of the loudspeaker.

8. A method comprising:

detecting noise in an environment with a microphone to produce a noise signal;

receiving a voice signal to be played into the environment through a loudspeaker;

determining a frequency analysis region for a multiband voice intelligibility computation based on a relationship between a microphone transfer function of the microphone and a loudspeaker transfer function of the loudspeaker;

performing multiband correction of the noise signal based on the microphone transfer function, to produce a corrected noise signal;

performing multiband correction of the voice signal based on the loudspeaker transfer function, to produce a corrected voice signal;

computing multiband voice intelligibility results over the frequency analysis region based on the corrected noise signal and the corrected voice signal;

determining whether a start frequency of a loudspeaker passband is greater than a start frequency of a microphone passband; and

when the start frequency of the loudspeaker passband is greater, attenuating the voice signal in bands below the start frequency of the microphone passband.

9. An apparatus comprising:

a microphone to detect noise in an environment, to produce a noise signal;

a loudspeaker to play a voice signal into the environment; and

a controller coupled to the microphone and the loudspeaker and configured to perform:

multiband correction of the noise signal based on a microphone transfer function of the microphone, to produce a corrected noise signal;

multiband correction of the voice signal based on a loudspeaker transfer function of the loudspeaker to produce a corrected voice signal;

computing a global speech-to-noise ratio of (i) voice power based on the voice signal across a voice analysis bands limited to an overlap passband to (ii) noise power based on the noise signal across a microphone passband;

computing multiband voice intelligibility results based on the corrected noise signal and the corrected voice signal, wherein the multiband voice intelligibility results include long segments analyzed by a long-term voice and noise profiling obtained based on an accumulation of short-term voice intelligibility results over time with a sliding window;

computing multiband gain values based on the multiband voice intelligibility results; and

enhancing the voice signal based on the multiband gain values.

10. The apparatus of claim 9 , wherein the controller is further configured to perform:

enhancing intelligibility of the voice signal using the multiband voice intelligibility results.

11. The apparatus of claim 9 , wherein the controller is further configured to perform:

determining the overlap passband over which the microphone passband of the microphone transfer function and a loudspeaker passband of the loudspeaker transfer function overlap,

wherein the controller is configured to perform computing by computing per-band voice intelligibility values across voice analysis bands limited to the overlap passband.

12. The apparatus of claim 11 , wherein the controller is further configured to perform:

determining whether a start frequency of the loudspeaker passband is greater than a start frequency of the microphone passband; and

when the start frequency of the loudspeaker passband is greater, attenuating the voice signal in bands below the start frequency of the microphone passband.

13. The apparatus of claim 9 , wherein:

the controller is configured to perform computing the multiband voice intelligibility results by computing per-band voice intelligibility values and a global speech-to-noise ratio.

14. The apparatus of claim 9 , wherein computing the multiband voice intelligibility results includes:

performing multiband voice intelligibility analysis on short/medium length segments of the corrected voice signal and the corrected noise signal, to produce short term voice intelligibility results; and

performing multiband voice intelligibility analysis on long segments of the corrected voice signal and the corrected noise signal that are longer than the short/medium length segments of the corrected voice signal and the corrected noise signal, to produce long term voice intelligibility results.

15. The apparatus of claim 9 , further comprising:

prior to multiband correction of the noise signal, performing digital-to-acoustic level conversion of the noise signal based on a sensitivity of the microphone; and

prior to multiband correction of the voice signal, performing digital-to-acoustic level conversion of the voice signal.

16. A non-transitory computer readable medium encoded with instructions that, when executed by a processor, cause the processor to perform:

receiving, from a microphone, a noise signal representative of noise in an environment;

receiving a voice signal to be played into the environment through a loudspeaker;

digital-to-acoustic level conversion of the noise signal, and multiband correction of the noise signal based on a microphone transfer function, to produce a corrected noise signal;

digital-to-acoustic level conversion of the voice signal, and multi band correction of the voice signal based on a loudspeaker transfer function, to produce a corrected voice signal;

computing a global speech-to-noise ratio of (i) voice power based on the voice signal across a voice analysis bands limited to an overlap passband to (ii) noise power based on the noise signal across a microphone passband; and

computing, based on the corrected noise signal and the corrected voice signal, multiband voice intelligibility results, wherein the multiband voice intelligibility results include long segments analyzed by a long-term voice and noise profiling obtained based on an accumulation of short-term voice intelligibility results over time with a sliding window.

17. The non-transitory computer readable medium of claim 16 , wherein the instructions to cause the processor to perform computing include instructions to cause the processor to perform a speech intelligibility index (SII) analysis of the corrected noise signal and the corrected voice signal across voice analysis bands.

18. The non-transitory computer readable medium of claim 16 , further comprising instructions to cause the processor to perform:

determining the overlap passband over which the microphone passband of the microphone transfer function and a loudspeaker passband of the loudspeaker transfer function overlap,

wherein the instructions to cause the processor to perform the computing include instructions to cause the processor to perform computing a per-band voice intelligibility values across voice analysis bands limited to the overlap passband.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2023
From: NOH, DAEKYOUNG; CHUBAREV, PAVEL; GUO, XIAOYU
To: DTS, INC.
Reel/Frame 063869/0735 →
Continuity (3)
Continuation PCTUS2020049933 · Sep 9, 2020
Provisional Application 62898977 · Sep 11, 2019
Related Publication 20220165287A1 · May 26, 2022
References Cited (50)
US 6249237B1 · Prater · 2001 [cited by applicant]
US 6993480B1 · Klayman · 2006 [cited by applicant]
US 8204742B2 · Yang et al. · 2012 [cited by applicant]
US 8386247B2 · Yang et al. · 2013 [cited by applicant]
US 9117455B2 · Tracey et al. · 2015 [cited by applicant]
US 9947337B1 · Wang · 2018 [cited by examiner]
US 20020075856A1 · LeBlanc · 2002 [cited by examiner]
US 20030112261A1 · Zhang · 2003 [cited by examiner]
US 20040057586A1 · Licht · 2004 [cited by examiner]
US 20060293882A1 · Giesbrecht · 2006 [cited by examiner]
US 20070112563A1 · Krantz et al. · 2007 [cited by applicant]
US 20090007596A1 · Goldstein · 2009 [cited by examiner]
US 20090010443A1 · Ahnert · 2009 [cited by examiner]
US 20090191817A1 · Gupta · 2009 [cited by examiner]
US 20090225980A1 · Schmidt · 2009 [cited by examiner]
US 20110066428A1 · Yang et al. · 2011 [cited by applicant]
US 20110176687A1 · Birkenes · 2011 [cited by examiner]
US 20120259625A1 · Yang et al. · 2012 [cited by applicant]
US 20130030800A1 · Tracey · 2013 [cited by examiner]
US 20130218560A1 · Hsiao · 2013 [cited by examiner]
US 20140056435A1 · Kjems · 2014 [cited by examiner]
US 20150011266A1 · Feldt · 2015 [cited by examiner]
US 20150019213A1 · Nongpiur · 2015 [cited by examiner]
US 20150226831A1 · Nakamura · 2015 [cited by examiner]
US 20160029142A1 · Isaac · 2016 [cited by examiner]
US 20160078879A1 · Lu · 2016 [cited by examiner]
US 20160260423A1 · Rosenkranz · 2016 [cited by examiner]
US 20160287180A1 · Ansari · 2016 [cited by examiner]
US 20170265010A1 · Kavalekalam · 2017 [cited by examiner]
US 20170345440A1 · Matsuo · 2017 [cited by examiner]
US 20180077507A1 · Bernal Castillo · 2018 [cited by examiner]
US 20180359560A1 · Defraene · 2018 [cited by examiner]
US 20190206420A1 · Kandade Rajan · 2019 [cited by examiner]
US 20190318757A1 · Chen · 2019 [cited by examiner]
US 20200020315A1 · Tachi · 2020 [cited by examiner]
US 20200194006A1 · Grancharov · 2020 [cited by examiner]
CN 105489224A · 2016 [cited by examiner]
CN 105702262A · 2016 [cited by examiner]
JP 4482247B2 · 2010 [cited by examiner]
KR 20110043699A · 2011 [cited by applicant]
WO 2010009414A1 · 2010 [cited by applicant]
WO 2011019339A1 · 2011 [cited by applicant]
Ansi, S3 22-1997. “Methods for calculation of the speech intelligibility index.” American National Standard Institute, pp. 1-28 (1997). [cited by applicant]
Choo, K., et al., “Enhancement of Voice Intelligibility for mobile speech communication in noisy environments”, Convention Paper 9810, Audio Engineering Society, Presented at the 143rd Convention Oct. 18-21, 2017, New Y… [cited by applicant]
Faiget, L., et al., “Speech Intelligibility Model Including Room and Loudspeaker Influences”, J. Acoust. Soc. Am., vol. 105 (6), pp. 3345-3354 (Jun. 1999). [cited by applicant]
Hu, Y, and P.C. Loizou, “A Perceptually Motivated Approach for Speech Enhancement”, IEEE Transactions on Speech and Audio Processing, vol. 11(5), pp. 457-465 (2003). [cited by applicant]
Postma Barteld NJ, et al., “Perceptive and Objective Evaluation of Calibrated Room Acoustic Simulation Auralizations”, J. Acoust. Soc. Am., vol. 140 (6), pp. 4326-4337 (Dec. 2016). [cited by applicant]
Rhebergen, K.S., and N.J. Versfeld, “A Speech Intelligibility Index-based approach to predict the speech reception threshold for sentences in fluctuating noise for normal-hearing listeners”, The Journal of the Acoustica… [cited by applicant]
Thomas, I.B, “The Influence of First and Second Formants on the Intelligibility of Clipped Speech”, Journal of the Audio Engineering Society, vol. 16(2), 182-185 (1968). [cited by applicant]
Search Report and Written Opinion from corresponding International Patent Application No. PCT/US2020/049933, mailed Jan. 26, 2021. [cited by applicant]