IP Library Granted Patent US 12,412,595
Granted Patent B2
US 12,412,595 · App. 17/915,074 · Granted Sep 9, 2025

Automatic leveling of speech content

Inventors: Chunghsin Yeh (Barcelona, ES); Giulio Cengarle (Barcelona, ES); Mark David de Burgh (San Francisco, CA)
Assignees: Dolby Laboratories Licensing Corporation; Dolby International AB
G10L21/0364G10L17/20G10L21/028G10L21/034G10L25/21G10L25/30G10L25/84G10L2025/786
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,595
App. No.
17/915,074
Granted
Sep 9, 2025
Kind
B2
Abstract

Embodiments are disclosed for automatic leveling of speech content. In an embodiment, a method comprises: receiving, using one or more processors, frames of an audio recording including speech and non-speech content; for each frame: determining, using the one or more processors, a speech probability; analyzing, using the one or more processors, a perceptual loudness of the frame; obtaining, using the one or more processors, a target loudness range for the frame; computing, using the one or more processors, gains to apply to the frame based on the target loudness range and the perceptual loudness analysis, where the gains include dynamic gains that change frame-by-frame and that are scaled based on the speech probability; and applying the gains to the frame so that a resulting loudness range of the speech content in the audio recording fits within the target loudness range.

Claims (39)

1. A method comprising:

receiving, using one or more processors, frames of an audio recording including speech and non-speech, where the speech includes speech from multiple speakers with different signal-to-noise ratios (SNRs);

detecting speech frames in the speech content and their respective speech probabilities;

segmenting the speech frames into separate streams according to the identities of the multiple speakers, wherein segmenting is based on time indices to indicate where each of the multiple speakers is actively dominant in the audio recording;

for each frame in each one of the separate streams:

analyzing, using the one or more processors, a perceptual loudness of the frame, the analyzing including determining an integrated loudness and a momentary loudness;

obtaining, using the one or more processors, a target loudness range for the frame;

in response to a first condition being met,

setting a dynamic gain for the frame as a function of the integrated loudness, the target loudness range, the momentary loudness and a non-linear function of the speech probability for the frame, wherein the non-linear function of the speech probability is configured to avoid boosting non-speech parts of the audio;

in response to a second condition being met, setting the dynamic gain for the frame as a function of the integrated loudness, the target loudness range, and the momentary loudness for the frame, wherein the second condition is different than the first condition;

and otherwise, when both the first and second conditions are not met, setting the dynamic range to zero; and

applying, using the one or more processors, the dynamic gains to the frame so that a resulting loudness range of the speech content in the audio recording fits within the target loudness range.

2. The method of claim 1 , further comprising:

computing, using the one or more processors, a static gain that is applied to all the frames.

3. The method of claim 2 , where the static gain is the difference between an integrated loudness and a target loudness.

4. The method of claim 2 , where the gain applied to each frame is the sum of the static gain and the dynamic gain for the frame.

5. The method of claim 1 , wherein the dynamic gains are computed as a continuous function of a distance between the perceptual loudness of each frame and an integrated loudness.

6. The method of claim 1 , wherein the dynamic gains of the frames within a desired loudness range with respect to the integrated loudness are unity, and the dynamic gains applied to frames outside the desired loudness range are computed as the difference between the frame's loudness value and the nearest boundary of the desired loudness range.

7. The method of claim 1 , where the dynamic gains are multiplied by a coefficient between 0.0 and 1.0.

8. The method of claim 1 , where the speech probability is computed by a neural network.

9. The method of claim 1 , where the speech probability is a function of a broadband energy level of each frame.

10. The method of claim 1 , further comprising:

estimating a signal-to-noise ratio (SNR); and

modifying the speech probability based at least in part on the estimated SNR.

11. The method of claim 10 , where the speech probability is determined by a voice activity detector (VAD), and the method further comprises:

adjusting a sensitivity of the VAD to increase discrimination between speech and non-speech when the estimated SNR indicates the speech content is clean.

12. The method of claim 1 , further comprising:

estimating a signal-to-noise ratio (SNR); and

adjusting the target loudness based on the estimated SNR so that a small dynamic range is only achieved when the speech content is clean.

13. The method of claim 10 , where the dynamic gains are multiplied by a coefficient between 0 and 1, and the coefficient is a function of the SNR.

14. The method of claim 1 , where the speech probability can be modified through a sigmoid function, and wherein a parameter of the sigmoid function is either manually fixed or automatically adapted based on the estimated SNR of the speech content.

15. The method of claim 1 , where the speech probability is a function of the energy level of each frame in a specific frequency band.

16. The method of claim 1 , where the perceptual loudness of each frame is computed at recording time and stored.

17. The method of claim 1 , where the speech probability is computed at recording time and stored.

18. The method of claim 12 , wherein the estimated SNR is used to control a parameter of the non-linear function of the speech probability.

19. A system comprising:

one or more processors; and

a non-transitory computer-readable medium storing instructions that, upon execution by the one or more processors, cause the one or more processors to perform operations of the method of claim 1 .

20. A non-transitory, computer-readable medium storing instructions that, upon execution by one or more processors, cause the one or more processors to perform operations of the method of claim 1 .

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR PREVIOUSLY RECORDED AT REEL: 062176 FRAME: 0901. ASSIGNOR(S) HEREBY CONFIRMS THE NEW ASSIGNMENT. Recorded Mar 6, 2023
From: YEH, CHUNGHSIN; CENGARLE, GIULIO; DE BURGH, MARK DAVID
To: DOLBY LABORATORIES LICENSING CORPORATION; DOLBY INTERNATIONAL AB
Reel/Frame 062961/0691 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2022
From: YEH, CHUNGHSIN; CENGARLE, GIULIO
To: DOLBY LABORATORIES LICENSING CORPORATION; DOLBY INTERNATIONAL AB
Reel/Frame 062176/0901 →
Priority Claims (1)
ES ES202000051 · Mar 27, 2020 · national
Continuity (3)
Provisional Application 63126889 · Dec 17, 2020
Provisional Application 63032158 · May 29, 2020
Related Publication 20230162754A1 · May 25, 2023
References Cited (61)
US 6415253B1 · Johnson · 2002 [cited by applicant]
US 6651040B1 · Bakis · 2003 [cited by applicant]
US 7742746B2 · Xiang · 2010 [cited by applicant]
US 7848531B1 · Vickers · 2010 [cited by applicant]
US 8121835B2 · Archibald · 2012 [cited by applicant]
US 8407045B2 · Erell · 2013 [cited by applicant]
US 8428949B2 · Neoran · 2013 [cited by applicant]
US 9350312B1 · Wishnick · 2016 [cited by applicant]
US 9495970B2 · Dickins · 2016 [cited by applicant]
US 9509267B2 · Patwardhan · 2016 [cited by applicant]
US 9559651B2 · Baumgarte · 2017 [cited by applicant]
US 9841941B2 · Riedmiller · 2017 [cited by applicant]
US 9842596B2 · Riedmiller · 2017 [cited by applicant]
US 9842605B2 · Lu · 2017 [cited by applicant]
US 10147436B2 · Riedmiller · 2018 [cited by applicant]
US 10355658B1 · Yang · 2019 [cited by examiner]
US 10448161B2 · Xiang · 2019 [cited by applicant]
US 10546590B2 · Sharma · 2020 [cited by applicant]
US 11164592B1 · Wu · 2021 [cited by examiner]
US 20060184366A1 · Hidaka · 2006 [cited by examiner]
US 20090009251A1 · Spielbauer · 2009 [cited by examiner]
US 20090257335A1 · Lin · 2009 [cited by examiner]
US 20100017203A1 · Archibald · 2010 [cited by applicant]
US 20100094625A1 · Mohammad · 2010 [cited by examiner]
US 20110071825A1 · Emori · 2011 [cited by examiner]
US 20120123769A1 · Urata · 2012 [cited by applicant]
US 20120221328A1 · Muesch · 2012 [cited by examiner]
US 20120263317A1 · Shin · 2012 [cited by applicant]
US 20130304464A1 · Wang · 2013 [cited by examiner]
US 20140074467A1 · Ziv · 2014 [cited by examiner]
US 20150332685A1 · Bleidt · 2015 [cited by examiner]
US 20160042746A1 · Fujieda · 2016 [cited by examiner]
US 20160322067A1 · Sehlstedt · 2016 [cited by examiner]
US 20160358618A1 · Chen · 2016 [cited by examiner]
US 20170345439A1 · Jensen · 2017 [cited by examiner]
US 20180234069A1 · De Burgh · 2018 [cited by examiner]
US 20190318758A1 · Ma · 2019 [cited by examiner]
US 20200112294A1 · Mabande · 2020 [cited by examiner]
US 20200176012A1 · Herbig · 2020 [cited by examiner]
US 20210074282A1 · Borgstrom · 2021 [cited by examiner]
US 20210152934A1 · Tu · 2021 [cited by examiner]
US 20210352408A1 · Honma · 2021 [cited by examiner]
US 20210367574A1 · Chon · 2021 [cited by examiner]
US 20220148571A1 · Wang · 2022 [cited by examiner]
CN 101647059A · 2010 [cited by applicant]
CN 102804261A · 2012 [cited by applicant]
CN 105190750A · 2015 [cited by applicant]
CN 107994879B · 2018 [cited by applicant]
EP 3340657A1 · 2018 [cited by examiner]
KR 101704926B1 · 2017 [cited by examiner]
WO 2018188812A1 · 2018 [cited by applicant]
EBU-Recommendation, R. (2011). Loudness normalisation and permitted maximum level of audio signals. European Broadcasting Union. (Year: 2011). [cited by examiner]
Molinder, H. (2016). Adaptive Normalisation of Programme Loudness in Audiovisual Broadcasts. (Year: 2016). [cited by examiner]
Plapous, C., Marro, C., & Scalart, P. (2006). Improved signal-to-noise ratio estimation for speech enhancement. IEEE transactions on audio, speech, and language processing, 14(6), 2098-2108. (Year: 2006). [cited by examiner]
Zheng, Y., & Lin, Z. (2003). Recursive adaptive algorithms for fast and rapidly time-varying systems. IEEE Transactions on Circuits and Systems II: Analog and Digital Signal Processing, 50(9), 602-614. (Year: 2003). [cited by examiner]
https://www.waves.com/plugins/vocal-rider?gclid=EAlalQobChMlirgk_vu5wIV0_ZRCh1eDQErEAAYASAAEgKjlvD_BwE#achieving-perfect-vocallevels-with-vocal-rider. [cited by applicant]
https://www.izotope.com/en/products/rx/features/leveler.html. [cited by applicant]
ITU-R BS. 1770-4, “Algorithms to Measure Audio Programme Loudness and True-Peak Audio Level” Oct. 2015, pp. 1-25. [cited by applicant]
Maddams, J. A. et al “An Autonomous Method for Multi-Track Dynamic Range Compression” Proc. of the 15th Int. Conference on Digital Audio Effects, Sep. 17-21, 2012. [cited by applicant]
Moore, B. C.J. et al.“A Model for the Prediction of Thresholds, Loudness, and Partial Loudness” 1997, AES. [cited by applicant]
Zwicker, E. et al “Program for Calculating Loudness according to DIN 45631” JST, 1991, vol. 12 Issue 1 pp. 39-42. [cited by applicant]