IP Library Granted Patent US 12,382,234
Granted Patent B2
US 12,382,234 · App. 18/008,431 · Granted Aug 5, 2025

Perceptual optimization of magnitude and phase for time-frequency and softmask source separation systems

Inventors: Aaron Steven Master (San Francisco, CA); Lie Lu (San Francisco, CA); Heiko Purnhagen (San Francisco, CA)
Assignees: Dolby Laboratories Licensing Corporation; DOLBY INTERNATIONAL AB
H04S7/30G10L21/0308G10L25/18H04S1/007H04S2400/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,382,234
App. No.
18/008,431
Granted
Aug 5, 2025
Kind
B2
Abstract

A method comprises: obtaining softmask values for frequency bins of time-frequency tiles representing an audio signal; reducing, or expanding and limiting, the softmask values; and applying the reduced, or expanded and limited, softmask values to the frequency bins to create a time-frequency representation of an estimated target source. An alternative method comprises, for each time-frequency tile: obtaining softmask values; applying the softmask values to the frequency bins to create a time-frequency domain representation of an estimated target source; obtaining a panning parameter and a source concentration estimates for the target source; determining, using the panning parameter estimate and the softmask values, a magnitude for the time-frequency representation of the estimated target source; determining, using the panning parameter estimate and the source phase concentration estimate, a phase for the time-frequency representation of the estimated target source; and combining the magnitude and the phase.

Claims (59)

1. A method comprising:

obtaining softmask values for frequency bins of time-frequency tiles representing an audio signal, the audio signal including a target source and one or more backgrounds;

reducing the softmask values; and

applying the reduced softmask values to the frequency bins to create a time-frequency representation of an estimated target source,

wherein reducing the softmask values comprises:

estimating a bulk reduction threshold, the bulk reduction threshold representing a balance point between softmask values that correlate with target dominant time-frequency tiles and softmask values that correlate with background dominant time-frequency tiles; and

multiplying each softmask value that falls below the bulk reduction threshold by a fractional value.

2. The method of claim 1 , further comprising, prior to obtaining the softmask values,

transforming, using one or more processors, one or more frames of a time domain audio signal into a time-frequency domain representation including the time-frequency tiles, wherein the time-frequency domain representation includes the target source and the one or more backgrounds, and wherein the frequency domain of the time-frequency domain representation includes the frequency bins grouped into a plurality of subbands.

3. The method of claim 2 , wherein the time domain audio signal is a multiple-channel audio signal, further comprising:

for each time-frequency tile:

calculating spatial parameters and a level for the time-frequency tile, and

obtaining the softmask values using the spatial parameters, the level and a subband information.

4. The method of claim 1 , further comprising:

setting to zero or near-zero the softmask values in the frequency bins that are outside a specified frequency range.

5. A method comprising:

obtaining softmask values for frequency bins of time-frequency tiles representing an audio signal, the audio signal including a target source and one or more backgrounds;

expanding and limiting the softmask values; and

applying the expanded and limited, softmask values to the frequency bins to create a time-frequency representation of an estimated target source,

wherein expanding and limiting the softmask values, further comprises:

adding a fixed expansion addition value to the softmask values;

multiplying the softmask values by an expansion multiplier constant; and

limiting any softmask values that are above 1.0 to 1.0.

6. The method of claim 5 , further comprising, prior to obtaining the softmask values,

transforming, using one or more processors, one or more frames of a time domain audio signal into a time-frequency domain representation including the time-frequency tiles, wherein the time-frequency domain representation includes the target source and the one or more backgrounds, and wherein the frequency domain of the time-frequency domain representation includes the frequency bins grouped into a plurality of subbands.

7. The method of claim 6 , wherein the time domain audio signal is a multiple-channel audio signal, further comprising:

for each time-frequency tile:

calculating spatial parameters and a level for the time-frequency tile, and obtaining the softmask values using the spatial parameters, the level and a subband information.

8. The method of claim 5 , further comprising:

setting to zero or near-zero the softmask values in the frequency bins that are outside a specified frequency range.

9. A method comprising:

obtaining softmask values for frequency bins of time-frequency tiles representing an audio signal, the audio signal including a target source and one or more backgrounds, wherein the time-frequency tiles represent a multiple channels audio signal and the frequency bins of the time-frequency tiles are organized into a plurality of subbands, the method further comprising, for each time-frequency tile:

obtaining softmask values for frequency bins of time-frequency tiles representing the multiple channels audio signal;

applying the softmask values to the frequency bins to create a time-frequency domain representation of an estimated target source; wherein the method further comprises:

obtaining a panning parameter estimate for the target source;

obtaining a source phase concentration estimate for the target source, wherein the source phase concentration estimate is obtained by estimating a statistical distribution of phase differences between the multiple channels in the time-frequency tiles for capturing a predetermined amount of audio energy of the target source;

determining, using the panning parameter estimate and the softmask values, a magnitude for the time-frequency domain representation of the estimated target source;

determining, using the panning parameter estimate and the source phase concentration estimate, a phase for the time-frequency domain representation of the estimated target source based; and

combining the magnitude and the phase to create a modified time-frequency domain representation of the estimated target source.

10. The method of claim 9 , further comprising, prior to obtaining the softmask values,

transforming, using one or more processors, one or more frames of a time domain audio signal into a time-frequency domain representation including the time-frequency tiles, wherein the time-frequency domain representation includes the target source and the one or more backgrounds, and wherein the frequency domain of the time-frequency domain representation includes the frequency bins grouped into the plurality of subbands.

11. The method of claim 10 , further comprising:

for each time-frequency tile:

calculating spatial parameters and a level for the time-frequency tile, and

obtaining the softmask values using the spatial parameters, the level and a subband information.

12. The method of claim 9 , wherein determining, using the panning parameter estimate and the source phase concentration estimate, a phase for the time-frequency domain representation of the estimated target source based, further comprises:

computing, using the panning parameter estimate, a first weight for a left channel phase and a second weight for a right channel phase;

computing a weighted average of the left and right channel phases using the first weight and the second weight, respectively; and

adjusting a phase parameter of the time-frequency tile for the time-frequency domain representation of the estimated target source to be the weighted average of the left and right channel phases.

13. The method of claim 9 , wherein determining, using the panning parameter estimate and the softmask values, a magnitude for the time-frequency domain representation of the estimated target source, further comprises:

computing a left channel ratio as a function of the panning parameter estimate;

computing a right channel ratio as a function the panning parameter estimate;

computing a left channel magnitude for the left channel based on a product of the left channel ratio, a softmask value and a level of the frequency bin; and

computing a right channel magnitude based on the product of the right channel ratio, the softmask value for the frequency bin and the level of the frequency bin.

14. The method of claim 9 , wherein estimating the statistical distribution of the phase differences between the multiple channels in the time-frequency tiles further comprises:

determining a peak value of the statistical distribution;

determining a phase difference corresponding to the peak value; and

determining a width of the statistical distribution around the peak value for capturing the amount of audio energy.

15. The method of claim 9 , wherein the predetermined amount of audio energy is at least eighty percent of a total energy in the statistical distribution of the phase differences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 9, 2023
From: MASTER, AARON STEVEN; LU, LIE; PURNHAGEN, HEIKO
To: DOLBY LABORATORIES LICENSING CORPORATION; DOLBY INTERNATIONAL AB
Reel/Frame 062934/0856 →
Priority Claims (1)
EP 20179450 · Jun 11, 2020 · regional
Continuity (2)
Provisional Application 63038052 · Jun 11, 2020
Related Publication 20230232176A1 · Jul 20, 2023
References Cited (84)
US 8180062B2 · Turku · 2012 [cited by applicant]
US 8238569B2 · Jeong · 2012 [cited by applicant]
US 8321214B2 · Chan et al. · 2012 [cited by applicant]
US 8379868B2 · Goodwin et al. · 2013 [cited by applicant]
US 8472631B2 · Klayman · 2013 [cited by applicant]
US 8509464B1 · Kato · 2013 [cited by applicant]
US 8600533B2 · Berchin · 2013 [cited by applicant]
US 9008329B1 · Mandel · 2015 [cited by applicant]
US 9053697B2 · Park · 2015 [cited by applicant]
US 9111526B2 · Visser · 2015 [cited by applicant]
US 9324337B2 · Brown · 2016 [cited by applicant]
US 9438992B2 · Every · 2016 [cited by applicant]
US 9699554B1 · Choi · 2017 [cited by applicant]
US 9881631B2 · Erdogan · 2018 [cited by applicant]
US 10043527B1 · Gurijala · 2018 [cited by applicant]
US 10075797B2 · Thompson · 2018 [cited by applicant]
US 10123134B2 · Jensen · 2018 [cited by applicant]
US 10192568B2 · Wang · 2019 [cited by applicant]
US 10321241B2 · Lunner · 2019 [cited by applicant]
US 10347271B2 · Nesta · 2019 [cited by applicant]
US 10373623B2 · Dittmar · 2019 [cited by applicant]
US 10430154B2 · Gillespie · 2019 [cited by applicant]
US 11115774B2 · Zhang · 2021 [cited by applicant]
US 11184709B2 · Purnhagen · 2021 [cited by applicant]
US 11190900B2 · McElveen · 2021 [cited by applicant]
US 20040062401A1 · Davis · 2004 [cited by applicant]
US 20070076902A1 · Master · 2007 [cited by applicant]
US 20090097670A1 · Jeong · 2009 [cited by examiner]
US 20090279715A1 · Jeong · 2009 [cited by examiner]
US 20100183158A1 · Haykin · 2010 [cited by applicant]
US 20150271620A1 · Lando et al. · 2015 [cited by applicant]
US 20150312663A1 · Traa · 2015 [cited by applicant]
US 20160071526A1 · Wingate · 2016 [cited by applicant]
US 20160203829A1 · Vishnubhotla · 2016 [cited by examiner]
US 20170061978A1 · Wang · 2017 [cited by applicant]
US 20170178664A1 · Wingate · 2017 [cited by applicant]
US 20170345433A1 · Dittmar · 2017 [cited by examiner]
US 20180088899A1 · Gillespie · 2018 [cited by applicant]
US 20180122689A1 · Adkisson · 2018 [cited by applicant]
US 20180299527A1 · Helwani · 2018 [cited by applicant]
US 20180308502A1 · Parekh · 2018 [cited by applicant]
US 20190043491A1 · Kupryjanow · 2019 [cited by applicant]
US 20190066713A1 · Mesgarani · 2019 [cited by applicant]
US 20190122689A1 · Jain · 2019 [cited by applicant]
US 20190132687A1 · Santos · 2019 [cited by applicant]
US 20190139562A1 · Fueg · 2019 [cited by applicant]
US 20190139563A1 · Chen · 2019 [cited by applicant]
US 20190164052A1 · Sung · 2019 [cited by applicant]
US 20190172476A1 · Wung · 2019 [cited by applicant]
US 20190180142A1 · Lim · 2019 [cited by applicant]
US 20190246203A1 · Elko · 2019 [cited by applicant]
US 20190251985A1 · Yu · 2019 [cited by applicant]
US 20220223144A1 · Sun · 2022 [cited by examiner]
US 20230215423A1 · Master · 2023 [cited by applicant]
US 20230232176A1 · Master · 2023 [cited by applicant]
US 20230245664A1 · Master · 2023 [cited by applicant]
US 20230245671A1 · Master · 2023 [cited by applicant]
AU 2004286507B2 · 2010 [cited by applicant]
AU 2012241166B2 · 2016 [cited by applicant]
CA 2649911C · 2013 [cited by applicant]
CA 2794946C · 2017 [cited by applicant]
CA 2983359C · 2019 [cited by applicant]
WO 2015024940A1 · 2015 [cited by applicant]
WO 2019106221A1 · 2019 [cited by applicant]
WO 2020232180A1 · 2020 [cited by applicant]
WO 2021252823A1 · 2021 [cited by applicant]
WO 2023172852A1 · 2023 [cited by applicant]
WO 2023192036A1 · 2023 [cited by applicant]
WO 2024167785A1 · 2024 [cited by applicant]
Aaron Master et al., Stereo Speech Enhancement Using Custom Mid-Side Signals and Monaural Processing ARIXIV.Org, Cornell University Library, 201 Olin Library Cornell university Ithaca, NY 14853. XP091379367. [cited by applicant]
Aaron Master et al., DeepSpace: Dynamic Spatial and Source Cue Based Source Separation for Dialog Enhancement, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY, 14853 Feb. 16, 2003. … [cited by applicant]
Davila-Chacon, J., et al., Neural and Statistical Processing of Spatial Cues for Sound Source Localisation, The 2013 International Joint Conference on Neural Networks (IJCNN), 2013, pp. 1-8, doi: 10.1109/IJCNN.2013.6706… [cited by applicant]
Tan Ke et al., Deep Learning Based Real-Time Speech Enhancement for Dual-Microphone Mobile Phones vol. 29, May 21, 2021, pp. 1853-1863. IEEE/ACM Transactions on Audio, Speech, and Language Processing, IEEEE, USA. XP0118… [cited by applicant]
Master, S., et al., Dialog Enhancement via Spatio-Level Filtering and Classification, Audio Engineering Society, Convention Paper 10427, Oct. 2020. [cited by applicant]
Gu, R., et al., Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature Learning, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, 2020, pp. 7319-7323, XP0… [cited by applicant]
Fabian-Robert Stöter, et al., The 2018 Signal Separation Evaluation Campaign. arXiv:1804.06267v3 [eess.AS]. https://arxiv.org/abs/1804.0626. [cited by applicant]
Han, C., et al., Real-Time Binaural Speech Separation With Preserved Spatial Cues, Dept. of Electrical Engineering, Columbia University, New York, NY, Feb. 16, 2020, https://arxiv.org/abs/2002.06637. [cited by applicant]
Le Roux, J., et al., The Phasebook: Building Complex Masks Via Discrete Representations for Source Separation, ICASSP 2019. https://ieeexplore.ieee.org/document/8682587. [cited by applicant]
Master, Steven A., Stereo Music Source Separation via Bayesian Modeling, Doctoral thesis, Stanford University, http://citeseerx.ist.psu.edu/viewdoc/download?doi=I0.1.1.81.7477&rep=repl&type=pdf. [cited by applicant]
Reddy, A. M., et al., Soft Mask Methods for Single-Channel Speaker Separation, IEEE Transactions on Audio, Speech, and Language Processing, vol. 15, No. 6, pp. 1766-1776, Aug. 2007. [cited by applicant]
Kim Biho et al., Speech enhancement based on soft-masking exploiting both output SNR and selectivity of spatial filtering Electronics Letters, The Institution of Engineering and Technology, GB vol. 50, No. 12, Jun. 5, 2… [cited by applicant]
Toroghi, R.M., Blind Speech Separation in Distant Speech Recognition Front-end Processing, Doctoral dissertation, Saarland University, Saarbrücken, Germany, Nov. 16, 2016. [cited by applicant]
Wang, De Liang, Time-Frequency Masking for Speech Separation and its Potential for Hearing Aid Design, Trends Amplif. 2008 Fall; 12(4): pp. 332-353. [cited by applicant]
Xia, S., et al., Using Optimal Ratio Mask as Training Target for Supervised Speech Separation, 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) (pp. 163-166). IEE… [cited by applicant]