IP Library Granted Patent US 12,367,887
Granted Patent B2
US 12,367,887 · App. 18/009,651 · Granted Jul 22, 2025

Separation of panned sources from generalized stereo backgrounds using minimal training

Inventor: Aaron Steven Master (San Francisco, CA)
Assignee: DOLBY LABORATORIES LICENSING CORPORATION
G10L19/02G06F3/165G10L25/21H04S7/30H04S2400/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,887
App. No.
18/009,651
Granted
Jul 22, 2025
Kind
B2
Abstract

In an embodiment, a spatio-level filter (SLF) is created by obtaining a first set of samples from a plurality of target source level and spatial distributions in frequency subbands in a frequency domain, obtaining a second set of samples from a plurality of background level and spatial distributions in frequency subbands in a frequency domain, adding the first and second sets of samples to create a combined set of samples, detecting level and spatial parameters for each sample in the combined set of samples for each subband, within subbands, weighting the detected level and spatial parameters by their respective level and spatial distributions for the target source and backgrounds; storing the weighted level, spatial parameters and signal-to-noise ratio (SNR) within subbands for each sample in the combined set of samples in a table; and re-indexing the table by the weighted level and spatial parameters and subband.

Claims (29)

1. A method comprising:

sampling, using one or more processors, a first spatial and level distribution of sound sources to obtain a first set of samples representing a target source;

sampling, using the one or more processors, a second spatial and level distribution of sound sources to obtain a second set of samples representing background sources;

obtaining, using the one or more processors, a frequency domain representation of the first set of samples in a plurality of frequency subbands;

obtaining, using the one or more processors, a frequency domain representation of the second set of samples in the plurality of frequency subbands;

pairwise adding, using the one or more processors, the frequency domain representations of the first and second sets of samples to create a plurality of training data sets;

detecting, using the one or more processors, respective parameter values for each of the training data sets for each frequency subband in the plurality of frequency subbands;

for each frequency subband of the plurality of frequency subbands, weighting, using the one or more processors, a contribution to a popularity array of each combination of the respective parameter values by the first and second spatial and level distributions;

storing, using the one or more processors, the popularity array and corresponding signal-to-noise ratio (SNR) values for the plurality of frequency subbands in a table, wherein each of the corresponding SNR values represents a level ratio of the target source to the background sources in a respective training data set; and

re-indexing, using the one or more processors, the table based on a Bayesian relationship between distributions of the respective parameter values and of the corresponding SNR values such that, for a given input set of parameter values, estimated SNR values associated with the given input set of parameter values for the plurality of frequency subbands are obtained from the re-indexed table via a lookup operation.

2. The method of claim 1 , further comprising:

smoothing data in the popularity array over one or more dimensions of the table.

3. The method of claim 1 , wherein the frequency domain representation is a short-time Fourier transform (STFT) domain representation.

4. The method of claim 1 , wherein spatial parameters represented by the respective parameter values include panning and a phase difference between two channels of a mixed audio signal corresponding to the plurality of training data sets.

5. The method of claim 1 , wherein the target source is amplitude panned using a constant power law.

6. A method comprising:

transforming, using one or more processors, one or more frames of a two-channel time domain audio signal into a time-frequency domain representation including a plurality of time-frequency tiles, wherein the time-frequency domain representation includes a plurality of frequency bins grouped into a plurality of subbands;

for each tile of the time-frequency domain representation:

calculating, using the one or more processors, a respective set of spatial parameter values and a respective level;

generating, using the one or more processors and a lookup table, a respective percentile signal-to-noise ratio (SNR) value for each frequency subband of the plurality of frequency subbands, wherein the lookup table is trained based on a Bayesian relationship between distributions of parameter values and corresponding SNR values such that, in response to a given input set of parameter values, the lookup table returns estimated SNR values associated with the given input set of parameter values for the plurality of frequency subbands, and wherein each of the corresponding SNR values represents a level ratio of the target source to background sources in a respective frequency subband of the plurality of frequency subbands;

generating, using the one or more processors, a softmask of fractional values based on the respective percentile SNR values, with each fractional value representing an audio portion of a corresponding subband attributed to the target source in a respective tile of the time-frequency domain representation; and

applying, using the one or more processors, the softmask of fractional values for to the plurality of frequency subbands in the respective tile to generate a modified time-frequency tile corresponding to the target source.

7. The method of claim 6 , further comprising:

transforming, using the one or more processors, the modified time-frequency tile into a corresponding time domain audio signal.

8. The method of claim 6 , wherein the transforming comprises applying a short-time frequency transform (STFT) to the two-channel time domain audio signal.

9. The method of claim 6 , wherein the plurality of frequency bins is grouped to form octave subbands or approximately octave subbands.

10. An apparatus comprising:

one or more processors; and

a memory storing instructions that when executed by the one or more processors, cause the apparatus to perform the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2023
From: MASTER, AARON STEVEN
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 064440/0306 →
Priority Claims (1)
EP 20179449 · Jun 11, 2020 · regional
Continuity (2)
Provisional Application 63038046 · Jun 11, 2020
Related Publication 20230245664A1 · Aug 3, 2023
References Cited (41)
US 7454333B2 · Ramakrishnan · 2008 [cited by applicant]
US 7912232B2 · Master · 2011 [cited by applicant]
US 7970144B1 · Avendano · 2011 [cited by applicant]
US 8027478B2 · Barry · 2011 [cited by applicant]
US 8280077B2 · Avendano · 2012 [cited by applicant]
US 8812322B2 · Mysore · 2014 [cited by applicant]
US 9564144B2 · Nesta · 2017 [cited by applicant]
US 10049678B2 · Nesta · 2018 [cited by applicant]
US 10433075B2 · Crow · 2019 [cited by applicant]
US 20070076902A1 · Master · 2007 [cited by applicant]
US 20090238371A1 · Rumsey · 2009 [cited by applicant]
US 20120128164A1 · Blamey · 2012 [cited by examiner]
US 20150071446A1 · Sun · 2015 [cited by applicant]
US 20160261961A1 · Andersen · 2016 [cited by examiner]
US 20170206907A1 · Wang · 2017 [cited by applicant]
US 20180144759A1 · Lu · 2018 [cited by applicant]
US 20190098399A1 · Lashkari · 2019 [cited by applicant]
US 20200075030A1 · Tsilfidis · 2020 [cited by applicant]
US 20200143819A1 · Delcroix · 2020 [cited by applicant]
US 20200191913A1 · Zhang · 2020 [cited by examiner]
US 20220408180A1 · Lepoutre · 2022 [cited by examiner]
US 20230245671A1 · Master · 2023 [cited by examiner]
US 20230333201A1 · Regani · 2023 [cited by examiner]
US 20230335141A1 · Pihlajakuja · 2023 [cited by examiner]
JP 2006349840A · 2006 [cited by applicant]
JP 2009218663A · 2009 [cited by applicant]
JP 2018031910A · 2018 [cited by applicant]
KR 20140042900A · 2014 [cited by applicant]
RU 2596592C2 · 2016 [cited by applicant]
RU 2640742C1 · 2018 [cited by applicant]
RU 2673390C1 · 2018 [cited by applicant]
WO 2017141542A1 · 2017 [cited by applicant]
WO 2019193070A1 · 2019 [cited by applicant]
Guo, Haiyan et al; “Single-Channel Speaker Separation Based on Sub-Spectrum GMM and Bayesian Theory”; ICSP2008 Proceedings; pp. 701-704. [cited by applicant]
Itakura, Kousuke et al.; “Bayesian Multichannel Nonnegative Matrix Factorization for Audio Source Separation and Localization”; Graduate School of Informatics; Kyoto University; Sakyo-ku, Kyoto 606-8501; Japan; ICASSP 2… [cited by applicant]
Itakura, Kousuke et al.; “Bayesian Multichannel Audio Source Separation Based on Integrated Source and Spatial Models”; IEEE/ACM Transactions on Audio, Speech, and Language Processing; vol. 26; No. 4; Apr. 2018; pp. 831… [cited by applicant]
Master, Steven A., Stereo Music Source Separation via Bayesian Modeling, Doctoral thesis, Stanford University, http://citeseerx.ist.psu.edu/viewdoc/download?doi=l0.1.1.81.7477&rep=repl&type=pdf. [cited by applicant]
Ozerov, Alexey et al.; “Adaption of Bayesian Models for Single Channel Source Separation and its Application to Voice/Music Separation in Popular Songs”; IEEE Transactions on Audio, Speech and Language Processing; Insti… [cited by applicant]
Sekiguchi, Kouhei et al.; “Bayesian Multichannel Speech Enhancement with a Deep Speech Prior”; Proceedings, APSIPA Annual Summit and Conference 2018; Nov. 12-15, 2018; Hawaii; pp. 1233-1239. [cited by applicant]
Wang, Wenwu et al.; “Video Assisted Speech Source Separation”; ICASSP 2005; pp. V-425-V-428. [cited by applicant]
Xu, Tao et al.; “Methods for Learning Adaptive Dictionary in Underdetermined Speech Separation”; Machine Learning for Signal Processing (MLSP); 2011 IEEE International Workshop ON, IEEE; Sep. 18, 2011; pp. 1-6. [cited by applicant]