IP Library › Granted Patent US 12,670,924
Granted Patent B2
US 12,670,924 · App. 18/525,307 · Granted Jun 30, 2026

Spatial region based audio separation

Inventors: George Chiachi Sung (San Diego, CA); Yang Yang (San Diego, CA); Shao-Fu Shih (San Jose, CA); Hakan Erdogan (Lexington, MA); Kevin Lee (Mountan View, CA)
Assignee: Google LLC
G10L21/057G10L15/063G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,924
App. No.
18/525,307
Filed
Nov 30, 2023
Granted
Jun 30, 2026
Kind
B2
Art Unit
2653
USPC
704/232
Abstract

A method includes receiving target audio data captured by a first audio input device, the target audio data comprising a target audio signal and a first version of an interfering audio signal, and receiving reference audio data captured by a second audio input device different from the first audio input device, the reference audio data comprising a second version of the interfering audio signal. The method also includes processing, using a trained neural network, the target audio data and the reference audio data to generate enhanced audio data, the neural network attenuating the interfering audio signal in the enhanced audio data.

Claims (72)

1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving target audio data captured by a first audio input device, the target audio data comprising a target audio signal and a first version of an interfering audio signal; receiving reference audio data captured by a second audio input device different from the first audio input device, the reference audio data comprising a second version of the interfering audio signal;

processing, using a trained neural network, the target audio data and the reference audio data to generate enhanced audio data;

determining at least one of a delay contrast or a magnitude contrast between the first version of the interfering audio signal and the second version of the interfering audio signal; and

attenuating, using the trained neural network, at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio signal based on the at least one of the delay contrast or the magnitude contrast.

2 . The computer-implemented method of claim 1 , wherein:

the target audio signal originates from a first region, the first region defined by a first set of angles; and

the neural network is configured to attenuate interfering audio signals originating in a second region different from the first region, the second region defined by a second set of angles different from the first set of angles.

3 . The computer-implemented method of claim 2 , wherein the first audio input device and the second audio input device are symmetrically arranged relative to the first region.

4 . The computer-implemented method of claim 1 , wherein the delay contrast represents an angular separation between a source of the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal and a target signal reception region.

5 . The computer-implemented method of claim 1 , wherein:

attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio data based on the at least one of the delay contrast or the magnitude contrast comprises attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal when the delay contrast satisfies a threshold; and

the operations further comprise:

receiving an input representing a time shift; and

time shifting the reference audio data by the time shift to effectively adjust a value of the threshold.

6 . The computer-implemented method of claim 1 , wherein:

the target audio signal originates from a first region, the first region defined by a first set of distances; and

the neural network is configured to attenuate interfering audio signals originating in a second region different from the first region, the second region defined by a second set of distances different from the first set of distances.

7 . The computer-implemented method of claim 1 , wherein the magnitude contrast represents a distance separation between a source of the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal and a target signal reception region.

8 . The computer-implemented method of claim 1 , wherein:

attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio data based on the at least one of the delay contrast or the magnitude contrast comprises attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal when the magnitude contrast satisfies a threshold; and

the operations further comprise:

receiving an input representing a scalar; and

multiplying the reference audio data by the scalar to effectively change the threshold.

9 . The computer-implemented method of claim 1 , wherein a training process trains the neural network by:

obtaining training target audio data comprising sampled speech of interest and a sampled first version of an interfering audio signal;

obtaining training reference audio data comprising a sampled second version of the interfering audio signal;

processing, using the neural network, the training target audio data and the training reference audio data to generate predicted enhanced audio data; and

training the neural network based on a loss term computed based on the predicted enhanced audio data and the sampled speech of interest.

10 . The computer-implemented method of claim 1 , wherein the neural network comprises a U-net model architecture.

11 . The computer-implemented method of claim 10 , wherein the U-net model architecture comprises:

a Fourier transform layer;

a contracting path comprising a plurality of two-dimensional (2D) convolution layers trained to successively reduce spatial information while increasing feature information;

an expansion path comprising a plurality of 2D convolution layers trained to combine feature and spatial information through a sequence of up-convolutions and concatenations with high-resolution features from the contracting path; and

and an inverse Fourier transform layer.

12 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving target audio data captured by a first audio input device, the target audio data comprising a target audio signal and a first version of an interfering audio signal;

receiving reference audio data captured by a second audio input device different from the first audio input device, the reference audio data comprising a second version of the interfering audio signal;

processing, using a trained neural network, the target audio data and the reference audio data to generate enhanced audio data;

determining at least one of a delay contrast or a magnitude contrast between the first version of the interfering audio signal and the second version of the interfering audio signal; and

attenuating, using the trained neural network, at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio signal based on the at least one of the delay contrast or the magnitude contrast.

13 . The system of claim 12 , wherein:

the target audio signal originates from a first region, the first region defined by a first set of angles; and

the neural network is configured to attenuate interfering audio signals originating in a second region different from the first region, the second region defined by a second set of angles different from the first set of angles.

14 . The system of claim 13 , wherein the first audio input device and the second audio input device are symmetrically arranged relative to the first region.

15 . The system of claim 12 , wherein the delay contrast represents an angular separation between a source of the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal and a target signal reception region.

16 . The system of claim 12 , wherein:

attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio data based on the at least one of the delay contrast or the magnitude contrast comprises attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal when the delay contrast satisfies a threshold; and

the operations further comprise:

receiving an input representing a time shift; and time shifting the reference audio data by the time shift to effectively adjust a value of the threshold.

17 . The system of claim 12 , wherein:

the target audio signal originates from a first region, the first region defined by a first set of distances; and

the neural network is configured to attenuate interfering audio signals originating in a second region different from the first region, the second region defined by a second set of distances different from the first set of distances.

18 . The system of claim 12 , wherein the magnitude contrast represents a distance separation between a source of the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal and a target signal reception region.

19 . The system of claim 12 , wherein:

attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal in the enhanced audio data based on the at least one of the delay contrast or the magnitude contrast comprises attenuating the at least one of the first version of the interfering audio signal or the second version of the interfering audio signal when the magnitude contrast satisfies a threshold; and

the operations further comprise:

receiving an input representing a scalar; and

multiplying the reference audio data by the scalar to effectively change the threshold.

20 . The system of claim 12 , wherein a training process trains the neural network by:

obtaining training target audio data comprising sampled speech of interest and a sampled first version of an interfering audio signal;

obtaining training reference audio data comprising a sampled second version of the interfering audio signal;

processing, using the neural network, the training target audio data and the training reference audio data to generate predicted enhanced audio data; and

training the neural network based on a loss term computed based on the predicted enhanced audio data and the sampled speech of interest.

21 . The system of claim 12 , wherein the neural network comprises a U-net model architecture.

22 . The system of claim 21 , wherein the U-net model architecture comprises:

a Fourier transform layer;

a contracting path comprising a plurality of two-dimensional (2D) convolution layers trained to successively reduce spatial information while increasing feature information;

an expansion path comprising a plurality of 2D convolution layers trained to combine feature and spatial information through a sequence of up-convolutions and concatenations with high-resolution features from the contracting path; and

and an inverse Fourier transform layer.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2024
From: SUNG, GEORGE CHIACHI; YANG, YANG; SHIH, SHAO-FU; ERDOGAN, HAKAN; LEE, KEVIN
To: GOOGLE LLC
Reel/Frame 066901/0285 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2023
From: SUNG, GEORGE CHIACHI; YANG, YANG; SHIH, SHAO-FU; ERDOGAN, HAKAN; LEE, KEVIN
To: GOOGLE LLC
Reel/Frame 065840/0705 →
Continuity (1)
Related Publication 20250182775A1 · Jun 5, 2025
References Cited (17)
US 11699454B1 · D Souza · 2023 [cited by examiner]
US 12169663B1 · Nagisetty · 2024 [cited by examiner]
US 20140241529A1 · Lee · 2014 [cited by applicant]
US 20200349954A1 · Yoshioka et al. · 2020 [cited by applicant]
US 20210043220A1 · Baek et al. · 2021 [cited by applicant]
US 20220068257A1 · Biadsy · 2022 [cited by examiner]
US 20230032280A1 · Celestinos Arroyo et al. · 2023 [cited by applicant]
US 20250118318A1 · Montazeri · 2025 [cited by examiner]
CN 117499852A · 2024 [cited by examiner]
T. S. Gunawan, M. R. M. Sarif, M. Kartiwi and Y. A. Ahmad, “Development of U-Net Architecture for Audio Super Resolution,” 2023 9th International Conference on Computer and Communication Engineering (ICCCE), Kuala Lumpu… [cited by examiner]
Nugraha, Aditya Arie, Antoine Liutkus, and Emmanuel Vincent. “Multichannel audio source separation with deep neural networks.” IEEE/ACM Transactions on Audio, Speech, and Language Processing 24.9 (2016): 1652-1664. [cited by examiner]
Xu, Alan, and Romit Roy Choudhury. “Learning to separate voices by spatial regions.” International Conference on Machine Learning. PMLR, 2022. [cited by examiner]
International Search Report and Written Opinion issued in related PCT Application No. PCT/US2024/054507, dated Dec. 23, 2024. [cited by applicant]
Dejan Markovic et al: “Implicit Neural Spatial Filtering for Multichannel Source Separation in the Waveform Domain”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 30, … [cited by applicant]
Teerapat Jenrungrot et al: “The Cone of Silence: Speech Separation by Localization”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 12, 2020 (Oct. 12, 2020), XP08178496… [cited by applicant]
Tan Ke et al: “Neural Spectrospatial Filtering”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, Jan. 25, 2022 (Jan. 25, 2022), pp. 605-621, XP093232280, ISSN: 2329-9290, DOI:10.1109/TASLP.2022… [cited by applicant]
Yang Yang et al: “Binaural Angular Separation Network”, ICASSP 2024—2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Apr. 14, 2024 (Apr. 14, 2024), pp. 12-01-1205, XP03465454… [cited by applicant]