IP Library Granted Patent US 12,469,513
Granted Patent B2
US 12,469,513 · App. 18/075,573 · Granted Nov 11, 2025

System and method for replicating background acoustic properties using neural networks

Inventors: Dushyant Sharma (Tracy, CA); James Wellford Fosburgh (Syracuse, NY); Patrick Aubrey Naylor (Reading, GB)
Assignee: Microsoft Technology Licensing, LLC
G10L21/0232G10L21/0264G10L21/034G10L25/18G10L25/30G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,513
App. No.
18/075,573
Granted
Nov 11, 2025
Kind
B2
Abstract

A method, computer program product, and computing system for estimating noise spectrum from a target audio signal segment. An acoustic neural embedding is generated from the target audio signal segment. An augmented audio signal segment is generated with background acoustic properties of the target audio signal segment by processing an input audio signal segment with the noise spectrum and the acoustic neural embedding using a neural network.

Claims (78)

1 . A computer-implemented method, executed on a computing device, comprising:

receiving a target audio signal segment recorded in a target acoustic environment, wherein a speech processing system is deployed in the target acoustic environment;

receiving an input audio signal segment generated by a text-to-speech (TTS) system, wherein background acoustic properties of the input audio signal segment mismatch background acoustic properties of the target audio signal segment;

estimating noise spectrum from the target audio signal segment;

generating an acoustic neural embedding from the target audio signal segment;

estimating loss associated with processing the target audio signal segment with the speech processing system;

generating an augmented audio signal segment with background acoustic properties matching the background acoustic properties of the target audio signal segment by processing the input audio signal segment to add noise and reverberation in accordance with the noise spectrum, the acoustic neural embedding, and the estimated loss associated with processing the target audio signal segment with the speech processing system, wherein a loss associated with processing the augmented audio signal segment with the speech processing system is within a threshold difference of the estimated loss associated with processing the target audio signal segment with the speech processing system; and

training the speech processing system based on training data that includes the augmented audio signal segment.

2 . The computer-implemented method of claim 1 , wherein generating the acoustic neural embedding includes extracting the acoustic neural embedding from the target audio signal segment using a Non-Intrusive Speech Assessment (NISA) system.

3 . The computer-implemented method of claim 1 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a neural filter from the input audio signal segment, wherein the neural filter is used to generate the augmented audio signal segment.

4 . The computer-implemented method of claim 1 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a filter mask for the acoustic neural embedding, wherein the filter mask is used to generate the augmented audio signal segment.

5 . The computer-implemented method of claim 1 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a neural filter from the input audio signal segment;

estimating a filter mask for the acoustic neural embedding; and

generating a multiplied filter in a frequency domain by multiplying the neural filter from the input audio signal segment and the filter mask for the acoustic neural embedding, wherein the multiplied filter is used to generate the augmented audio signal segment.

6 . The computer-implemented method of claim 1 , wherein generating the augmented audio signal segment further includes performing de-noising and de-reverberation on the input audio signal.

7 . The computer-implemented method of claim 1 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a neural filter from the input audio signal segment;

estimating a filter mask for the acoustic neural embedding;

generating a multiplied filter in a frequency domain by multiplying the neural filter from the input audio signal segment and the filter mask for the acoustic neural embedding;

generating a filtered audio signal segment by convolving the multiplied filter with the input audio signal segment a time domain or the frequency domain;

estimating a noise gain level using the filtered audio signal segment, the acoustic neural embedding, and the noise spectrum;

generating a noise signal segment by multiplying the noise spectrum by the noise gain level; and

generating the augmented audio signal segment by applying the noise signal segment to the filtered audio signal segment.

8 . A computing system comprising:

at least one processor; and

memory storing programming instructions for execution by the at least one processor, wherein the programming instructions, upon execution by the at least on processor, causes the computing system to perform the following operations:

receiving a target audio signal segment recorded in a target acoustic environment, wherein a speech processing system is deployed in the target acoustic environment;

receiving an input audio signal segment generated by a text-to-speech (TTS) system, wherein background acoustic properties of the input audio signal segment mismatch background acoustic properties of the target audio signal segment;

estimating a noise spectrum from the target audio signal segment;

generating an acoustic neural embedding from the target audio signal segment;

estimating loss associated with processing the target audio signal segment with the speech processing system;

generating an augmented audio signal segment with background acoustic properties matching the background acoustic properties of the target audio signal segment by processing the input audio signal segment to add noise and reverberation in accordance with the noise spectrum, the noise neural embedding, and the estimated loss associated with processing the target audio signal segment with the speech processing system, wherein a loss associated with processing the augmented audio signal segment with the speech processing system is within a threshold difference of the estimated loss associated with processing the target audio signal segment with the speech processing system; and

training the speech processing system based on training data that includes the augmented audio signal segment.

9 . The computing system of claim 8 , wherein generating the noise neural embedding includes extracting the noise neural embedding from the target audio signal segment using a Non-Intrusive Speech Assessment (NISA) system.

10 . The computing system of claim 8 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a neural filter from the input audio signal segment, wherein the neural filter is used to generate the augmented audio signal segment.

11 . The computing system of claim 8 , wherein generating the augmented audio signal segment further includes performing de-noising and de-reverberation on the input audio signal.

12 . The computing system of claim 8 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a neural filter from the input audio signal segment;

estimating a filter mask for the acoustic neural embedding; and

generating a multiplied filter in a frequency domain by multiplying the neural filter from the input audio signal segment and the filter mask for the acoustic neural embedding, wherein the multiplied filter is used to generate the augmented audio signal segment.

13 . The computing system of claim 8 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a neural filter from the input audio signal segment;

estimating a filter mask for the acoustic neural embedding;

generating a multiplied filter in a frequency domain by multiplying the neural filter from the input audio signal segment and the filter mask for the acoustic neural embedding;

generating a filtered audio signal segment by convolving the multiplied filter with the input audio signal segment a time domain or the frequency domain;

estimating a noise gain level using the filtered audio signal segment, the acoustic neural embedding, and the noise spectrum;

generating a noise signal segment by multiplying the noise spectrum by the noise gain level; and

generating the augmented audio signal segment by applying the noise signal segment to the filtered audio signal segment.

14 . A computer program product residing on a non-transitory computer readable medium having programming instructions stored thereon which, when executed by at least one processor of a system, cause the system to perform the following operations:

receiving a target audio signal segment recorded in a target acoustic environment, wherein a speech processing system is deployed in the target acoustic environment;

receiving an input audio signal segment generated by a text-to-speech (TTS) system, wherein background acoustic properties of the input audio signal segment mismatch background acoustic properties of the target audio signal segment;

estimating noise spectrum from the target audio signal segment;

generating an acoustic neural embedding from the target audio signal segment;

estimating loss associated with processing the target audio signal segment with the speech processing system; and

generating an augmented audio signal segment with background acoustic properties matching the background acoustic properties of the target audio signal segment by processing the input audio signal segment to add noise and reverberation in accordance with the noise spectrum, the acoustic neural embedding, and the estimated loss associated with processing the target audio signal segment with the speech processing system, wherein a loss associated with processing the augmented audio signal segment with the speech processing system is within a threshold difference of the estimated loss associated with processing the target audio signal segment with the speech processing system; and

training the speech processing system based on training data that includes the augmented audio signal segment.

15 . The computer program product of claim 14 , wherein generating the augmented audio signal segment further includes performing de-noising and de-reverberation on the input audio signal.

16 . The computer program product of claim 14 , wherein generating the acoustic neural embedding includes extracting the acoustic neural embedding from the target audio signal segment using a Non-Intrusive Speech Assessment (NISA) system.

17 . The computer program product of claim 14 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a neural filter from the input audio signal segment, wherein the neural filter is used to generate the augmented audio signal segment.

18 . The computer program product of claim 14 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a filter mask for the acoustic neural embedding, wherein the filter mask is used to generate the augmented audio signal segment.

19 . The computer program product of claim 14 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a neural filter from the input audio signal segment;

estimating a filter mask for the acoustic neural embedding; and

generating a multiplied filter in a frequency domain by multiplying the neural filter from the input audio signal segment and the filter mask for the acoustic neural embedding, wherein the multiplied filter is used to generate the augmented audio signal segment.

20 . The computer program product of claim 14 , wherein processing the input audio signal segment to add noise and reverberation includes:

estimating a neural filter from the input audio signal segment;

estimating a filter mask for the acoustic neural embedding;

generating a multiplied filter in a frequency domain by multiplying the neural filter from the input audio signal segment and the filter mask for the acoustic neural embedding;

generating a filtered audio signal segment by convolving the multiplied filter with the input audio signal segment a time domain or the frequency domain;

estimating a noise gain level using the filtered audio signal segment, the acoustic neural embedding, and the noise spectrum;

generating a noise signal segment by multiplying the noise spectrum by the noise gain level; and

generating the augmented audio signal segment by applying the noise signal segment to the filtered audio signal segment.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2025
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 070765/0651 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2022
From: SHARMA, DUSHYANT; FOSBURGH, JAMES WELLFORD; NAYLOR, PATRICK AUBREY
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 061990/0840 →