IP Library Granted Patent US 12,154,541
Granted Patent B2
US 12,154,541 · App. 17/197,717 · Granted Nov 26, 2024

System and method for data augmentation of feature-based voice data

Inventors: Dushyant Sharma (Mountain House, CA); Patrick A. Naylor (Reading, GB); James W. Fosburgh (Baltimore, MD); Do Yeong Kim (Lexington, MA)
Assignee: Microsoft Technology Licensing, LLC
G10L13/02G06F3/165G06N5/02G06N20/00G10K15/08G10L13/033G10L15/02G10L15/063G10L15/065G10L21/0224G10L25/03H04S7/30H04S7/302H04S7/303
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,154,541
App. No.
17/197,717
Granted
Nov 26, 2024
Kind
B2
Abstract

A method, computer program product, and computing system for receiving feature-based voice data associated with a first acoustic domain. One or more reverberation-based augmentations may be performed on at least a portion of the feature-based voice data, thus defining reverberation-augmented feature-based voice data.

Claims (31)

1. A computer-implemented method, executed on a computing device, comprising:

receiving feature-based voice data associated with a first acoustic domain, wherein the feature-based voice data is converted from an audio signal in the first acoustic domain to a feature domain;

extracting acoustic metadata from the audio signal before the audio signal is converted from the first acoustic domain to the feature domain;

processing the feature-based voice data associated with the audio signal based upon, at least in part, the acoustic metadata by performing one or more gain-based augmentations on at least a portion of the feature-based voice data associated with the audio signal in the first acoustic domain based upon, at least in part, the acoustic metadata;

receiving selections of various characteristics of the acoustic domain to define a target acoustic domain from a library of predefined acoustic domains;

determining a distribution of one or more reverberation levels associated with the target acoustic domain; and

performing one or more reverberation-based augmentations on at least a portion of the feature-based voice data converted from the audio signal in the first acoustic domain to the feature domain based upon, at least in part the distribution of the one or more reverberation levels associated with the target acoustic domain, thus defining reverberation-augmented feature-based voice data.

2. The computer-implemented method of claim 1 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain.

3. The computer-implemented method of claim 1 , further comprising: training a machine learning model with one or more room impulse responses associated with the target acoustic domain.

4. The computer-implemented method of claim 3 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data using the trained machine learning model configured to model the reverberation associated with the target acoustic domain.

5. The computer-implemented method of claim 1 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes adding reverberation to the at least a portion of the feature-based voice data.

6. The computer-implemented method of claim 1 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes removing reverberation to the at least a portion of the feature-based voice data.

7. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

receiving feature-based voice data associated with a first acoustic domain, wherein the feature-based voice data is converted from an audio signal in the first acoustic domain to a feature domain;

extracting acoustic metadata from the audio signal before the audio signal is converted from the first acoustic domain to the feature domain;

processing the feature-based voice data associated with the audio signal based upon, at least in part, the acoustic metadata by performing one or more gain-based augmentations on at least a portion of the feature-based voice data associated with the audio signal in the first acoustic domain based upon, at least in part, the acoustic metadata;

receiving selections of various characteristics of the acoustic domain to define a target acoustic domain from a library of predefined acoustic domains;

determining a distribution of one or more reverberation levels associated with the target acoustic domain; and

performing one or more reverberation-based augmentations on at least a portion of the feature-based voice data converted from the audio signal in the first acoustic domain to the feature domain based upon, at least in part the distribution of the one or more reverberation levels associated with the target acoustic domain, thus defining reverberation-augmented feature-based voice data.

8. The computer program product of claim 7 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain.

9. The computer program product of claim 7 , wherein the operations further comprise: training a machine learning model with one or more room impulse responses associated with the target acoustic domain.

10. The computer program product of claim 9 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data using the trained machine learning model configured to model the reverberation associated with the target acoustic domain.

11. The computer program product of claim 7 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes adding reverberation to the at least a portion of the feature-based voice data.

12. The computer program product of claim 7 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes removing reverberation to the at least a portion of the feature-based voice data.

13. A computing system comprising:

a memory; and

a processor configured to receive feature-based voice data associated with a first acoustic domain, wherein the feature-based voice data is converted from an audio signal in the first acoustic domain to a feature domain, wherein the processor is further configured to extract acoustic metadata from the audio signal before the audio signal is converted from the first acoustic domain to the feature domain, wherein the processor is further configured to process the feature-based voice data associated with the audio signal based upon, at least in part, the acoustic metadata by performing one or more gain-based augmentations on at least a portion of the feature-based voice data associated with the audio signal in the first acoustic domain based upon, at least in part, the acoustic metadata, wherein the processor is further configured to receive selections of various characteristics of the acoustic domain to define a target acoustic domain from a library of predefined acoustic domains, wherein the processor is further configured to determine a distribution of one or more reverberation levels associated with the target acoustic domain, and wherein the processor is further configured to perform one or more reverberation-based augmentations on at least a portion of the feature-based voice data converted from the audio signal in the first acoustic domain to the feature domain based upon, at least in part the distribution of the one or more reverberation levels associated with the target acoustic domain, thus defining reverberation-augmented feature-based voice data.

14. The computing system of claim 13 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain.

15. The computing system of claim 13 , wherein the processor is further configured to: train a machine learning model with one or more room impulse responses associated with the target acoustic domain.

16. The computing system of claim 15 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data using the trained machine learning model configured to model the reverberation associated with the target acoustic domain.

17. The computing system of claim 13 , wherein performing the one or more reverberation-based augmentations to the at least a portion of the feature-based voice data includes one or more of: adding reverberation to the at least a portion of the feature-based voice data; and removing reverberation to the at least a portion of the feature-based voice data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2021
From: SHARMA, DUSHYANT; NAYLOR, PATRICK A.; FOSBURGH, JAMES W.; KIM, DO YEONG
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 055553/0953 →
Continuity (2)
Provisional Application 62988337 · Mar 11, 2020
Related Publication 20210287653A1 · Sep 16, 2021