IP Library Granted Patent US 12,073,818
Granted Patent B2
US 12,073,818 · App. 17/197,740 · Granted Aug 27, 2024

System and method for data augmentation of feature-based voice data

Inventors: Dushyant Sharma (Mountain House, CA); Patrick A. Naylor (Reading, GB); James W. Fosburgh (Baltimore, MD); Do Yeong Kim (Lexington, MA)
Assignee: Microsoft Technology Licensing, LLC
G10L13/02G06F3/165G06N5/02G06N20/00G10K15/08G10L13/033G10L15/02G10L15/063G10L15/065G10L21/0224G10L25/03H04S7/30H04S7/302H04S7/303
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,073,818
App. No.
17/197,740
Granted
Aug 27, 2024
Kind
B2
Abstract

A method, computer program product, and computing system for receiving feature-based voice data. One or more data augmentation characteristics may be received. One or more augmentations of the feature-based voice data may be generated, via a machine learning model, based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

Claims (39)

1. A computer-implemented method, executed on a computing device, comprising:

receiving feature-based voice data associated with a first acoustic domain, the first acoustic domain defined, at least in part, by signal processing characteristics associated with a microphone associated with the first acoustic domain, wherein receiving feature-based voice data includes extracting acoustic metadata from a signal, the acoustic metadata describing an acoustic characteristic of at least a portion of the feature-based voice data including one or more of presence of a speech component, speaking rate, and reverberation level;

qualifying at least a portion of the feature-based voice data for one or more of training data and adaptation data based upon, at least in part, the acoustic metadata, including:

receiving one or more constraints associated with processing feature-based voice data; and

comparing the one or more constraints to at least a portion of the acoustic metadata to determine whether the feature-based voice data is qualified for a particular task;

receiving a selection of a target acoustic domain, the target acoustic domain defined, at least in part, by a characteristic associated with a microphone associated with the target acoustic domain;

receiving one or more data augmentation characteristics; and

generating, via a machine learning model, one or more augmentations of the feature-based voice data based upon, at least in part, the feature-based voice data, selected target acoustic domain, and the one or more data augmentation characteristics.

2. The computer-implemented method of claim 1 , wherein receiving the feature-based voice data includes:

converting the signal from the time domain to the feature domain, thus defining the feature-based voice data.

3. The computer-implemented method of claim 2 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes generating, via the machine learning model, the one or more augmentations of the feature-based voice data based upon, at least in part, the feature-based voice data, the one or more data augmentation characteristics, and the acoustic metadata.

4. The computer-implemented method of claim 1 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more gain-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

5. The computer-implemented method of claim 1 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more rate-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

6. The computer-implemented method of claim 1 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more audio feature-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

7. The computer-implemented method of claim 1 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more reverberation-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

8. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

receiving feature-based voice data associated with a first acoustic domain, the first acoustic domain defined, at least in part, by signal processing characteristics associated with a microphone associated with the first acoustic domain, wherein receiving feature-based voice data includes extracting acoustic metadata from a signal, the acoustic metadata describing an acoustic characteristic of at least a portion of the feature-based voice data including one or more of presence of a speech component, speaking rate, and reverberation level;

qualifying at least a portion of the feature-based voice data for one or more of training data and adaptation data based upon, at least in part, the acoustic metadata, including:

receiving one or more constraints associated with processing feature-based voice data; and

comparing the one or more constraints to at least a portion of the acoustic metadata to determine whether the feature-based voice data is qualified for a particular task;

receiving a selection of a target acoustic domain, the target acoustic domain defined, at least in part, by a characteristic associated with a microphone associated with the target acoustic domain;

receiving one or more data augmentation characteristics; and

generating, via a machine learning model, one or more augmentations of the feature-based voice data based upon, at least in part, the feature-based voice data, selected target acoustic domain, and the one or more data augmentation characteristics.

9. The computer program product of claim 8 , wherein receiving the feature-based voice data includes:

converting the signal from the time domain to the feature domain, thus defining the feature-based voice data.

10. The computer program product of claim 9 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes generating, via the machine learning model, the one or more augmentations of the feature-based voice data based upon, at least in part, the feature-based voice data, the one or more data augmentation characteristics, and the acoustic metadata.

11. The computer program product of claim 8 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more gain-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

12. The computer program product of claim 8 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more rate-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

13. The computer program product of claim 8 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more audio feature-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

14. The computer program product of claim 11 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more reverberation-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

15. A computing system comprising:

a memory; and

a processor configured to receive feature-based voice data associated with a first acoustic domain, the first acoustic domain defined, at least in part, by signal processing characteristics associated with a microphone associated with the first acoustic domain, wherein the processor configured to receive feature-based voice data is configured to extract acoustic metadata from a signal, the acoustic metadata describing an acoustic characteristic of at least a portion of the feature-based voice data including one or more of presence of a speech component, speaking rate, and reverberation level, wherein the processor is further configured to qualify at least a portion of the feature-based voice data for one or more of training data and adaptation data based upon, at least in part, the acoustic metadata, including: receiving one or more constraints associated with processing feature-based voice data; and comparing the one or more constraints to at least a portion of the acoustic metadata to determine whether the feature-based voice data is qualified for a particular task, wherein the processor is further configured to receive a selection of a target acoustic domain, the target acoustic domain defined, at least in part, by a characteristic associated with a microphone associated with the target acoustic domain, wherein the processor is further configured to receive one or more data augmentation characteristics, and wherein the processor is further configured to generate, via a machine learning model, one or more augmentations of the feature-based voice data based upon, at least in part, the feature-based voice data, the selected target acoustic domain, and the one or more data augmentation characteristics.

16. The computing system of claim 15 , wherein receiving the feature-based voice data includes:

converting the signal from the time domain to the feature domain, thus defining the feature-based voice data.

17. The computing system of claim 16 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes generating, via the machine learning model, the one or more augmentations of the feature-based voice data based upon, at least in part, the feature-based voice data, the one or more data augmentation characteristics, and the acoustic metadata.

18. The computing system of claim 15 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more gain-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

19. The computing system of claim 15 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more rate-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

20. The computing system of claim 15 , wherein generating, via the machine learning model, the one or more augmentations of the feature-based voice data includes performing, via the machine learning model, one or more audio feature-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the feature-based voice data and the one or more data augmentation characteristics.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2021
From: SHARMA, DUSHYANT; NAYLOR, PATRICK A.; FOSBURGH, JAMES W.; KIM, DO YEONG
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 055554/0133 →
Continuity (2)
Provisional Application 62988337 · Mar 11, 2020
Related Publication 20210287654A1 · Sep 16, 2021