IP Library Granted Patent US 12,014,722
Granted Patent B2
US 12,014,722 · App. 17/197,587 · Granted Jun 18, 2024

System and method for data augmentation of feature-based voice data

Inventors: Dushyant Sharma (Woburn, MA); Patrick A. Naylor (Reading, GB); James W. Fosburgh (Winchester, MA)
Assignee: Microsoft Technology Licensing, LLC
G10L13/02G06F3/165G06N5/02G06N20/00G10K15/08G10L13/033G10L15/02G10L15/063G10L15/065G10L21/0224G10L25/03H04S7/30H04S7/302H04S7/303
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,014,722
App. No.
17/197,587
Granted
Jun 18, 2024
Kind
B2
Abstract

A method, computer program product, and computing system for receiving feature-based voice data associated with a first acoustic domain. One or more gain-based augmentations may be performed on at least a portion of the feature-based voice data, thus defining gain-augmented feature-based voice data.

Claims (42)

1. A computer-implemented method, executed on a computing device, comprising:

receiving feature-based voice data associated with a first acoustic domain, wherein the feature-based voice data is converted from a signal in the first acoustic domain to a feature domain;

performing one or more gain-based augmentations on at least a portion of the feature-based voice data converted from the signal in the first acoustic domain to the feature domain, thus defining gain-augmented feature-based voice data;

receiving a selection of a target acoustic domain;

determining a distribution of gain levels from training data associated with the target acoustic domain varies over time for one or more of particular frequencies and particular frequency bands;

mapping the gain-augmented feature-based voice data from the first acoustic domain to the target acoustic domain, and

training a speech processing system for the target acoustic domain based on the gain-augmented feature-based voiced data mapped from the first acoustic domain to the target acoustic domain.

2. The computer-implemented method of claim 1 , further comprising:

receiving the selection of the target acoustic domain via at least one of a user interface and a database.

3. The computer-implemented method of claim 2 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain.

4. The computer-implemented method of claim 2 , further comprising:

determining a distribution of gain levels associated with the target acoustic domain.

5. The computer-implemented method of claim 4 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the distribution of gain levels associated with the target acoustic domain.

6. The computer-implemented method of claim 1 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes amplifying the at least a portion of the feature-based voice data.

7. The computer-implemented method of claim 1 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes attenuating the at least a portion of the feature-based voice data.

8. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

receiving feature-based voice data associated with a first acoustic domain, wherein the feature-based voice data is converted from a signal in the first acoustic domain to a feature domain;

performing one or more gain-based augmentations on at least a portion of the feature-based voice data converted from the signal in the first acoustic domain to the feature domain, thus defining gain-augmented feature-based voice data;

receiving a selection of a target acoustic domain;

determining a distribution of gain levels from training data associated with the target acoustic domain varies over time for one or more of particular frequencies and particular frequency bands;

mapping the gain-augmented feature-based voice data from the first acoustic domain to the target acoustic domain, and

training a speech processing system for the target acoustic domain based on the gain-augmented feature-based voiced data mapped from the first acoustic domain to the target acoustic domain.

9. The computer program product of claim 8 , wherein the operations further comprise:

receiving the selection of the target acoustic domain via at least one of a user interface and a database.

10. The computer program product of claim 9 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain.

11. The computer program product of claim 9 , wherein the operations further comprise:

determining a distribution of gain levels associated with the target acoustic domain.

12. The computer program product of claim 11 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the distribution of gain levels associated with the target acoustic domain.

13. The computer program product of claim 8 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes amplifying the at least a portion of the feature-based voice data.

14. The computer program product of claim 8 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes attenuating the at least a portion of the feature-based voice data.

15. A computing system comprising:

a memory; and

a processor configured to receive feature-based voice data associated with a first acoustic domain, wherein the feature-based voice data is converted from a signal in the first acoustic domain to a feature domain, wherein the processor is further configured to perform one or more gain-based augmentations on at least a portion of the feature-based voice data converted from the signal in the first acoustic domain to the feature domain, thus defining gain-augmented feature-based voice data, wherein the processor is further configured to receive a selection of a target acoustic domain, wherein the processor is further configured to determine a distribution of gain levels from training data associated with the target acoustic domain varies over time for one or more of particular frequencies and particular frequency bands; wherein the processor is further configured to map the gain-augmented feature-based voice data from the first acoustic domain to the target acoustic domain; and wherein the processor is further configured to train a speech processing system for the target acoustic domain based on the gain-augmented feature-based voiced data mapped from the first acoustic domain to the target acoustic domain.

16. The computing system of claim 15 , wherein the processor is further configured to:

receive the selection of the target acoustic domain via at least one of a user interface and a database.

17. The computing system of claim 16 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain.

18. The computing system of claim 16 , wherein the processor is further configured to:

determine a distribution of gain levels associated with the target acoustic domain.

19. The computing system of claim 18 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the distribution of gain levels associated with the target acoustic domain.

20. The computing system of claim 15 , wherein performing the one or more gain-based augmentations to the at least a portion of the feature-based voice data includes one or more of:

amplifying the at least a portion of the feature-based voice data; and

attenuating the at least a portion of the feature-based voice data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2021
From: SHARMA, DUSHYANT; NAYLOR, PATRICK A.; FOSBURGH, JAMES W.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 055553/0608 →
Continuity (2)
Provisional Application 62988337 · Mar 11, 2020
Related Publication 20210287660A1 · Sep 16, 2021
Cited By (1)
US 12,288,602