IP Library Granted Patent US 11,769,486
Granted Patent B2
US 11,769,486 · App. 17/178,734 · Granted Sep 26, 2023

System and method for data augmentation and speech processing in dynamic acoustic environments

Inventors: Patrick A. Naylor (Reading, GB); Dushyant Sharma (Woburn, MA); Uwe Helmut Jost (Groton, MA); William F. Ganong, III (Brookline, MA)
Assignee: Nuance Communications, Inc.
G10L15/063G06N20/00G10L25/03H04R1/406
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,769,486
App. No.
17/178,734
Granted
Sep 26, 2023
Kind
B2
Abstract

A method, computer program product, and computing system for defining model representative of a plurality of acoustic variations to a speech signal, thus defining a plurality of time-varying spectral modifications. The plurality of time-varying spectral modifications may be applied to a plurality of feature coefficients of a target domain of a reference signal, thus generating a plurality of time-varying spectrally-augmented feature coefficients of the reference signal.

Claims (42)

1. A computer-implemented method, executed on a computing device, comprising:

defining a model representative of a plurality of acoustic variations to a speech signal, thus defining a plurality of time-varying spectral modifications, wherein defining the model representative of the plurality of acoustic variations to the speech signal includes defining a model representative of a plurality of acoustic variations to the speech signal dependent upon variations in a beampattern of an adaptive beamformer and movement of a plurality of sound sources relative to a beam of the adaptive beamformer; and

applying the plurality of time-varying spectral modifications to a plurality of feature coefficients of a target domain of a reference signal, thus generating a plurality of time-varying spectrally-augmented feature coefficients of the reference signal.

2. The computer-implemented method of claim 1 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes:

defining a model representative of a plurality of acoustic variations to the speech signal associated with a change in a relative position of a speaker and a microphone.

3. The computer-implemented method of claim 1 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes one or more of:

modeling the plurality of acoustic variations to the speech signal using a statistical distribution; and

modeling the plurality of acoustic variations to the speech signal using a mathematical model representative of the plurality of acoustic variations to the speech signal associated with a particular use-case scenario.

4. The computer-implemented method of claim 1 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes receiving one or more inputs associated with one or more of speaker movement and speaker orientation.

5. The computer-implemented method of claim 4 , further comprising:

training a speech processing system using the plurality of time-varying spectrally-augmented feature coefficients of the reference signal and the one or more inputs associated with one or more of speaker location and speaker orientation.

6. The computer-implemented method of claim 1 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes generating, via a machine learning model, a mapping of the plurality of acoustic variations to one or more feature coefficients of a target domain.

7. The computer-implemented method of claim 1 , wherein applying the plurality of time-varying spectral modifications to the plurality of feature coefficients of the target domain of the reference signal includes simultaneously generating, via a machine learning model, the mapping of the plurality of acoustic variations to the one or more feature coefficients of the target domain and applying, via the machine learning model, the plurality of time-varying spectral modifications to the plurality of feature coefficients of the reference signal.

8. The computer-implemented method of claim 1 , further comprising:

training a speech processing system using the plurality of time-varying spectrally-augmented feature coefficients of the reference signal, thus defining a trained speech processing system.

9. The computer-implemented method of claim 8 , further comprising:

performing speech processing via the trained speech processing system, wherein the trained speech processing system is executed on at least one computing device.

10. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

defining a model representative of a plurality of acoustic variations to a speech signal, thus defining a plurality of time-varying spectral modifications, wherein defining the model representative of the plurality of acoustic variations to the speech signal includes defining a model representative of a plurality of acoustic variations to the speech signal dependent upon variations in a beampattern of an adaptive beamformer and movement of a plurality of sound sources relative to a beam of the adaptive beamformer; and

applying the plurality of time-varying spectral modifications to a plurality of feature coefficients of a target domain of a reference signal, thus generating a plurality of time-varying spectrally-augmented feature coefficients of the reference signal.

11. The computer program product of claim 10 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes:

defining a model representative of a plurality of acoustic variations to the speech signal associated with a change in a relative position of a speaker and a microphone.

12. The computer program product of claim 10 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes one or more of:

modeling the plurality of acoustic variations to the speech signal using a statistical distribution; and

modeling the plurality of acoustic variations to the speech signal using a mathematical model representative of the plurality of acoustic variations to the speech signal associated with a particular use-case scenario.

13. The computer program product of claim 10 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes receiving one or more inputs associated with one or more of speaker movement and speaker orientation.

14. The computer program product of claim 13 , wherein the operations further comprise:

training a speech processing system using the plurality of time-varying spectrally-augmented feature coefficients of the reference signal and the one or more inputs associated with one or more of speaker location and speaker orientation.

15. The computer program product of claim 10 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes generating, via a machine learning model, a mapping of the plurality of acoustic variations to one or more feature coefficients of a target domain.

16. The computer program product of claim 10 , wherein applying the plurality of time-varying spectral modifications to the plurality of feature coefficients of the target domain of the reference signal includes simultaneously generating, via a machine learning model, the mapping of the plurality of acoustic variations to the one or more feature coefficients of the target domain and applying, via the machine learning model, the plurality of time-varying spectral modifications to the plurality of feature coefficients of the reference signal.

17. The computer program product of claim 10 , further comprising:

training a speech processing system using the plurality of time-varying spectrally-augmented feature coefficients of the reference signal, thus defining a trained speech processing system.

18. The computer program product of claim 17 , further comprising:

performing speech processing via the trained speech processing system, wherein the trained speech processing system is executed on at least one computing device.

19. A computing system comprising:

a memory; and

a processor configured to define a model representative of a plurality of acoustic variations to a speech signal, thus defining a plurality of time-varying spectral modifications, wherein defining the model representative of the plurality of acoustic variations to the speech signal includes defining a model representative of a plurality of acoustic variations to the speech signal dependent upon variations in a beampattern of an adaptive beamformer and movement of a plurality of sound sources relative to a beam of the adaptive beamformer and wherein the processor is further configured to apply the plurality of time-varying spectral modifications to a plurality of feature coefficients of a target domain of a reference signal, thus generating a plurality of time-varying spectrally-augmented feature coefficients of the reference signal.

20. The computing system of claim 19 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes:

defining a model representative of a plurality of acoustic variations to the speech signal associated with a change in a relative position of a speaker and a microphone.

21. The computing system of claim 19 , wherein defining the model representative of the plurality of acoustic variations to the speech signal includes one or more of:

modeling the plurality of acoustic variations to the speech signal using a statistical distribution; and

modeling the plurality of acoustic variations to the speech signal using a mathematical model representative of the plurality of acoustic variations to the speech signal associated with a particular use-case scenario.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2021
From: NAYLOR, PATRICK A.; SHARMA, DUSHYANT; JOST, UWE HELMUT; GANONG III, WILLIAM F.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 055320/0047 →
Continuity (1)
Related Publication 20220262343A1 · Aug 18, 2022