IP Library Granted Patent US 10,096,318
Granted Patent B2
US 10,096,318 · App. 15/689,837 · Granted Oct 9, 2018

System and method of using neural transforms of robust audio features for speech processing

Inventors: Enrico Luigi Bocchieri (Chatham, NJ); Dimitrios Dimitriadis (Rutherford, NJ)
Assignee: NUANCE COMMUNICATIONS, INC.
G10L15/20G10L15/02G10L15/144G10L25/15G10L25/24G10L15/142G10L15/16G10L21/0208
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,096,318
App. No.
15/689,837
Granted
Oct 9, 2018
Kind
B2
Abstract

A system and method for processing speech includes receiving a first information stream associated with speech, the first information stream comprising micro-modulation features and receiving a second information stream associated with the speech, the second information stream comprising features. The method includes combining, via a non-linear multilayer perceptron, the first information stream and the second information stream to yield a third information stream. The system performs automatic speech recognition on the third information stream. The third information stream can also be used for training HMMs.

Claims (36)

1. A method comprising:

receiving, via a communication network, a first information stream associated with speech, wherein the first information stream comprises features modeled according to a first time scale;

receiving, via the communication network, a second information stream associated with the speech;

performing, via at least one hardware processor, automatic speech recognition on a third information stream formed by combining the first information stream and the second information stream, to yield a recognition result; and

outputting, via the communication network, the recognition result comprising text representing the speech.

2. The method of claim 1 , wherein the second information stream comprises cepstral features modeled in a second time scale.

3. The method of claim 2 , wherein the first time scale is distinct from the second time scale.

4. The method of claim 1 , further comprising filtering out noise from the third information stream prior to performing automatic speech recognition.

5. The method of claim 1 , wherein the third information stream comprises less features than raw features in the first information stream and the second information stream.

6. The method of claim 1 , further comprising training a Hidden Markov model using the third information stream.

7. A system comprising:

a processor; and

a non-transitory computer-readable storage medium storing instructions which, when executed by the processor, cause the processor to perform operations comprising:

receiving, via a communication network, a first information stream associated with speech, wherein the first information stream comprises features modeled according to a first time scale;

receiving, via the communication network, a second information stream associated with the speech;

performing automatic speech recognition on a third information stream formed by combining the first information stream and the second information stream, to yield a recognition result; and

outputting, via the communication network, the recognition result comprising text representing the speech.

8. The system of claim 7 , wherein the second information stream comprises cepstral features modeled in a second time scale.

9. The system of claim 8 , wherein the first time scale is distinct from the second time scale.

10. The system of claim 7 , wherein the non-transitory computer-readable storage medium stores additional instructions which, when executed by the processor, cause the processor to perform operations further comprising:

filtering out noise from the third information stream prior to performing automatic speech recognition.

11. The system of claim 7 , wherein the third information stream comprises less features than raw features in the first information stream and the second information stream.

12. The system of claim 7 , wherein the non-transitory computer-readable storage medium stores additional instructions which, when executed by the processor, cause the processor to perform operations further comprising:

training a Hidden Markov model using the third information stream.

13. A non-transitory computer-readable storage device storing instructions, which, when executed by a processor, cause the processor to perform operations comprising:

receiving, via a communication network, a first information stream associated with speech, wherein the first information stream comprises features modeled according to a first time scale;

receiving, via the communication network, a second information stream associated with the speech;

performing automatic speech recognition on a third information stream formed by combining the first information stream and the second information stream, to yield a recognition result; and

outputting, via the communication network, the recognition result comprising text representing the speech.

14. The non-transitory computer-readable storage device of claim 13 , wherein the second information stream comprises cepstral features modeled in a second time scale.

15. The non-transitory computer-readable storage device of claim 14 , wherein the first time scale is distinct from the second time scale.

16. The non-transitory computer-readable storage device of claim 13 , wherein the non-transitory computer-readable storage device stores additional instructions, which, when executed by the processor, cause the processor to perform operations further comprising:

filtering out noise from the third information stream prior to performing automatic speech recognition.

17. The non-transitory computer-readable storage device of claim 13 , wherein the third information stream comprises less features than raw features in the first information stream and the second information stream.

18. The non-transitory computer-readable storage device of claim 13 , wherein the non-transitory computer-readable storage device stores additional instructions, which, when executed by the processor, cause the processor to perform operations further comprising:

training a Hidden Markov model using the third information stream.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2020
From: BOCCHIERI, ENRICO LUIGI; DIMITRIADIS, DIMITRIOS
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 053543/0403 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2020
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 053555/0393 →
Continuity (3)
Continuation 15056000 · Feb 29, 2016
Continuation 14046393 · Oct 4, 2013
Related Publication 20170358298A1 · Dec 14, 2017