IP Library Granted Patent US 9,824,684
Granted Patent B2
US 9,824,684 · App. 14/578,938 · Granted Nov 21, 2017

Prediction-based sequence recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,824,684
App. No.
14/578,938
Granted
Nov 21, 2017
Kind
B2
Abstract

A sequence recognition system comprises a prediction component configured to receive a set of observed features from a signal to be recognized and to output a prediction output indicative of a predicted recognition based on the set of observed features. The sequence recognition system also comprises a classification component configured to receive the prediction output and to output a label indicative of recognition of the signal based on the prediction output.

Claims (69)

1. A computing system comprising:

at least one processor; and

memory storing instructions executable by the at least one processor, wherein the instructions configure the computing system to provide a prediction component, a classification component, and a synchronization component;

wherein the prediction component is configured to:

identify a speech signal to be recognized, the speech signal comprising at least a first frame and a second frame that is subsequent to the first frame in the signal, wherein the first frame comprises a first set of observed features and the second frame comprises a second set of observed features;

based on the first set of observed features, generate a prediction output indicative of a predicted phoneme label that is assigned to the second frame, the predicted phoneme label being indicative of a predicted phoneme for the second frame;

wherein the classification component is configured to:

receive the prediction output indicative of the predicted phoneme label assigned to the second frame;

receive the second set of observed features;

based on the predicted phoneme label and the second set of observed features, correct the predicted phoneme label assigned to the second frame; and

output a recognition result for the speech signal based on the phoneme label; and

wherein the synchronization component is configured to:

receive the prediction output from the prediction component;

perform a synchronization operation that implements a delay function on the prediction output, wherein the delay function synchronizes the prediction output and the second set of observed features received by the classification component; and

provide the prediction output to the classification component.

2. The computing system of claim 1 , wherein the classification component is configured to:

classify the signal by assigning the phoneme label to the second frame.

3. The computing system of claim 2 , wherein the classification component is configured to estimate a state posterior probability for the signal.

4. The computing system of claim 1 , wherein the first and second frames comprise non-contiguons frames in the signal.

5. The computing system of claim 1 , wherein the classification component is configured to:

receive the first set of observed features from the first frame;

generate feedback information based on classifying the first frame of the signal; and

generate the recognition result pertaining to the second frame based on the feedback information.

6. The computing system of claim 1 , wherein the prediction component comprises a first neural network and the classification component comprises a second neural network.

7. The computing system of claim 6 , wherein the predicted phoneme label is obtained from a bottleneck layer of the first neural network.

8. The computing system of claim 6 , wherein the classification component is configured to output feedback information to the first neural network, the feedback information being obtained from a bottleneck layer of the second neural network.

9. A computing system comprising:

at least one processor; and

memory storing instructions executable by the at least one processor, wherein the instructions, when executed, provide:

a sequence recognizer comprising:

a prediction component configured to:

receive a first set of observed features from a first frame of a speech signal to be recognized; and

base on the first set of observed features, generate a prediction output indicative of a predicted phoneme label for a second frame of the signal;

a synchronization component configured to:

receive the prediction output from the prediction component;

perform a synchronization operation that implements a delay function on the prediction output; and

provide the prediction output to the classification component; and

a classification component configured to:

receive a second set of observed features from the second frame of the speech signal:

receive the prediction output from the synchronization component, wherein the delay function synchronizes the prediction output and the second set of observed features received by the classification component;

based on the prediction output and the second set of observed features, correct the prediction output; and

output a recognition result based on the corrected recognition, the recognition result being indicative of a phoneme label assigned to the second frame; and

a training component configured to:

obtain labeled training data; and

apply the labeled training data as input to the prediction component and the classification component, to train the prediction component and classification component using a multi-objective training function that incorporates a prediction objective and a classification objective into an objective function.

10. The computing system of claim 9 , wherein the objective function is optimized by the training component.

11. A computing system comprising:

at least one processor; and

memory storing instructions executable by the at least one processor, wherein the instructions configure the computing system to provide a prediction component, a classification component, and a synchronization component;

wherein the prediction component is configured to:

receive a speech signal to be recognized, the speech signal comprising a plurality of frames, each frame having a set of observed features from a speech input; and

generate a prediction output based on the set of observed features for a first one of the frames, the prediction output being indicative of a predicted phoneme label for a second one of the frames that is temporally subsequent to the first frame in the speech signal; and

wherein the classification component is configured to:

receive the prediction output indicative of the predicted phoneme label for the second frame;

receive the set of observed features for the second frame;

based on the predicted phoneme label and the set of observed features for the second frame, estimate a state posterior probability for the second frame;

based on the estimated state posterior probability, generate a recognition output by assigning a state to the second frame;

generate feedback indicative of the generation of the recognition output; and

provide the feedback to the prediction component, wherein machine learning is performed on the prediction component, using the feedback, to dynamically adjust performance of the prediction component in generating a subsequent prediction output;

wherein the synchronization component is configured to:

receive the prediction output from the prediction component;

perform a synchronization operation on the prediction output;

implement a delay function on the prediction output, wherein the delay function synchronizes the prediction output and the set of observed features for the second frame received by the classification component; and

provide the prediction output to the classification component.

12. The computer-implemented method of claim 11 , wherein the prediction component is configured to:

based on the set of observed features for the first frame, predict a posterior probability of the phoneme label for the second frame; and

assign the predicted phoneme label to the second frame based on the predicted posterior probability, and

wherein generating a subsequent prediction output comprises generating a second prediction output based on the feedback and the set of observed features for the second frame, the second prediction output being indicative of a predicted phoneme label for a third one of the frames.

13. The computer-implemented method of claim 11 , wherein the first and second frames comprise non-contiguous frames in the speech signal.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2015
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034819/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2014
From: YU, DONG; ZHANG, YU; SELTZER, MICHAEL L; DROPPO, JAMES G
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034567/0718 →