IP Library › Granted Patent US 11,996,088
Granted Patent B2
US 11,996,088 · App. 16/918,669 · Granted May 28, 2024

Setting latency constraints for acoustic models

Inventors: Andrew W. Senior (New York, NY); Hasim Sak (Santa Clara, CA); Kanury Kanishka Rao (Santa Clara, CA)
Assignee: Google LLC
G10L15/16G06N3/044G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,996,088
App. No.
16/918,669
Granted
May 28, 2024
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media for acoustic modeling of audio data. One method includes receiving audio data representing a portion of an utterance, providing the audio data to a trained recurrent neural network that has been trained to indicate the occurrence of a phone at any of multiple time frames within a maximum delay of receiving audio data corresponding to the phone, receiving, within the predetermined maximum delay of providing the audio data to the trained recurrent neural network, output of the trained neural network indicating a phone corresponding to the provided audio data using output of the trained neural network to determine a transcription for the utterance, and providing the transcription for the utterance.

Claims (44)

1. A method of training a neural network, the method comprising:

receiving, at data processing hardware, training data comprising training audio data for an utterance and a sequence of phone labels identifying respective phones that occur in the utterance;

generating, by the data processing hardware, using an acoustic model, a reference alignment indicating ground truth audio frames that the phone labels in the sequence of phone labels occur in a sequence of audio frames representing the training audio data;

for each phone label, setting, by the data processing hardware, a respective constrained range of audio frames in the sequence of audio frames that ends a predetermined number of audio frames after a respective ground truth ending audio frame that the respective phone label last occurs in the reference alignment;

providing, by the data processing hardware, the training audio data to the neural network executing on the data processing hardware, the neural network configured to:

receive, as input, each audio frame in the sequence of audio frames representing the training audio data; and

generate, as output, a sequence of output labels indicating the occurrence of each respective phone;

determining, by the data processing hardware, for each particular output label of the sequence of output labels, a corresponding delay between output of the particular output label and a corresponding respective constrained range of audio frames; and

updating, by the data processing hardware, using the sequence of output labels generated and the corresponding delays, parameters of the neural network by applying a penalty when a corresponding delay satisfies a threshold constraint.

2. The method of claim 1 , further comprising, prior to generating the reference alignment, splitting, by the data processing hardware, the training audio data into a sequence of audio frames, each audio frame in the sequence of audio frames corresponding to a different respective time period of the training audio data.

3. The method of claim 2 , wherein sequence of audio frames comprise a sequence of fixed-length audio frames.

4. The method of claim 2 , wherein sequence of audio frames overlap.

5. The method of claim 2 , wherein splitting the training audio data into a sequence of audio frames further comprises generating a corresponding acoustic feature representation for each audio frame in the sequence of audio frames.

6. The method of claim 1 , wherein the neural network comprises a recurrent neural network having a convolutional layer, one or more long short-term memory layers, and a deep neural network.

7. The method of claim 6 , wherein:

providing the training audio data to the neural network comprises providing the training audio data as input to the convolutional layer;

output of the convolutional layer is provided as input to the one or more long short-term memory layers; and

output of the one or more long short-term memory layers is provided as input to the deep neural network.

8. The method of claim 6 , wherein the deep neural network comprises a connectionist temporal classification (CTC) output layer configured to generate a set of output scores for each audio frame in the sequence of audio frames received as input to the neural network.

9. The method of claim 8 , wherein each score in the set of scores indicates a likelihood that a particular phone in set of phones has occurred in the sequence of audio frames received as input to the neural network.

10. The method of claim 1 , wherein a number of output labels in the sequence of output labels is equal to a number of audio frames in the sequence of audio frames received as input to the neural network.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving training data comprising training audio data for an utterance and a sequence of phone labels identifying respective phones that occur in the utterance;

generating, using an acoustic model, a reference alignment indicating ground truth audio frames that the phone labels in the sequence of phone labels occur in a sequence of audio frames representing the training audio data;

for each phone label, setting a respective constrained range of audio frames in the sequence of audio frames that ends a predetermined number of audio frames after a respective ground truth ending audio frame that the respective phone label last occurs in the reference alignment;

providing the training audio data to the neural network executing on the data processing hardware, the neural network configured to:

receive, as input, each audio frame in the sequence of audio frames representing the training audio data; and

generate, as output, a sequence of output labels indicating the occurrence of each respective phone;

determining, by the data processing hardware, for each particular output label of the sequence of output labels, a corresponding delay between output of the particular output label and a corresponding respective constrained range of audio frames; and

updating, by the data processing hardware, using the sequence of output labels generated and the corresponding delays, parameters of the neural network by applying a penalty when a corresponding delay satisfies a threshold constraint.

12. The system of claim 11 , wherein the operations further comprise, prior to generating the reference alignment, splitting the training audio data into a sequence of audio frames, each audio frame in the sequence of audio frames corresponding to a different respective time period of the training audio data.

13. The system of claim 12 , wherein sequence of audio frames comprise a sequence of fixed-length audio frames.

14. The system of claim 12 , wherein sequence of audio frames overlap.

15. The system of claim 12 , wherein splitting the training audio data into a sequence of audio frames further comprises generating a corresponding acoustic feature representation for each audio frame in the sequence of audio frames.

16. The system of claim 11 , wherein the neural network comprises a recurrent neural network having a convolutional layer, one or more long short-term memory layers, and a deep neural network.

17. The system of claim 16 , wherein:

providing the training audio data to the neural network comprises providing the training audio data as input to the convolutional layer;

output of the convolutional layer is provided as input to the one or more long short-term memory layers; and

output of the one or more long short-term memory layers is provided as input to the deep neural network.

18. The system of claim 16 , wherein the deep neural network comprises a connectionist temporal classification (CTC) output layer configured to generate a set of output scores for each audio frame in the sequence of audio frames received as input to the neural network.

19. The system of claim 18 , wherein each score in the set of scores indicates a likelihood that a particular phone in set of phones has occurred in the sequence of audio frames received as input to the neural network.

20. The system of claim 11 , wherein a number of output labels in the sequence of output labels is equal to a number of audio frames in the sequence of audio frames received as input to the neural network.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AS 16918661 TO THE CORRECT APPLICATION NUMBER 16918669 PREVIOUSLY RECORDED ON REEL 053215 FRAME 0966. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jul 16, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 053232/0476 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AS 16918661 TO THE CORRECT APPLICATION NUMBER 16918669 PREVIOUSLY RECORDED ON REEL 053213 FRAME 0528. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jul 16, 2020
From: SENIOR, ANDREW W; SAK, HASIM; RAO, KANURY KANISHKA
To: GOOGLE LLC
Reel/Frame 053232/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2020
From: SENIOR, ANDREW W.; SAK, HASIM; RAO, KANURY KANISHKA
To: GOOGLE INC.
Reel/Frame 053213/0528 →
CONVERSION Recorded Jul 15, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 053215/0966 →
Continuity (2)
Continuation 14879225 · Oct 9, 2015
Related Publication 20200335093A1 · Oct 22, 2020