IP Library Granted Patent US 10,134,393
Granted Patent B2
US 10,134,393 · App. 15/664,153 · Granted Nov 20, 2018

Generating representations of acoustic sequences

Inventors: Hasim Sak (New York, NY); Andrew W. Senior (New York, NY)
Assignee: Google LLC
G10L15/16G10L15/02G10L15/142G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,134,393
App. No.
15/664,153
Granted
Nov 20, 2018
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating representation of acoustic sequences. One of the methods includes: receiving an acoustic sequence, the acoustic sequence comprising a respective acoustic feature representation at each of a plurality of time steps; processing the acoustic feature representation at an initial time step using an acoustic modeling neural network; for each subsequent time step of the plurality of time steps: receiving an output generated by the acoustic modeling neural network for a preceding time step, generating a modified input from the output generated by the acoustic modeling neural network for the preceding time step and the acoustic representation for the time step, and processing the modified input using the acoustic modeling neural network to generate an output for the time step; and generating a phoneme representation for the utterance from the outputs for each of the time steps.

Claims (35)

1. A method comprising:

receiving an acoustic sequence, the acoustic sequence representing an utterance, and the acoustic sequence comprising a respective acoustic feature representation at each of a plurality of time steps; and

generating, for use in a speech recognition system, a phoneme representation for the utterance, comprising:

at a particular time step of the plurality of time steps:

combining the acoustic feature representation at the particular time step and a preceding output generated by an acoustic modeling neural network for a preceding time step to generate a modified input; and

processing the modified input using the acoustic modeling neural network to generate an output for the particular time step.

2. The method of claim 1 , wherein the acoustic modeling neural network is a feed-forward neural network.

3. The method of claim 1 , wherein the acoustic modeling neural network is a recurrent neural network.

4. The method of claim 3 , wherein the acoustic modeling neural network is a long short-term memory (LSTM) neural network.

5. The method of claim 1 , wherein the output for the acoustic feature representation at the particular time step is a set of scores for a set of phonemes or phoneme subdivisions, wherein the score for each phoneme or phoneme subdivision represents a likelihood that the phoneme or phoneme subdivision is a representation of the utterance at the particular time step.

6. The method of claim 5 , wherein generating the modified input comprises appending the set of scores for the preceding time step to the acoustic feature representation for the particular time step.

7. The method of claim 5 , wherein generating the modified input comprises appending data identifying a highest-scoring phoneme or phoneme subdivision according to the set of scores for the preceding time step to the acoustic feature representation for the particular time step.

8. The method of claim 5 , wherein the set of scores defines a probability distribution over a set of Hidden Markov Model (HMM) states.

9. The method of claim 1 , wherein phoneme representation for the utterance is generated from the outputs generated for each of the respective acoustic feature representations.

10. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers causes the one or more computers to perform operations comprising:

receiving an acoustic sequence, the acoustic sequence representing an utterance, and the acoustic sequence comprising a respective acoustic feature representation at each of a plurality of time steps; and

generating, for use in a speech recognition system, a phoneme representation for the utterance, comprising:

at a particular time step of the plurality of time steps:

combining the acoustic feature representation at the particular time step and a preceding output generated by an acoustic modeling neural network for a preceding time step to generate a modified input; and

processing the modified input using the acoustic modeling neural network to generate an output for the particular time step.

11. The system of claim 10 , wherein the acoustic modeling neural network is a feed-forward neural network.

12. The system of claim 10 , wherein the acoustic modeling neural network is a recurrent neural network.

13. The system of claim 12 , wherein the acoustic modeling neural network is a long short-term memory (LSTM) neural network.

14. The system of claim 10 , wherein the output for the acoustic feature representation at the particular time step is a set of scores for a set of phonemes or phoneme subdivisions, wherein the score for each phoneme or phoneme subdivision represents a likelihood that the phoneme or phoneme subdivision is a representation of the utterance at the particular time step.

15. The system of claim 14 , wherein generating the modified input comprises appending the set of scores for the preceding time step to the acoustic feature representation for the particular time step.

16. The system of claim 14 , wherein generating the modified input comprises appending data identifying a highest-scoring phoneme or phoneme subdivision according to the set of scores for the preceding time step to the acoustic feature representation for the particular time step.

17. A non-transitory computer storage medium encoded with a computer program, the computer program comprising instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving an acoustic sequence, the acoustic sequence representing an utterance, and the acoustic sequence comprising a respective acoustic feature representation at each of a plurality of time steps; and

generating, for use in a speech recognition system, a phoneme representation for the utterance, comprising:

at a particular time step of the plurality of time steps:

combining the acoustic feature representation at the particular time step and a preceding output generated by an acoustic modeling neural network for a preceding time step to generate a modified input; and

processing the modified input using the acoustic modeling neural network to generate an output for the particular time step.

18. The non-transitory computer storage medium of claim 17 , wherein the output for the acoustic feature representation at the particular time step is a set of scores for a set of phonemes or phoneme subdivisions, wherein the score for each phoneme or phoneme subdivision represents a likelihood that the phoneme or phoneme subdivision is a representation of the utterance at the particular time step.

19. The non-transitory computer storage medium of claim 18 , wherein generating the modified input comprises appending the set of scores for the preceding time step to the acoustic feature representation for the particular time step.

20. The non-transitory computer storage medium of claim 18 , wherein generating the modified input comprises appending data identifying a highest-scoring phoneme or phoneme subdivision according to the set of scores for the preceding time step to the acoustic feature representation for the particular time step.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2017
From: SAK, HASIM; SENIOR, ANDREW W.
To: GOOGLE INC.
Reel/Frame 043145/0938 →
Continuity (3)
Continuation 14559113 · Dec 3, 2014
Provisional Application 61917089 · Dec 17, 2013
Related Publication 20170330558A1 · Nov 16, 2017