IP Library › Granted Patent US 12,670,356
Granted Patent B2
US 12,670,356 · App. 17/643,736 · Granted Jun 30, 2026

Generating representations of input sequences using neural networks

Inventors: Oriol Vinyals (London, GB); Quoc V. Le (Sunnyvale, CA); Ilya Sutskever (San Francisco, CA)
Assignee: Google LLC
G06N3/02G06F40/40G06N3/0442G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,356
App. No.
17/643,736
Filed
Dec 10, 2021
Granted
Jun 30, 2026
Kind
B2
Examiner
CHEN, ALAN S
Art Unit
2125
USPC
706/15
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating representations of input sequences. One of the methods includes obtaining an input sequence, the input sequence comprising a plurality of inputs arranged according to an input order; processing the input sequence using a first long short term memory (LSTM) neural network to convert the input sequence into an alternative representation for the input sequence; and processing the alternative representation for the input sequence using a second LSTM neural network to generate a target sequence for the input sequence, the target sequence comprising a plurality of outputs arranged according to an output order.

Claims (36)

1 . A method performed by one or more computers, the method comprising:

obtaining an input sequence, the input sequence comprising a plurality of inputs arranged according to an input order;

processing the input sequence with an encoder network to convert the input sequence into an alternative representation of the input sequence;

processing the alternative representation of the input sequence with a decoder network in accordance with an initial version of a hidden state of the decoder network to predict a first output of a target sequence for the input sequence, wherein the target sequence comprises a plurality of outputs arranged according to an output order, wherein the hidden state of the decoder network is updated as a result of processing the alternative representation for the input sequence;

for each output of the target sequence after the first output, processing a preceding output that was predicted for a preceding position of the target sequence with the decoder network in accordance with a current version of the hidden state of the decoder network to predict a next output of the target sequence, wherein the hidden state of the decoder network is updated as a result of processing each preceding output to predict each next output of the target sequence, wherein the alternative representation of the input sequence is processed with the decoder network just once to predict only the first output of the target sequence.

2 . The method of claim 1 , wherein the input sequence is a variable length input sequence.

3 . The method of claim 2 , wherein the alternative representation is a vector of fixed dimensionality.

4 . The method of claim 1 , comprising:

adding an end-of-sentence token to the end of the input sequence to generate a modified input sequence; and

processing the modified input sequence with the encoder network to generate the alternative representation.

5 . The method of claim 1 , wherein processing the alternative representation of the input sequence with the decoder network comprises initializing the hidden state of the decoder network to the alternative representation of the input sequence.

6 . The method of claim 1 , wherein processing the alternative representation of the input sequence with the decoder network comprises using a left to right beam search decoding technique.

7 . The method of claim 1 , comprising training the encoder network and the decoder network using Stochastic Gradient Descent.

8 . The method of claim 1 , wherein the input sequence is a sequence of words in a first language and the target sequence is a translation of the sequence of words into a second language.

9 . The method of claim 1 , wherein the input sequence is a sequence of words and the target sequence is an autoencoding of the input sequence.

10 . The method of claim 1 , wherein the input sequence is a sequence of graphemes and the target sequence is a phoneme representation of the sequence of graphemes.

11 . The method of claim 1 , wherein the encoder network comprises a first long short-term memory (LSTM) neural network and the decoder network comprises a second LSTM neural network.

12 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:

obtaining an input sequence, the input sequence comprising a plurality of inputs arranged according to an input order;

processing the input sequence with an encoder network to convert the input sequence into an alternative representation of the input sequence;

processing the alternative representation of the input sequence with a decoder network in accordance with an initial version of a hidden state of the decoder network to predict a first output of a target sequence for the input sequence, wherein the target sequence comprises a plurality of outputs arranged according to an output order, wherein the hidden state of the decoder network is updated as a result of processing the alternative representation for the input sequence;

for each output of the target sequence after the first output, processing a preceding output that was predicted for a preceding position of the target sequence with the decoder network in accordance with a current version of the hidden state of the decoder network to predict a next output of the target sequence, wherein the hidden state of the decoder network is updated as a result of processing each preceding output to predict each next output of the target sequence, wherein the alternative representation of the input sequence is processed with the decoder network just once to predict only the first output of the target sequence.

13 . The system of claim 12 , wherein the input sequence is a variable length input sequence.

14 . The system of claim 13 , wherein the alternative representation is a vector of fixed dimensionality.

15 . The system of claim 12 , wherein the operations comprise:

adding an end-of-sentence token to the end of the input sequence to generate a modified input sequence; and

processing the modified input sequence with the encoder network to generate the alternative representation.

16 . The system of claim 12 , wherein processing the alternative representation of the input sequence with the decoder network comprises initializing the hidden state of the decoder network to the alternative representation of the input sequence.

17 . The system of claim 12 , wherein the input sequence is a sequence of words in a first language and the target sequence is a translation of the sequence of words into a second language.

18 . The system of claim 12 , wherein the input sequence is a sequence of words and the target sequence is an autoencoding of the input sequence.

19 . The system of claim 12 , wherein the input sequence is a sequence of graphemes and the target sequence is a phoneme representation of the sequence of graphemes.

20 . A computer program product encoded on one or more non-transitory storage media, the computer program product comprising instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:

obtaining an input sequence, the input sequence comprising a plurality of inputs arranged according to an input order;

processing the input sequence with an encoder network to convert the input sequence into an alternative representation of the input sequence;

processing the alternative representation of the input sequence with a decoder network in accordance with an initial version of a hidden state of the decoder network to predict a first output of a target sequence for the input sequence, wherein the target sequence comprises a plurality of outputs arranged according to an output order, wherein the hidden state of the decoder network is updated as a result of processing the alternative representation for the input sequence;

for each output of the target sequence after the first output, processing a preceding output that was predicted for a preceding position of the target sequence with the decoder network in accordance with a current version of the hidden state of the decoder network to predict a next output of the target sequence, wherein the hidden state of the decoder network is updated as a result of processing each preceding output to predict each next output of the target sequence, wherein the alternative representation of the input sequence is processed with the decoder network just once to predict only the first output of the target sequence.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2021
From: VINYALS, ORIOL; LE, QUOC V.; SUTSKEVER, ILYA
To: GOOGLE INC.
Reel/Frame 058363/0653 →
ENTITY CONVERSION Recorded Dec 10, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 058479/0141 →
Continuity (4)
Continuation 16211635 · Dec 6, 2018
Continuation 14731326 · Jun 4, 2015
Provisional Application 62009121 · Jun 6, 2014
Related Publication 20220101082A1 · Mar 31, 2022
References Cited (37)
US 8682643B1 · Hafez · 2014 [cited by applicant]
US 9484023B2 · Arisoy · 2016 [cited by applicant]
US 10181098B2 · Vinyals · 2019 [cited by applicant]
US 20040002848A1 · Zhou et al. · 2004 [cited by applicant]
US 20060136193A1 · Lux-Pogodalla et al. · 2006 [cited by applicant]
US 20140229158A1 · Zweig · 2014 [cited by applicant]
CN 1387651 · 2002 [cited by applicant]
CN 1475907 · 2004 [cited by applicant]
CN 101077011 · 2007 [cited by applicant]
EP 0094293 · 1983 [cited by applicant]
EP 0875832 · 1998 [cited by applicant]
WO WO200137128 · 2001 [cited by applicant]
Sutskever et al., Sequence to Sequence Learning with Neural Networks, arXiv:1409.3215v1 [cs. CL] Sep. 10, 2014; Total Pages: 10 (Year: 2014). [cited by examiner]
Jordan B. Pollack, Recursive Distributed Representations, Elsevier Science Publishers B.V. (North-Holland), Artificial Intelligence 46 (1990)l; pp. 77-105 (Year: 1990). [cited by examiner]
Rao et al., Grapheme-To-Phoneme Conversion Using Long Short-Term Memory Recurrent Neural Networks, ICASSP Conference, Apr. 2015; pp. 4225-4229 (Year: 2015). [cited by examiner]
Cho et al., Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation, arXiv:1406.1078v1 [cs.CL] Jun. 3, 2014; Total Pages: 14 (Year: 2014). [cited by examiner]
Bengio et al., “Neural probabilistic language models,” In Innovations in Machine Learning, vol. 194, pp. 137-186. Springer, 2006. [cited by applicant]
Cho et al. “Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation,” arXiv preprint arXiv Jun. 3, 2014, 14 pages. [cited by applicant]
CN Office Action in Chinese Application No. 2015104264010.8, dated Nov. 3, 2020, 11 pages (with English translation). [cited by applicant]
CN Office Action in Chinese Appln. No. 201510426401.8, dated Feb. 3, 2020, 19 pages (with English translation). [cited by applicant]
CN Office Action in Chinese Appln. No. 201510426401.8, dated May 20, 2019, 17 pages (with English translation). [cited by applicant]
CN Office Action issued in Chinese Application No. 201510426401.8, mailed on Jul. 30, 2018, 13 pages (English Translation). [cited by applicant]
EP Office Action in European Application No. 15170815.3, dated Apr. 23, 2020, 8 pages. [cited by applicant]
EP Office Action in European Application No. 20201300.9, dated Feb. 11, 2021, 9 pages. [cited by applicant]
EP Office Action issued in European Application No. 15170815.3, mailed on Jul. 10, 2018, 8 pages. [cited by applicant]
EP Result of consultation in European Appln. No. 15170815.3, dated Feb. 19, 2020, 4 pages. [cited by applicant]
EP Summons to attend oral proceedings pursuant to Rule 115(1) EPC in European Appln. No. 15170815.3, dated Sep. 5, 2019, 9 pages. [cited by applicant]
Extended European Search Report in European Application No. 15170815.3-1951/2953065, mailed on Nov. 8, 2016, 12 pages. [cited by applicant]
Graves, “Generating sequences with recurrent neural networks,” arXiv:1308.0850v5 [cs.NE], Jun. 2014, pp. 1-43. [cited by applicant]
Graves. “Sequence Transduction with Recurrent Neural Networks,” arXiv preprint arXiv, Nov. 14, 2012, 9 pages. [cited by applicant]
Hermann and Blunsom, “Multilingual distributed representations without word alignment,” In ICLR, 2014, Mar. 2014, pp. 1-9. [cited by applicant]
Hochreiter and Schmidhuber, “Long Short-Term Memory,” Neural Computation 9(8):1735-1780, 1997. [cited by applicant]
Mikolov et al., “Extensions of recurrent neural network language model,” In ICASSP, May 2011, pp. 5528-5531. [cited by applicant]
Mikolov et al., “Recurrent neural network based language model,” In Interspeech, pp. 1045-1048, Sep. 2010. [cited by applicant]
Qi et al. “RBM-Based Phoneme Recognition by Deep Neural Network Based on RBM,” Journal of Information Engineering University, vol. 14(5), 6 pages (with English Abstract). [cited by applicant]
Rumelhart et al., “Learning representations by back-propagating errors,” Nature, 323(6088):533-536, Oct. 1986. [cited by applicant]
Socher et al. “Dynamic Pooling and Unfolding Recursive Autoencoders for Paraphrase Detection,” Advances in neural information processing systems, vol. 24, Jan. 1, 2011, 9 pages. [cited by applicant]