IP Library › Granted Patent US 12,198,047
Granted Patent B2
US 12,198,047 · App. 17/122,894 · Granted Jan 14, 2025

Quasi-recurrent neural network based encoder-decoder model

Inventors: James Bradbury (San Francisco, CA); Stephen Joseph Merity (San Francisco, CA); Caiming Xiong (Palo Alto, CA); Richard Socher (Menlo Park, CA)
Assignee: Salesforce, Inc.
G06N3/08G06F17/16G06F40/216G06F40/30G06F40/44G06N3/04G06N3/044G06N3/045G06N3/10G06F40/00G10L15/16G10L15/18G10L15/1815G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,047
App. No.
17/122,894
Granted
Jan 14, 2025
Kind
B2
Abstract

The technology disclosed provides a quasi-recurrent neural network (QRNN) encoder-decoder model that alternates convolutional layers, which apply in parallel across timesteps, and minimalist recurrent pooling layers that apply in parallel across feature dimensions.

Claims (68)

1. A method for generating an output from a quasi-recurrent neural network (QRNN), the method including:

receiving a time series of input vectors;

passing the time series of input vectors to a plurality of QRNN layers, each QRNN layer including a respective convolution layer and a respective pooling layer;

generating, by each respective convolution layer, an activation vector from the input vectors through a hyperbolic tangent nonlinearity activation and one or more gate vectors though a convolutional filter bank;

applying, by each respective pooling layer, a QRNN fo-pooling operation to the one or more gate vectors by:

computing a current state vector based on a current forget gate vector from the one or more gate vectors, a previous state vector and the activation vector, and

computing a current hidden state vector by taking an elementwise multiplication of an output gate vector from the one or more gate vectors and the computed current state vector;

passing the computed current hidden state vector to a next QRNN layer from the plurality of QRNN layer; and

sequentially outputting an encoded hidden state vector from a last QRNN layer of the plurality of QRNN layers for each successive time series window among a plurality of times series windows corresponding to the time series of input vectors.

2. The method of claim 1 , wherein the one or more gate vectors correspond to a dimensionality that is augmented relative to dimensionality of the input vectors in dependence upon a number of convolutional filters in the convolutional filter bank.

3. The method of claim 1 , wherein the input vectors represent elements of an input sequence selected from the group consisting of a word-level sequence and a character-level sequence.

4. The method of claim 1 , wherein the one or more gate vectors comprises a forget gate vector, and

wherein the respective pooling layer uses a forget gate vector for a current time series window to control accumulation of information from the previous state vector accumulated for a prior time series window and information from an activation vector for the current time series window.

5. The method of claim 1 , wherein the one or more gate vectors comprises an input gate vector, and

wherein the respective pooling layer uses an input gate vector for a current time series window to control accumulation of information from an activation vector for the current time series window.

6. The method of claim 1 , wherein the one or more gate vectors comprises an output gate vector, and

wherein the pooling layer uses an output gate vector for a current time series window to control accumulation of information from the current state vector for the current time series window.

7. The method of claim 1 , further comprising:

receiving as input a preceding output generated by a preceding QRNN layer of the plurality of QRNN layers;

processing the preceding output through the respective convolutional layer to produce an alternative representation of the preceding output; and

processing the alternative representation through the respective pooling layer to produce an output.

8. The method of claim 1 , further comprising:

including skip connections between the plurality of QRNN layers,

wherein the skip connections concatenate output of a preceding QRNN layer with output of a current QRNN layer and provide a concatenation to a following layer as input.

9. The method of claim 1 , wherein the activation vector is generated by the respective convolutional layer by:

applying the convolutional filter bank to a number of time series windows that are aligned in parallel, corresponding to the time series of the input vectors;

concurrently outputting convolutional vectors in parallel corresponding to the number of the time series windows from the convolutional filter bank,

each of the convolution vectors comprising feature values in an activation vector and in one or more gate vectors, and

the feature values in the gate vectors are parameters that, respectively, apply element-wise by ordinal position to the feature values in the activation vector.

10. The method of claim 9 , wherein the current state vectors are generated by:

applying accumulators in parallel over feature values of a convolutional vector to concurrently accumulate ordinal position-wise, in the current state vector for a current time series window, an ordered set of feature sums for all ordinal positions in the current state vector in dependence upon a feature value at a given ordinal position in an activation vector outputted for the current time series window,

one or more feature values at the given ordinal position in one or more gate vectors outputted for the current time series window, and

a feature sum at the given ordinal position in the previous state vector accumulated for a prior time series window; and

sequentially outputting the current state vector for each successive time series window among the time series windows.

11. A system of a quasi-recurrent neural network (QRNN), the system including:

a memory storing parameters of the QRNN comprising an input layer, a plurality of QRNN layers and an output layer, and a plurality of processor-executable instructions; and

one or more processor executing the plurality of processor-executable instructions to perform operations comprising:

receiving at the input layer a time series of input vectors;

generating, at each respective convolution layer comprised in each QRNN layer, an activation vector from the input vectors through a hyperbolic tangent nonlinearity activation and one or more gate vectors though a convolutional filter bank,

applying, at each respective pooling layer comprised in the each QRNN layer, a QRNN fo-pooling operation to the one or more gate vectors by:

computing a current state vector based on a current forget gate vector from the one or more gate vectors, a previous state vector and the activation vector, and

computing a current hidden state vector by taking an elementwise multiplication of an output gate vector from the one or more gate vectors and the computed current state vector;

wherein the respective QRNN layer passes the computed current hidden state vector to a next QRNN layer from the plurality of QRNN layer; and

sequentially outputs, at the output layer, an encoded hidden state vector from a last QRNN layer of the plurality of QRNN layers for each successive time series window among a plurality of times series windows corresponding to the time series of input vectors.

12. The system of claim 11 , wherein the one or more gate vectors correspond to a dimensionality that is augmented relative to dimensionality of the input vectors in dependence upon a number of convolutional filters in the convolutional filter bank.

13. The system of claim 11 , wherein the input vectors represent elements of an input sequence selected from the group consisting of a word-level sequence and a character-level sequence.

14. The system of claim 11 , wherein the one or more gate vectors comprises a forget gate vector, and

wherein the respective pooling layer uses a forget gate vector for a current time series window to control accumulation of information from the previous state vector accumulated for a prior time series window and information from an activation vector for the current time series window.

15. The system of claim 11 , wherein the one or more gate vectors comprises an input gate vector, and

wherein the respective pooling layer uses an input gate vector for a current time series window to control accumulation of information from an activation vector for the current time series window.

16. The system of claim 11 , wherein the one or more gate vectors comprises an output gate vector, and

wherein the pooling layer uses an output gate vector for a current time series window to control accumulation of information from the current state vector for the current time series window.

17. The system of claim 11 , wherein each QRNN layer in the plurality of QRNN layers receives as input a preceding output generated by a preceding QRNN layer of the plurality of QRNN layers, processes the preceding output through the respective convolutional layer to produce an alternative representation of the preceding output, and processes the alternative representation through the respective pooling layer to produce an output.

18. The system of claim 11 , wherein the plurality of QRNN layers include skip connections between layers, and the skip connections concatenate output of a preceding QRNN layer with output of a current QRNN layer and provide a concatenation to a following layer as input.

19. The system of claim 11 , wherein the activation vector is generated by the respective convolutional layer by:

applying the convolutional filter bank to a number of time series windows that are aligned in parallel, corresponding to the time series of the input vectors;

concurrently outputting convolutional vectors in parallel corresponding to the number of the time series windows from the convolutional filter bank,

each of the convolution vectors comprising feature values in an activation vector and in one or more gate vectors, and

the feature values in the gate vectors are parameters that, respectively, apply element-wise by ordinal position to the feature values in the activation vector.

20. A non-transitory processor-readable medium storing processor-executable instructions for generating an output from a quasi-recurrent neural network (QRNN), the instructions executable by a processor to perform operations comprising:

receiving a time series of input vectors;

passing the time series of input vectors to a plurality of QRNN layers, each QRNN layer including a respective convolution layer and a respective pooling layer;

generating, by each respective convolution layer, an activation vector from the input vectors through a hyperbolic tangent nonlinearity activation and one or more gate vectors though a convolutional filter bank;

applying, by each respective pooling layer, a QRNN fo-pooling operation to the one or more gate vectors by:

computing a current state vector based on a current forget gate vector from the one or more gate vectors, a previous state vector and the activation vector, and

computing a current hidden state vector by taking an elementwise multiplication of an output gate vector from the one or more gate vectors and the computed current state vector;

passing the computed current hidden state vector to a next QRNN layer from the plurality of QRNN layer; and

sequentially outputting an encoded hidden state vector from a last QRNN layer of the plurality of QRNN layers for each successive time series window among a plurality of times series windows corresponding to the time series of input vectors.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2024
From: BRADBURY, JAMES; MERITY, STEPHEN JOSEPH; XIONG, CAIMING; SOCHER, RICHARD
To: SALESFORCE.COM, INC.
Reel/Frame 068516/0833 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2021
From: BRADBURY, JAMES; MERITY, STEPHEN JOSEPH; XIONG, CAIMING; SOCHER, RICHARD
To: SALESFORCE.COM, INC.
Reel/Frame 055463/0092 →
Continuity (4)
Continuation 15420801 · Jan 31, 2017
Provisional Application 62417333 · Nov 4, 2016
Provisional Application 62418075 · Nov 4, 2016
Related Publication 20210103816A1 · Apr 8, 2021
References Cited (80)
US 6128606A · Bengio et al. · 2000 [cited by applicant]
US 10282663B2 · Socher et al. · 2019 [cited by applicant]
US 10346721B2 · Albright et al. · 2019 [cited by applicant]
US 20020038294A1 · Matsugu · 2002 [cited by applicant]
US 20110218950A1 · Mirowski et al. · 2011 [cited by applicant]
US 20140079297A1 · Tadayon et al. · 2014 [cited by applicant]
US 20140229164A1 · Martens et al. · 2014 [cited by applicant]
US 20160099010A1 · Sainath et al. · 2016 [cited by applicant]
US 20160180215A1 · Vinyals et al. · 2016 [cited by applicant]
US 20160321784A1 · Annapureddy · 2016 [cited by applicant]
US 20160350653A1 · Socher et al. · 2016 [cited by applicant]
US 20170024645A1 · Socher et al. · 2017 [cited by applicant]
US 20170032280A1 · Socher · 2017 [cited by applicant]
US 20180129931A1 · Bradbury et al. · 2018 [cited by applicant]
CN 1674153A · 2005 [cited by applicant]
CN 105868829A · 2016 [cited by applicant]
WO 2016160237A1 · 2016 [cited by applicant]
WO 2018085724A1 · 2018 [cited by applicant]
Examiner's Report for Canadian Application No. 3,040,153 dated Feb. 22, 2021, 4 pages. [cited by applicant]
Povey, D., et al., “Parallel training of DNNs with natural gradient and parameter averaging”, 2014, arXiv:1410.7455, <URL: https://arxiv.org/pdf/1410.7455.pdf>, 28 pages. [cited by applicant]
James Bradbury et al: “Quasi-Recurrent Neural Networks”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 5, 2016 (Nov. 5, 2016), XP080729557. [cited by applicant]
Bahdanau et al., “Neural Machine Translation by Jointly Learning to Align and Translate,” Published as a conference paper at International Conference on Learning Representations (ICLR) 2015 (May 19, 2016) pp. 1.15 arXiv… [cited by applicant]
Balduzzi et al., Strongly-Typed Recurrent Neural Networks, Proceedings of the 33rd International Conference on Machine Learning, JMLR: W&CP, vol. 48 (May 24, 2016) pp. 1-10 arxiv.org/abs/1602.02218v2. [cited by applicant]
Bradbury et al., “MetaMind Neural Machine Translation System for WMT 2016,” Proceedings of the First Conference on Machine Translation, vol. 2: Shared Task Papers, pp. 264-267, Berlin, Germany (Aug. 11-12, 2016) pp. 1-4… [cited by applicant]
Cho et al., “Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation,” Conference on Empirical Methods in Natural Language Processing (Sep. 3, 2014) pp. 1-15 arXiv:1406.1078v3. [cited by applicant]
Chopra et al., “Abstractive Sentence Summarization with Attentive Recurrent Neural Networks,” Proceedings of NAACL-HLT 2016, Conference of the North American Chapter of the Association for Computational Linguistics: Hum… [cited by applicant]
Chung et al., “Gated Feedback Recurrent Neural Networks” (Nov. 17, 2015) pp. 1-9 efarXiv:1502.02367v4. [cited by applicant]
Gal et al., “A Theoretically Grounded Application of Dropout in Recurrent Neural Networks,” 30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain. (Jan. 1, 2016) pp. 1-9 arxiv.org/abs/15… [cited by applicant]
Greff et al., “LSTM: A Search Space Odyssey,” IEEE Transactions on Neural Networks and Learning Systems, vol. 28, Issue: 10 (Oct. 10, 2017 ) pp. 2222-2232 pp. 1-12 arXiv:1503.04069v2. [cited by applicant]
Hochreiter et al., “Long Short-Term Memory,” Neural Computation 9(8) (Jan. 1, 1997) pp. 1735-1780 pp. 1-32 www7.informatik.tu-muenchen.de/˜hochreit. [cited by applicant]
Huang et al., “Densely Connected Convolutional Networks” (Jan. 28, 2018) pp. 1-9, arXiv:1608.06993v5. [cited by applicant]
Johnson et al., “Effective Use of Word Order for Text Categorization with Convolutional Neural Networks,” to appear in North American Chapter of the Association for Computational Linguistics HLT 2015 (Mar. 26, 2015) pp.… [cited by applicant]
Kalchbrenner et al., “Neural Machine Translation in Linear Time” Google Deepmind, London UK (Mar. 15, 2017) pp. 1-9 arXiv:1610.10099. [cited by applicant]
Kim et al., “Character-Aware Neural Language Models,” Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI-16), Association for the Advancement of Artificial Intelligence (Dec. 1, 2015) pp. 2741… [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization,” Published as a conference paper at International Conference on Learning Representations (ICLR) (Jan. 30, 2017) pp. 1-15 arXiv:1412.6980. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” Conference on Neural Information Processing Systems (Jan. 1, 2012) pp. 1-9 http://papers.nips.cc/paper/4824-imagenet-classification-w… [cited by applicant]
Krueger et al., “Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations,” Under review as a conference paper at International Conference on Learning Representations (ICLR) 2017 (Sep. 22, 2017) pp. 1-11 arX… [cited by applicant]
Kumar et al. “Ask Me Anything: Dynamic Memory Networks for Natural Language Processing,” [email protected], MetaMind, Palo Alto, CA USA (Mar. 5, 2016) pp. 1-10 arXiv:1506.07285v5. [cited by applicant]
Lee et al., “Fully Character-Level Neural Machine Translation Without Explicit Segmentation,” Transactions of the Association for Computational Linguistics (TACL) 2017 (Jun. 13, 2017) pp. 1-13 arXiv:1610.03017v3. [cited by applicant]
Longpre et al., “A Way Out of the Odyssey: Analyzing and Combining Recent Insights for LSTMS,” Under review as a conference paper at International Conference on Learning Representations (ICLR) 2017 (Dec. 17, 2016) pp. 1… [cited by applicant]
Luong et al., “Effective Approaches to Attention-Based Neural Machine Translation,” Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lisbon, Portugal. (Sep. 20, 2015) pp. 1-11 arXi… [cited by applicant]
Maas et al., “Multi-Dimensional Sentiment Analysis with Learned Representations” Technical Report (2011) pp. 1-10. [cited by applicant]
Merity et al., “Pointer Sentinel Mixture Models,” MetaMind—A Salesforce Company, Palo Alto, CA (Sep. 26, 2016) pp. 1-13 arXiv:1609.07843. [cited by applicant]
Mesnil et al., “Ensemble of Generative and Discriminative Techniques for Sentiment Analysis of Movie Reviews” Accepted as a workshop contribution at International Conference on Learning Representations (ICLR) 2015 (May … [cited by applicant]
Mikolov et al., “Recurrent Neural Network Based Language Model,” Conference: INTERSPEECH 2010, 11th Annual Conference of the International Speech Communication Association, (Jul. 20, 2010) pp. 1-25 https://www.researchg… [cited by applicant]
Miyato et al., “Adversarial Training Methods for Semi-Supervised Text Classification,” Published as a conference paper at ICLR 2017, NIPS 2016, Deep Learning Symposium recommendation readers (May 6, 2017) pp. 1-11 arXiv… [cited by applicant]
Nallapati et al., “Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond,” Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning (Aug. 26, 2016) pp. 1-12 SecarXiv:1602.… [cited by applicant]
Pennington et al., “GloVe: Global Vectors for Word Representation”, Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar (Oct. 25-29, 2014) pp. 1532-1543, pp. 1-12 … [cited by applicant]
Ranzato et al., Sequence Level Training With Recurrent Neural Networks, Published as a Conference Paper at International Conference on Learning Representations (ICLR) 2016, (May 6, 2016) pp. 1-16 arXiv:1511.06732. [cited by applicant]
Tieleman et al. “Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude,” Coursera: Neural Networks for Machine Learning (Jan. 1, 2012) pp. 1-4. [cited by applicant]
Tokui et al., “Chainer: A Next-Generation Open Source Framework for Deep Learning,” Preferred Networks America. San Mateo, CA. (Jan. 1, 2015) pp. 1-6 [email protected]. [cited by applicant]
Van den Oord et al., “Pixel Recurrent Neural Networks,” JMLR: W&CP, vol. 48, Proceedings of the 33rd International Conference on Machine, New York, NY (Aug. 19, 2016) pp. 1-11 arXiv:1601.06759. [cited by applicant]
Wang et al., “Baselines and Bigrams: Simple, Good Sentiment and Topic Classification,” Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics, pp. 90-94, pp. 1-5 http://www.aclweb.org/an… [cited by applicant]
Wang et al., “Predicting Polarities of Tweets by Composing Word Embeddings with Long Short-Term Memory,” Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International … [cited by applicant]
Wiseman et al., “Sequence-to-Sequence Learning as Beam-Search Optimization,” 2016 Conference on Empirical Methods in Natural Language Processing (Nov. 10, 2016) pp. 1-11 arXiv:1606.02960. [cited by applicant]
Wu et al., “Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation,” arXiv preprint arXiv:1609.08144, pp. 1-23 (Oct. 8, 2016). [cited by applicant]
Xiao et al., “Efficient Character-level Document Classification by Combining Convolution and Recurrent Layers,” (Feb. 1, 2016) pp. 1-10 arXiv:1602.00367. [cited by applicant]
Xiong et al., “Dynamic Memory Networks for Visual and Textual Question Answering,” Proceedings of the 33rd International Conference on Machine Learning, New York, NY, USA, JMLR: W&CP vol. 48, (Mar. 4, 2016) pp. 1-10 arX… [cited by applicant]
Zaremba et al., “Recurrent Neural Network Regularization,” Under review as a conference paper at International Conference on Learning Representations (ICLR) 2015 (Feb. 19, 2015) pp. 1-8 arXiv:1409.2329. [cited by applicant]
Zhang et al., “Character-level Convolutional Networks for Text Classification,” Conference on Neural Information Processing Systems (Sep. 4, 2015) pp. 1-9 arXiv:1509.01626. [cited by applicant]
Zhou et al., “A C-LSTM Neural Network for Text Classification” (Nov. 30, 2015) pp. 1-11 arXiv:1511.08630v2. [cited by applicant]
International Search Report issued by ISA/EP for PCT/US2017/060051 on Jan. 26, 2018; pp. 1-4. [cited by applicant]
Written Opinion of the International Searching Authority issued by ISA/EP for PCT/US2017/060051 on Jan. 26, 2018; pp. 1-8. [cited by applicant]
International Search Report issued by ISA/EP for International Application No. PCT/US2017/060049 on Jan. 31, 2018; pp. 1-3. [cited by applicant]
Written Opinion of the International Searching Authority issued by ISA/EP for PCT/US2017/060049 on Jan. 31, 2018; pp. 1-8. [cited by applicant]
Intemational Preliminary Report on Patentability for PCT/US2017/060049, dated May 7, 2019, pp. 1-9. [cited by applicant]
Intemational Preliminary Report on Patentability for PCT/US2017/060051, dated May 7, 2019, pp. 1-9. [cited by applicant]
Chrupala et al., “Learning Language through Pictures,” arXiv:1506.03694v2, Jun. 19, 2015, pp. 1-10. [cited by applicant]
Examination Report from Australian Patent Application No. 2017355537, Feb. 10, 2020, pp. 1-5. [cited by applicant]
Graves et al., “Speech Recognition with Deep Recurrent Neural Networks,” IEEE, 2013, pp. 6645-6649. [cited by applicant]
Kalchbrenner, A Convolutional Neural Network for Modeling Sentences, arXiv:1404.2188v1, Apr. 8, 2014, pp. 1-11. [cited by applicant]
Karpathy et al., “Large-scale Video Classification with Convolutional Neural Networks” 2014 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1725-1732, 2014. [cited by applicant]
Lai et al., “Recurrent Convolutional Neural Networks for Text Classification,” Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, pp. 2267-2273, 2015. [cited by applicant]
Liu et al., “Parallel Training of Convolutional Neural Networks for Small Sample Learning,” IEEE, 2015, pp. 1-6. [cited by applicant]
Lu et al., “Hierarchical Question-Image Co-Attention for Visual Question Answering,” arXiv:1606.00061v3, pp. 1-11, 2016. [cited by applicant]
Meng et al., “Encoding Source Language with Convolutional Neural Network for Machine Translation,” arXiv:1503.01838v5, pp. 1-12, 2015. [cited by applicant]
Office Action, including English translation, from Japan Patent Application No. 2019-522910, dated Sep. 8, 2020, pp. [cited by applicant]
Sak et al., “Long Short-Term Memory based Recurrent Neural Network Architectures for Large Vocabulary Speech Recognition,” arXiv:1402.1128v1, Feb. 5, 2014, pp. 1-5. [cited by applicant]
Van Den Oord et al., “Conditional Image Generation with PixelCNN Decoders,” arXiv: 1606.05328v2, 2016, pp. 1-13. [cited by applicant]
Zuo et al., “Convolutional Recurrent Neural Networks: Learning Spatial Dependencies for Image Representation,” Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CV PRW), 2015,… [cited by applicant]