IP Library › Granted Patent US 12,572,793
Granted Patent B1
US 12,572,793 · App. 17/068,691 · Granted Mar 10, 2026

Subroutine neural networks

Inventors: Yujun Yan (Ann Arbor, MI); Kevin Jordan Swersky (Toronto, CA); Milad Olia Hashemi (San Francisco, CA)
Assignee: Google LLC
G06N3/08G06N3/045G06N3/04G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,793
App. No.
17/068,691
Granted
Mar 10, 2026
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating an output sequence from an input sequence. In one aspect, one of the systems includes one or more subroutine neural networks each comprising: an encoder neural network configured to receive an encoder network input comprising i) a subroutine input sequence of the subroutine neural network and ii) an input mask that masks one or more subroutine input elements of the subroutine input sequence and to generate an encoded representation of the encoder network input; a decoder neural network configured to receive the encoded representation and to generate the subroutine network output; and a masking neural network configured to generate an output mask, wherein the output mask will be used as the input mask of the encoder neural network at a subsequence time step of the plurality of time steps.

Claims (59)

1 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to implement a sequence transduction neural network for transducing an input sequence comprising a respective input element at each of a plurality of input positions into an output sequence comprising a respective output element at each of a plurality of output positions, the sequence transduction neural network comprising;

one or more subroutine neural networks that are each configured to, at each of a plurality of time steps, receive a subroutine network input and to generate a subroutine network output, wherein the subroutine network input comprises a subroutine input sequence comprising a respective subroutine input element at each of a plurality of subroutine input positions,

wherein each subroutine neural network comprises:

an encoder neural network comprising a masked self-attention layer, wherein the encoder neural network is configured to, at each of the plurality of time steps, receive an encoder network input comprising i) the subroutine input sequence of the subroutine neural network and ii) an input mask that masks one or more subroutine input elements of the subroutine input sequence and to process the subroutine input sequence of the subroutine neural network while masking out the one or more subroutine input elements of the subroutine input sequence to generate an encoded representation of the encoder network input;

a decoder neural network configured, at each of the plurality of time steps, to receive the encoded representation and to generate the subroutine network output; and

a masking neural network configured to perform conditional masking, the masking neural network being configured to, at each of the plurality of time steps:

receive a masking network input comprising i) the input mask of the encoder neural network for the time step that masks the one or more subroutine input elements of the subroutine input sequence and ii) an output from a portion of the decoder neural network and to process the input mask of the encoder neural network for the time step and the output from the portion of the decoder neural network to generate a masking network output comprising an output mask that masks a different one or more subroutine input elements of the subroutine input sequence from the input mask; and

provide the output mask for the time step as an input mask of the encoder neural network at a subsequent time step of the plurality of time steps.

2 . The system of claim 1 , wherein:

the encoder neural network comprises a sequence of one or more encoder subnetworks, each encoder subnetwork configured, at each of the plurality of time steps, to receive a respective encoder subnetwork input element for each of the plurality of input positions and to generate a respective encoder subnetwork output element for each of the plurality of input positions; and

each encoder subnetwork comprises a masked self-attention neural network layer that is configured to receive the encoder subnetwork input element for each of the plurality of input positions and, for each particular input position, apply a self-attention mechanism over the encoder subnetwork input elements at the plurality of input positions to generate a respective attention output for the particular input position according to the input mask.

3 . The system of claim 2 , wherein the encoder neural network comprises an embedding layer configured to:

for each subroutine input element in the subroutine input sequence, map the subroutine input element to an embedded representation of the subroutine input element; and

provide the embedded representations of the subroutine input elements as the encoder subnetwork input elements for a first encoder subnetwork in the sequence of encoder subnetworks.

4 . The system of claim 3 , wherein the embedded representation for each subroutine input element is a bitwise embedding.

5 . The system of claim 1 , wherein:

the decoder neural network comprises a sequence of one or more decoder subnetworks, each decoder subnetwork configured, at each of the plurality of time steps, to receive a respective decoder subnetwork input element for each of the plurality of input positions and to generate a respective decoder subnetwork output element for each of the plurality of input positions; and

each decoder subnetwork comprises an encoder-decoder attention neural network layer that is configured to receive an input for each input position and, for each particular input position, apply a self-attention mechanism over the encoded representation at the input positions using one or more queries derived from the input for the particular input position to generate an updated representation for the particular input position.

6 . The system of claim 5 , wherein the decoder neural network further comprises a stack of one or more output neural network layers that each comprises a learned linear transformation and an activation function, where the stack of output neural network layers is configured to receive the decoder subnetwork output elements of a final decoder subnetwork of the sequence of decoder subnetworks and to generate a value of the subroutine network output.

7 . The system of claim 6 , wherein the subroutine network output comprises a value, a pointer, or both.

8 . The system of claim 7 , wherein the pointer is generated according to the updated representations of the encoder-decoder attention neural network layer of a final decoder subnetwork of the sequence of decoder subnetworks.

9 . The system of claim 5 , wherein the masking network input further comprises the updated representations of the encoder-decoder attention neural network layer of a final decoder subnetwork of the sequence of decoder subnetworks.

10 . The system of claim 9 , wherein the input mask and the updated representations of the encoder-decoder attention neural network layer of the final decoder subnetwork are concatenated.

11 . The system of claim 5 , wherein the decoder subnetwork input elements for a first decoder subnetwork of the sequence of decoder subnetworks is a zero-vector.

12 . The system of claim 1 , wherein the masking neural network comprises one or more of:

one or more one-dimensional convolutional neural network layers,

one or more feed-forward neural network layers, or

an activation function that is configured to generate the output mask.

13 . The system of claim 1 , wherein the output mask of a first subroutine neural network of the one or more subroutine neural networks in a current time step is the input mask of a second subroutine neural network of the one or more subroutine neural networks at a subsequent time step in the plurality of time steps.

14 . The system of claim 1 , wherein the one or more subroutine neural networks comprise one or more of:

a comparison subroutine neural network,

an arithmetic subroutine neural network, or

a pointer manipulation subroutine neural network.

15 . A method comprising:

receiving an input sequence comprising a respective input element at each of a plurality of input positions; and

processing the input sequence using sequence transduction neural network to generate an output sequence comprising a respective output element at each of a plurality of output positions, the sequence transduction neural network comprising:

one or more subroutine neural networks that are each configured to, at each of a plurality of time steps, receive a subroutine network input and to generate a subroutine network output, wherein the subroutine network input comprises a subroutine input sequence comprising a respective subroutine input element at each of a plurality of subroutine input positions,

wherein each subroutine neural network comprises:

an encoder neural network comprising a masked self-attention layer, wherein the encoder neural network is configured to, at each of the plurality of time steps, receive an encoder network input comprising i) the subroutine input sequence of the subroutine neural network and ii) an input mask that masks one or more subroutine input elements of the subroutine input sequence and to process the subroutine input sequence of the subroutine neural network while masking out the one or more subroutine input elements of the subroutine input sequence to generate an encoded representation of the encoder network input;

a decoder neural network configured, at each of the plurality of time steps, to receive the encoded representation and to generate the subroutine network output; and

a masking neural network configured to perform conditional masking, the masking neural network being configured to, at each of the plurality of time steps:

receive a masking network input comprising i) the input mask of the encoder neural network for the time step that masks the one or more subroutine input elements of the subroutine input sequence and ii) an output from the portion of the decoder neural network and to process the input mask of the encoder neural network for the time step and the output from the portion of the decoder neural network to generate a masking network output comprising an output mask that masks a different one or more subroutine input elements of the subroutine input sequence from the input mask; and

provide the output mask for the time step as an input mask of the encoder neural network at a subsequent time step of the plurality of time steps.

16 . The method of claim 15 , wherein:

the encoder neural network comprises a sequence of one or more encoder subnetworks, each encoder subnetwork configured, at each of the plurality of time steps, to receive a respective encoder subnetwork input element for each of the plurality of input positions and to generate a respective encoder subnetwork output element for each of the plurality of input positions; and

each encoder subnetwork comprises a masked self-attention neural network layer that is configured to receive the encoder subnetwork input element for each of the plurality of input positions and, for each particular input position, apply a self-attention mechanism over the encoder subnetwork input elements at the plurality of input positions to generate a respective attention output for the particular input position according to the input mask.

17 . The method of claim 15 , wherein:

the decoder neural network comprises a sequence of one or more decoder subnetworks, each decoder subnetwork configured, at each of the plurality of time steps, to receive a respective decoder subnetwork input element for each of the plurality of input positions and to generate a respective decoder subnetwork output element for each of the plurality of input positions; and

each decoder subnetwork comprises an encoder-decoder attention neural network layer that is configured to receive an input for each input position and, for each particular input position, apply a self-attention mechanism over the encoded representation at the input positions using one or more queries derived from the input for the particular input position to generate an updated representation for the particular input position.

18 . One or more non-transitory computer readable media storing instructions that when executed by one or more computers cause the one or more computers to implement a sequence transduction neural network for transducing an input sequence comprising a respective input element at each of a plurality of input positions into an output sequence comprising a respective output element at each of a plurality of output positions, the sequence transduction neural network comprising:

one or more subroutine neural networks that are each configured to, at each of a plurality of time steps, receive a subroutine network input and to generate a subroutine network output, wherein the subroutine network input comprises a subroutine input sequence comprising a respective subroutine input element at each of a plurality of subroutine input positions,

wherein each subroutine neural network comprises:

an encoder neural network comprising a masked self-attention layer, wherein the encoder neural network is configured to, at each of the plurality of time steps, receive an encoder network input comprising i) the subroutine input sequence of the subroutine neural network and ii) an input mask that masks one or more subroutine input elements of the subroutine input sequence and to process the subroutine input sequence of the subroutine neural network while masking out the one or more subroutine input elements of the subroutine input sequence to generate an encoded representation of the encoder network input;

a decoder neural network configured, at each of the plurality of time steps, to receive the encoded representation and to generate the subroutine network output; and

a masking neural network configured to perform conditional masking, the masking neural network being configured to, at each of the plurality of time steps:

receive a masking network input comprising i) the input mask of the encoder neural network for the time step that masks the one or more subroutine input elements of the subroutine input sequence and ii) an output from a portion of the decoder neural network and to process the input mask of the encoder neural network for the time step and the output from the portion of the decoder neural network to generate a masking network output comprising an output mask that masks a different one or more subroutine input elements of the subroutine input sequence from the input mask; and

provide the output mask for the time step as an input mask of the encoder neural network at a subsequent time step of the plurality of time steps.

19 . The system of claim 1 , wherein the output from the portion of the decoder neural network comprises an output generated by an attention layer of the decoder neural network.

20 . The system of claim 1 , wherein the masking neural network is configured to generate the output mask provided as an input mask to the encoder neural network based on a result generated based on an attention layer of the decoder neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2020
From: YAN, YUJUN; SWERSKY, KEVIN JORDAN; HASHEMI, MILAD OLIA
To: GOOGLE LLC
Reel/Frame 054645/0982 →
Continuity (1)
Provisional Application 63078305 · Sep 14, 2020
References Cited (44)
US 10635974B2 · Reed · 2020 [cited by examiner]
US 20200293828A1 · Wang · 2020 [cited by examiner]
US 20240362894A1 · Yu · 2024 [cited by examiner]
Zhou, L., Zhou, Y., Corso, J. J., Socher, R., & Xiong, C. (2018). End-to-end dense video captioning with masked transformer. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 8739-874… [cited by examiner]
Ghazvininejad et al. “Mask-predict: Parallel decoding of conditional masked language models.” arXiv preprint arXiv: 1904.09324 (Year: 2019). [cited by examiner]
Yan Y et al., Neural Execution Engines: Learning to Execute Subroutines. arXiv preprint arXiv:2006.08084. (Jun. 15, 2020). [cited by examiner]
Tripathi, A., Lu, H., & Sak, H. (May 2020). End-to-end multi-talker overlapping speech recognition. In ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6129-6133). … [cited by examiner]
Zhou, L., Zhou, Y., Corso, J. J., Sacher, R., & Xiong, C. (2018). End-to-end dense video captioning with masked transformer. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 8739-87 … [cited by examiner]
Germain, Mathieu, et al. “Made: Masked autoencoder for distribution estimation.” International conference on machine learning. PMLR, (Year: 2015). [cited by examiner]
Liu L, Huang Y. Masked pre-trained encoder base on joint ctc-transformer. arXiv preprint arXiv:2005.11978. May 25, 2020. [cited by examiner]
Bahdanau et al., “Neural machine translation by jointly learning to align and translate,” arXiv:1409.0473, Sep. 2014, 15 pages. [cited by applicant]
Barabási et al., “Emergence of scaling in random networks,” Science, Oct. 1999, 286(5439):509-12. [cited by applicant]
Bunel et al., “Leveraging grammar and reinforcement learning for neural program synthesis,” arXiv preprint arXiv:1805.04276, May 2018, 15 pages. [cited by applicant]
Cai et al., “Making neural programming architectures generalize via recursion,” arXiv preprint arXiv:1704.06611, Apr. 2017, 20 pages. [cited by applicant]
Dehghani et al., “Universal transformers,” arXiv preprint arXiv:1807.03819, Jul. 2018, 23 pages. [cited by applicant]
Devlin et al., “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, Oct. 2018, 16 pages. [cited by applicant]
Devlin et al., “Robustfill: Neural program learning under noisy i/o,” arXiv preprint arXiv:1703.07469, Mar. 2017, 18 pages. [cited by applicant]
Dong et al., “Neural logic machines,” arXiv preprint arXiv:1904.11694, Apr. 2019, 22 pages. [cited by applicant]
Erdős et al., “On the evolution of random graphs,” Publ. Math. Inst. Hung. Acad. Sci., Jan. 1960, 5(1):17-60. [cited by applicant]
Esmaeilzadeh et al., “Neural acceleration for general-purpose approximate programs,” 45th Annual IEEE/ACM International Symposium on Microarchitecture, Dec. 2012, 449-460. [cited by applicant]
Freivalds et al., “Neural Shuffle-Exchange Networks—Sequence Processing in O (n log n) Time,” Advances in Neural Information Processing Systems, 2019, 6630-6641. [cited by applicant]
Graves et al., “Hybrid computing using a neural network with dynamic external memory,” Nature, Oct. 2016, 538(7626):471-6. [cited by applicant]
Graves et al., “Neural turing machines,” arXiv preprint arXiv:1410.5401, Oct. 2014, 26 pages. [cited by applicant]
Joulin et al., “Inferring algorithmic patterns with stack-augmented recurrent nets,” Advances in Neural Information Processing Systems, 2015, 190-198. [cited by applicant]
Kaiser et al., “Neural gpus learn algorithms,” arXiv preprint arXiv:1511.08228, Nov. 2015, 9 pages. [cited by applicant]
Krizhevsky et al., “Imagenet classification with deep convolutional neural networks.” Communications of the ACM, May 2017, 60(6):84-90. [cited by applicant]
Kurach et al., “Neural random-access machines,” arXiv preprint arXiv:1511.06392, Nov. 2015, 17 pages. [cited by applicant]
Mikolov et al., “Distributed representations of words and phrases and their compositionality,” Advances in neural information processing systems, 2013, 26:3111-9. [cited by applicant]
Neelakantan et al., “Neural programmer: Inducing latent programs with gradient descent,” arXiv preprint arXiv:1511.04834, Nov. 2015, 18 pages. [cited by applicant]
Newman et al., “Renormalization group analysis of the small-world network model,” Physics Letters A., Dec. 1999, 263(4-6):341-6. [cited by applicant]
Paccanaro et al., “Learning distributed representations of concepts using linear relational embedding,” IEEE Transactions on Knowledge and Data Engineering, Mar. 2001, 13(2):232-44. [cited by applicant]
Parisotto et al., “Neuro-symbolic program synthesis,” arXiv preprint arXiv:1611.01855, Nov. 2016, 14 pages. [cited by applicant]
Radford et al., “Language models are unsupervised multitask learners,” OpenAI blog, Feb. 2019, 1(8):9. [cited by applicant]
Reed and De Freitas, “Neural Programmer—Interpreters,” arXiv preprint arXiv:1511.06279, Nov. 2015, 15 pages. [cited by applicant]
Shi et al., “Learning Execution through Neural Code Fusion,” arXiv preprint arXiv:1906.07181, Jun. 2019, 13 pages. [cited by applicant]
Sukhbaatar et al., “End-to-end memory networks,” Advances in neural information processing systems, 2015, 2440-2448. [cited by applicant]
Sutskever et al., “Sequence to sequence learning with neural networks,” Advances in neural information processing systems, 2014, 3104-3112. [cited by applicant]
Sutskever et al., “Using matrices to model symbolic relationship,” Advances in neural information processing systems, 2008, 21:1593-600. [cited by applicant]
Trask et al., “Neural arithmetic logic units,” Advances in Neural Information Processing Systems, 2018, 8035-8044. [cited by applicant]
Vaswani et al.. “Attention is all you need,” Advances in Neural Information Processing Systems, 2017, 30:5998-6008. [cited by applicant]
Veličković et al., “Graph attention networks,” arXiv preprint arXiv:1710.10903, Oct. 2017, 12 pages. [cited by applicant]
Vinyals et al., “Pointer networks,” Advances in neural information processing systems, 2015, 2692-2700. [cited by applicant]
Von Neumann et al., “First Draft of a Report on the EDVAC,” IEEE Annals of the History of Computing, 1993, 15(4):27-75. [cited by applicant]
Yan et al., “Neural Execution Engines: Learning to Execute Subroutines,” Advances in Neural Information Processing Systems, 2020, 33 pages. [cited by applicant]