IP Library › Granted Patent US 12,271,791
Granted Patent B2
US 12,271,791 · App. 17/308,033 · Granted Apr 8, 2025

Attention free transformer

Inventors: Shuangfei Zhai (Sunnyvale, CA); Walter A. Talbott (Cupertino, CA); Nitish Srivastava (San Francisco, CA); Chen Huang (Cupertino, CA); Hanlin Goh (Santa Clara, CA); Joshua M. Susskind (Campbell, CA)
Assignee: Apple Inc.
G06N20/00G06F17/16G06F40/58G06N5/04G06T3/4053
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,271,791
App. No.
17/308,033
Granted
Apr 8, 2025
Kind
B2
Abstract

Attention-free transformers are disclosed. Various implementations of attention-free transformers include a gating and pooling operation that allows the attention-free transformers to provide comparable or better results to those of a standard attention-based transformer, with improved efficiency and reduced computational complexity with respect to space and time.

Claims (63)

1. A method comprising:

processing an input data using a transformer, wherein the input data comprises a string of text;

generating, using the transformer responsive to processing the input data, an output comprising a translation of the input data, wherein generating the output comprises:

processing the input data to generate a key, a value, and a query, wherein the key, the value and the query are embedding matrices generated using linear transformations of the input data;

performing a gating and pooling operation of the transformer, wherein the gating and the pooling operation comprises:

applying a non-linearity to the key;

combining the key with the value using an element-wise multiplication;

reducing a spatial context by performing a global pooling operation on the combined key and value;

applying a non-linearity to the query; and

combining the query with the reduced spatial context.

2. The method of claim 1 , wherein the transformer is an attention-free transformer, and wherein generating the output comprises executing the transformer without performing an attention operation.

3. The method of claim 1 , wherein performing the gating and pooling operation comprises performing a first gating operation, performing a pooling operation using a result of the first gating operation, and performing a second gating operation using a result of the pooling operation.

4. The method of claim 3 , wherein first gating operation comprises an element-wise product using a key matrix and a value matrix.

5. The method of claim 4 , wherein the first gating operation further comprises computing a softmax function using the key matrix prior to performing the element-wise product.

6. The method of claim 4 , wherein the pooling operation comprises computing a sum over rows of the result of the element-wise product using the key matrix and the value matrix.

7. The method of claim 6 , wherein the second gating operation comprises computing an additional element-wise product using a query matrix and the result of the pooling operation.

8. The method of claim 1 , wherein the transformer is a causal attention-free transformer.

9. The method of claim 1 , wherein the transformer is a local causal attention-free transformer.

10. The method of claim 9 , wherein the local causal attention-free transformer is an AFT-local-hard transformer.

11. The method of claim 9 , wherein the local causal attention-free transformer is an AFT-local-learned transformer.

12. The method of claim 1 , further comprising training the transformer by providing input training data to the transformer, generating a training output from the transformer, and comparing the training output to output training data.

13. The method of claim 12 , wherein training the transformer comprises performing an in-place computation of the gating and pooling operation.

14. The method of claim 1 , wherein the input data comprises a string of characters or an image.

15. The method of claim 14 , wherein the output comprises a different string of characters, a different image, or a three-dimensional model.

16. The method of claim 1 , wherein applying the non-linearity comprises applying a relu operation.

17. The method of claim 1 , wherein the gating and pooling operation comprises:

applying a non-linearity to a key;

combining the key with a value using an element-wise multiplication;

reducing a spatial context with a global pooling operation;

applying a sigmoid operation to a query; and

combining each point of the query with the reduced spatial context.

18. A system comprising:

a processor; and

a memory device containing instructions, which when executed by the processor, cause the processor to:

process an input data using a transformer, wherein the input data comprises a string of text;

generate, using the transformer responsive to processing the input data, an output comprising a translation of the input data, wherein generating the output comprises:

process the input data to generate a key, a value, and a query, wherein the key, the value and the query are embedding matrices generated using linear transformations of the input data;

perform a gating and pooling operation of the transformer, wherein the gating and the pooling operation comprises:

applying a non-linearity to the key;

combining the key with the value using an element-wise multiplication;

reducing a spatial context by performing a global pooling operation on the combined key and value;

applying a non-linearity to the query; and

combining the query with the reduced spatial context.

19. The system of claim 18 , wherein the instructions, when executed by the processor, cause the processor to perform the gating and pooling operation by performing at least one gating and pooling operation of an encoder of the transformer and at least one gating and pooling operation of a decoder of the transformer.

20. The system of claim 18 , wherein the input data comprises at least one of a two-dimensional input data such as an image or an array of values, three-dimensional input data such as a three-dimensional model or point cloud, or other input data arranged in one, two, three, or more than three dimensions and the output of the transformer comprises at least one of a prediction based on the input data, or a super-resolution image corresponding to the input data.

21. The system of claim 18 , wherein applying the non-linearity comprises applying a relu operation.

22. A non-transitory machine-readable medium comprising code that, when executed by a processor, causes the processor to perform a method, comprising:

processing an input data using a transformer, wherein the input data comprises a string of text;

generating, using the transformer responsive to processing the input data, an output comprising a translation of the input data, wherein generating the output comprises:

processing the input data to generate a key, a value, and a query, wherein the key, the value and the query are embedding matrices generated using linear transformations of the input data;

performing a gating and pooling operation of the transformer, wherein the gating and the pooling operation comprises:

applying a non-linearity to the key;

combining the key with the value using an element-wise multiplication;

reducing a spatial context by performing a global pooling operation on the combined key and value;

applying a non-linearity to the query; and

combining the query with the reduced spatial context.

23. The non-transitory machine-readable medium of claim 22 , wherein:

the transformer is an attention-free transformer;

generating the output comprises executing the transformer without performing an attention operation; and

performing the gating and pooling operation comprises:

performing a first gating operation,

performing a pooling operation using a result of the first gating operation, and

performing a second gating operation using a result of the pooling operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2024
From: ZHAI, SHUANGFEI; TALBOTT, WALTER A.; SRIVASTAVA, NITISH; HUANG, CHEN; GOH, HANLIN; SUSSKIND, JOSHUA M.
To: APPLE INC.
Reel/Frame 068288/0827 →
Continuity (3)
Provisional Application 63145429 · Feb 3, 2021
Provisional Application 63086517 · Oct 1, 2020
Related Publication 20220108212A1 · Apr 7, 2022
References Cited (39)
US 20210232773A1 · Wang · 2021 [cited by examiner]
Tay et al. (Tay) “Compositional de-attention networks”, Advances in Neural Information Processing Systems 32 (NeurIPS 2019), ISBN: 9781713807933, 2019 “https://proceedings.neurips.cc/paper/2019/file/16fc18d787294ad51711… [cited by examiner]
Shi et al. “Weak-Attention Suppression for Transformer Based Speech Recognition”, Facebook AI, USA, https://arxiv.org/abs/2005.09137 in Audio and Speech Processing (eess.AS); Computation and Language (cs.CL), interspeec… [cited by examiner]
Pan et al. “X-Linear Attention Networks for Image Captioning” DOI: https://doi.org/10.48550/arXiv.2003.14080, Mar. 31, 2020, http://arxiv.org/pdf/2003.14080.pdf (Year: 2020). [cited by examiner]
Ba, ct al., “Laycr normalization,” 2016, rctricvcd from arXiv.org/pdf/1607.06450. [cited by applicant]
Bradbury, et al., “Quasi-recurrent neural networks,” 2016, retrieved from arXiv.org/pdf/1611.01576. [cited by applicant]
Chang, et al., ShapeNet: An Information-Rich 3D Model Repository, 2015 Technical Report, retrieved from arXiv.org/pdf/1512.0301. [cited by applicant]
Chen, ct al., “Gcncrativc prctraining from Pixcls,” 2020, rctricvcd from https://cdn.openai.com/papers/Generative_Pretraining_from_Pixels_V2.pdf. [cited by applicant]
Choromanski, et al., “Rethinking attention with performers,” 2020, retrieved from https://openreview.net/pdf?id=Ua6zuk0WRH. [cited by applicant]
Chung, et al., Empirical evaluation of gated recurrent neural networks on sequence modcling, 2014, rctricvcd from arXiv.org/pdf/1412.3555. [cited by applicant]
Dahl, et al., “Pixel recursive super resolution,” 2017, retrieved from https://arxiv.org/pdf/1702.00783.pdf. [cited by applicant]
Dai, et al., “Transformer-xl: Attentive language models beyond a fixed-length context,” 2019, retrieved from arXiv.org/pdf/1901.02860. [cited by applicant]
Devlin, ct al., Prc-training of dccp bidircctional transformcrs for languagc undcrstanding, 2018, retrieved from https://arxiv.org/pdf/1810.04805.pdf. [cited by applicant]
Hochreiter, et al., “Long short-term memory,” Neural computation, 1997, 9(8): 1735-7180. [cited by applicant]
Huang, et al., “CCNET: Criss-cross attention for semantic segmentation,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 603-612. [cited by applicant]
Huang, et al., Interlaced sparse self-attention for semantic segmentation, 2019, retrieved from ArXiv.org/pdf/1907.12273. [cited by applicant]
Katharopoulos, et al., “Transformers are rnns: Fast autoregressive transformers with linear attention,” 2020 Proceedings of the International Conference on Machine Learning (ICML), retrieved from https://fleuret.org/pap… [cited by applicant]
Kitaev, ct al., “Rcformcr: The cfficicnt transformcr,” 2020, rctricvcd from arXiv.org/pdf/abs/2001.04451. [cited by applicant]
Klein, et al., “OpenNMT: Opensource toolkit for neural machine translation,” Proceedings of ACL 2017, System Demonstrations, pp. 67-72, retrieved from https://www.aclweb.org/anthology/P17-4012. [cited by applicant]
Lee, et al., “Set transformer: A framework for attention-based permutation-invariant neural nctworks,” 2019, Intcrnational Confcrcnce on Machinc Lcarning, pp. 3744-3753. [cited by applicant]
Liu, et al., “Deep learning face attributes in the wild,” 2015, Proceedings of International Conference on Computer Vision (ICCV). [cited by applicant]
Mahoney, “Large text compression benchmark,” 2011, retrieved from http://www.mattmahoney.net/dc/text.html, 95 pages. [cited by applicant]
Nash, ct al., “Polygcn: An autorcgressivc gcncrativc modcl of 3d mcshcs,” ICML 2020. [cited by applicant]
Parmar, et al., “Image transformer,” 2018, retrieved from arXiv.org/pdf/1802.05751. [cited by applicant]
Radford, et al., “Improving language understanding by generative pre-training,” 2018, retrieved from https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pd… [cited by applicant]
Rae, ct al., “Comprcssivc transformcrs for long-rangc scqucncc modclling,” 2020, rctricvcd from arXiv.org/pdf/1911.05507. [cited by applicant]
Ramachandran, et al., “Stand-alone self-attention in vision models,” 2019, retrieved from ArXiv.org/pdf/1906.05909. [cited by applicant]
Rewon, et al., “Generating long sequences with sparse Transformers,” 2019, retrieved from http://arxiv.org/abs/1904.10509. [cited by applicant]
Roy, et al., “Efficient content-based sparse attention with routing transformers,” 2020, retrieved from arXiv.org/pdf/2003.05997. [cited by applicant]
Salimans, et al., “Pixelcnn++: Improving the pixelenn with discretized logistic mixture likclihood and othcr modifications,” 2017, rctricvcd from arXiv.org/pdf/1701.05517. [cited by applicant]
Sukhbaatar, et al., “Adaptive attention span in transformers,” 2019 Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. [cited by applicant]
Tay et al., “Efficient transformers: A survey,” 2020, retrieved from https://arxiv.org/pdf/2009.06732.pdf. [cited by applicant]
Tay ct al., “Sparsc sinkhorn attcntion,” 2020, rctricvcd from https://arxiv.org/pdf/2002.11296.pdf. [cited by applicant]
Tay, et al., “Synthesizer: Rethinking self-attention in transformer models,” 2020, retrieved from https://arxiv.org/pdf/2005.00743.pdf. [cited by applicant]
Vaswani, et al., “Attention is all you need,” 2017 Advances in neural information processing systems, pp. 5998-6008, 2017. [cited by applicant]
Wang, et al., “Axial-deeplab: Standalone axial-attention for panoptic segmentation,” 2020, retrieved from arxiv.org/pdf/2003.07853. [cited by applicant]
Wang, et al., “Linformer: Self-attention with linear complexity,” 2020, retrieved from ArXiv.org/pdf/2006.04768. [cited by applicant]
Wu, et al., “Pay less attention with lightweight and dynamic convolutions,” 2019, retrieved from ArXiv.org/pdf/1901.10430. [cited by applicant]
Zhu, et al., “Asymmetric non-local neural networks for semantic segmentation,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 593-602. [cited by applicant]