IP Library › Granted Patent US 12,373,698
Granted Patent B2
US 12,373,698 · App. 18/073,383 · Granted Jul 29, 2025

Small and fast transformer model for multi-modal or other tasks

Inventors: Qian Lou (Oviedo, FL); Yen-Chang Hsu (Fremont, CA); Burak Uzkent (Mountain View, CA); Ting Hua (Cupertino, CA); Yilin Shen (San Jose, CA); Hongxia Jin (San Jose, CA)
Assignee: Samsung Electronics Co., Ltd.
G06N3/082G06V10/772G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,698
App. No.
18/073,383
Granted
Jul 29, 2025
Kind
B2
Abstract

A method includes obtaining, using a first electronic device, a weight matrix associated with a trained transformer model. The method also includes factorizing the weight matrix into a dictionary weight matrix and an intermediate matrix. The method further includes pruning the intermediate matrix to generate a sparse intermediate matrix. The method also includes fine-tuning the sparse intermediate matrix based on a training dataset to generate a fine-tuned sparse intermediate matrix. The method further includes determining an index matrix and a coefficient matrix based on the fine-tuned sparse intermediate matrix. In addition, the method includes deploying the dictionary weight matrix, the index matrix, and the coefficient matrix to a second electronic device without deploying the weight matrix to the second electronic device. A number of parameters in the dictionary weight matrix, the index matrix, and the coefficient matrix is smaller than a number of parameters in the weight matrix.

Claims (80)

1. A method comprising:

obtaining, using at least one processing device of a first electronic device, a weight matrix associated with a trained transformer model;

factorizing, using the at least one processing device, the weight matrix into a dictionary weight matrix and an intermediate matrix;

pruning, using the at least one processing device, the intermediate matrix to generate a sparse intermediate matrix;

fine-tuning, using the at least one processing device, the sparse intermediate matrix based on a training dataset to generate a fine-tuned sparse intermediate matrix;

determining, using the at least one processing device, an index matrix and a coefficient matrix based on the fine-tuned sparse intermediate matrix; and

deploying the dictionary weight matrix, the index matrix, and the coefficient matrix to a second electronic device without deploying the weight matrix to the second electronic device;

wherein a number of parameters in the dictionary weight matrix, the index matrix, and the coefficient matrix is smaller than a number of parameters in the weight matrix.

2. The method of claim 1 , wherein the trained transformer model comprises a text transformer.

3. The method of claim 1 , wherein the trained transformer model comprises a multi-modal transformer.

4. The method of claim 1 , wherein:

the trained transformer model comprises a text transformer; and

the obtaining, factorizing, pruning, fine-tuning, determining, and deploying are repeated for a second weight matrix associated with a second trained transformer model to generate a second dictionary weight matrix, a second index matrix, and a second coefficient matrix, the second trained transformer model comprising a multi-modal transformer.

5. The method of claim 4 , further comprising:

deploying a visual encoder to the second electronic device;

wherein the second electronic device is configured to:

use the visual encoder to process a visual input in order to generate first features;

use the dictionary weight matrix, the index matrix, and the coefficient matrix to process a textual input in order to generate second features; and

use the second dictionary weight matrix, the second index matrix, and the second coefficient matrix to process the first and second features in order to generate an output based on the visual input and the textual input.

6. The method of claim 5 , wherein:

the visual input comprises one or more images;

the textual input comprises text associated with at least one object in the one or more images; and

the second dictionary weight matrix, the second index matrix, and the second coefficient matrix are configured to use the first and second features to identify the at least one object in the one or more images based on the text.

7. An apparatus comprising:

at least one processing device configured to:

obtain a weight matrix associated with a trained transformer model;

factorize the weight matrix into a dictionary weight matrix and an intermediate matrix;

prune the intermediate matrix to generate a sparse intermediate matrix;

fine-tune the sparse intermediate matrix based on a training dataset to generate a fine-tuned sparse intermediate matrix;

determine an index matrix and a coefficient matrix based on the fine-tuned sparse intermediate matrix; and

deploy the dictionary weight matrix, the index matrix, and the coefficient matrix to an electronic device without deploying the weight matrix to the electronic device;

wherein a number of parameters in the dictionary weight matrix, the index matrix, and the coefficient matrix is smaller than a number of parameters in the weight matrix.

8. The apparatus of claim 7 , wherein the trained transformer model comprises a text transformer.

9. The apparatus of claim 7 , wherein the trained transformer model comprises a multi-modal transformer.

10. The apparatus of claim 7 , wherein:

the trained transformer model comprises a text transformer; and

the at least one processing device is further configured to process a second weight matrix associated with a second trained transformer model and generate a second dictionary weight matrix, a second index matrix, and a second coefficient matrix, the second trained transformer model comprising a multi-modal transformer.

11. The apparatus of claim 10 , wherein:

the at least one processing device is further configured to deploy a visual encoder to the electronic device; and

the electronic device is configured to:

use the visual encoder to process a visual input in order to generate first features;

use the dictionary weight matrix, the index matrix, and the coefficient matrix to process a textual input in order to generate second features; and

use the second dictionary weight matrix, the second index matrix, and the second coefficient matrix to process the first and second features in order to generate an output based on the visual input and the textual input.

12. The apparatus of claim 11 , wherein:

the visual input comprises one or more images;

the textual input comprises text associated with at least one object in the one or more images; and

the second dictionary weight matrix, the second index matrix, and the second coefficient matrix are configured to use the first and second features to identify the at least one object in the one or more images based on the text.

13. A method comprising:

obtaining, by an electronic device that stores a trained machine learning model, an input, wherein the trained machine learning model comprises a dictionary weight matrix, an index matrix, and a coefficient matrix; and

determining, using the electronic device, an output corresponding to the input by performing a linear combination using the dictionary weight matrix, the index matrix, and the coefficient matrix;

wherein performing the linear combination comprises:

generating a dictionary based on a product of the input and the dictionary weight matrix; and

for each of one or more columns of the output, determining a weighted combination of columns in the dictionary, wherein the columns in the dictionary are selected based on the index matrix, and wherein weights applied to the selected columns in the dictionary are selected based on the coefficient matrix.

14. The method of claim 13 , further comprising:

using the output as a second input to a second trained machine learning model, the second trained machine learning model comprising a second dictionary weight matrix, a second index matrix, and a second coefficient matrix; and

determining a second output using the second trained machine learning model by performing a second linear combination using the second dictionary weight matrix, the second index matrix, and the second coefficient matrix.

15. The method of claim 14 , wherein:

the trained machine learning model comprises a textual encoder configured to process a textual input; and

the second trained machine learning model comprises a multi-model transformer configured to process a multi-modal input, the multi-modal input comprising the output of the trained machine learning model.

16. The method of claim 15 , further comprising:

using a visual encoder to process a visual input;

wherein the multi-modal input further comprises an output of the visual encoder.

17. The method of claim 16 , wherein:

the visual input comprises one or more images;

the textual input comprises text associated with at least one object in the one or more images; and

the multi-model transformer is configured identify the at least one object in the one or more images based on the text.

18. An apparatus comprising:

at least one memory configured to store a trained machine learning model, wherein the trained machine learning model comprises a dictionary weight matrix, an index matrix, and a coefficient matrix; and

at least one processing device configured to:

obtain an input; and

perform a linear combination using the dictionary weight matrix, the index matrix, and the coefficient matrix to determine an output corresponding to the input;

wherein, to perform the linear combination, the at least one processing device is configured to:

generate a dictionary based on a product of the input and the dictionary weight matrix; and

for each of one or more columns of the output, determine a weighted combination of columns in the dictionary, the at least one processing device configured to select the columns in the dictionary based on the index matrix, and the at least one processing device configured to select weights applied to the selected columns in the dictionary based on the coefficient matrix.

19. The apparatus of claim 18 , wherein the at least one processing device is further configured to:

use the output as a second input to a second trained machine learning model, the second trained machine learning model comprising a second dictionary weight matrix, a second index matrix, and a second coefficient matrix; and

perform a second linear combination using the second dictionary weight matrix, the second index matrix, and the second coefficient matrix to determine a second output.

20. The apparatus of claim 19 , wherein:

the trained machine learning model comprises a textual encoder configured to process a textual input; and

the second trained machine learning model comprises a multi-model transformer configured to process a multi-modal input, the multi-modal input comprising the output of the trained machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2022
From: LOU, QIAN; HSU, YEN-CHANG; UZKENT, BURAK; HUA, TING; SHEN, YILIN; JIN, HONGXIA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061947/0045 →
Continuity (2)
Provisional Application 63286470 · Dec 6, 2021
Related Publication 20230177338A1 · Jun 8, 2023
References Cited (73)
US 10482157B2 · Konoshima · 2019 [cited by applicant]
US 10635965B2 · Chen et al. · 2020 [cited by applicant]
US 11386326B2 · Lin et al. · 2022 [cited by applicant]
US 20190156248A1 · Togashi · 2019 [cited by applicant]
US 20190378015A1 · Lin et al. · 2019 [cited by applicant]
US 20210255862A1 · Volkovs et al. · 2021 [cited by applicant]
US 20220076112A1 · Wagner et al. · 2022 [cited by applicant]
US 20220198254A1 · Dalli et al. · 2022 [cited by applicant]
CN 114169345A · 2022 [cited by applicant]
CN 114267333A · 2022 [cited by applicant]
Cahyawijaya, “Greenformers: Improving Computation and Memory Efficiency in Transformer Models via Low-Rank Approximation”, Aug. 24, 2021, arXiv:2108.10808v1 [cs.LG] (81 Pages) (Year: 2021). [cited by examiner]
Lagunas et al, “Block Pruning For Faster Transformers”, Sep. 10, 2021, arXiv:2109.04838v1 [cs.LG] (11 Pages) (Year: 2021). [cited by examiner]
Ravishankar, “Magnetic Resonance Image Reconstruction From Highly Undersampled K-Space Data Using Dictionary Learning,” 2010, University of Illinois at Urbana-Champaign (85 Pages) (Year: 2010). [cited by examiner]
Merity et al., “Pointer Sentinel Mixture Models,” arXiv:1609.07843v1 [cs.CL], Sep. 2016, 13 pages. [cited by applicant]
Ott et al., “FAIRSEQ: A Fast, Extensible Toolkit for Sequence Modeling,” arXiv:1904.01038v1, Apr. 2019, 6 pages. [cited by applicant]
Prato et al., “Fully Quantized Transformer for Machine Translation,” arXiv:1910.10485v3, Mar. 2020, 13 pages. [cited by applicant]
Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” Journal of Machine Learning Research 21, Jun. 2020, 67 pages. [cited by applicant]
Reid et al., “Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers,” arXiv:2101.00234v3, Sep. 2021, 10 pages. [cited by applicant]
So et al., “The Evolved Transformer,” arXiv:1901.11117v4, May 2019, 14 pages. [cited by applicant]
Takase et al., “Lessons on Parameter Sharing across Layers in Transformers,” arXiv:2104.06022v3, Apr. 2022, 11 pages. [cited by applicant]
Vaswani et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 2017, 11 pages. [cited by applicant]
Xia et al., “Tied Transformers: Neural Machine Translation with Shared Encoder and Decoder,” The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), Jul. 2019, 8 pages. [cited by applicant]
Zhou et al., “Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting,” The Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI-21), Mar. 2021, 10 pages. [cited by applicant]
Lou et al., “DictFormer: Tiny Transformer with Shared Dictionary,” ICLR, 2022, 16 pages. [cited by applicant]
Tariyal et al., “Deep Dictionary Learning,” IEEE Access, 2016, 14 pages. [cited by applicant]
Noach et al., “Compressing Pre-trained Language Models by Matrix Decomposition,” Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International … [cited by applicant]
Mao et al., “LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression,” Proceedings of the 28th International Conference on Computational Linguistic, Dec. 2020, 10 pages. [cited by applicant]
Maalouf et al., “Deep Learning Meets Projective Clustering,” ICLR, 2021, 20 pages. [cited by applicant]
Kamath et al., “MDETR—Modulated Detection for End-to-End Multi-Modal Understanding,” arXiv:2104.12763v1, Apr. 2021, 21 pages. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision,” arXiv:2103.00020v1, Feb. 2021, 48 pages. [cited by applicant]
Argueta et al., “Accelerating Sparse Matrix Operations in Neural Networks on Graphics Processing Units,” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 2019, 10 pages. [cited by applicant]
Carion et al., “End-to-End Object Detection with Transformers,” arXiv:2005.12872v3, May 2020, 26 pages. [cited by applicant]
Chen et al., “UNITER: Universal Image-Text Representation Learning,” arXiv:1909.11740v3, Jul. 2020, 26 pages. [cited by applicant]
Choo et al., “Understanding and Optimizing GPU Cache Memory Performance for Compute Workloads,” 13th International Symposium on Parallel and Distributed Computing, 2014, 8 pages. [cited by applicant]
Dai et al., “Dynamic DETR: End-to-End Object Detection with Dynamic Attention,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021, 10 pages. [cited by applicant]
Dai et al., “UP-DETR: Unsupervised Pre-training for Object Detection with Transformers,” arXiv:2011.09094v2, Apr. 2021, 11 pages. [cited by applicant]
Deng et al., “TransVG: End-to-End Visual Grounding with Transformers,” arXiv:2104.08541v1, Apr. 2021, 10 pages. [cited by applicant]
Gan et al., “Large-Scale Adversarial Training for Vision-and-Language Representation Learning,” arXiv:2006.06195v2, Oct. 2020, 16 pages. [cited by applicant]
Golub et al., “Handbook Series Linear Albegra: Singular Value Decomposition and Least Squares Solutions,” Numer. Math., 1970, 18 pages. [cited by applicant]
Hudson et al., “GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering,” arXiv:1902.09506v2, Apr. 2019, 19 pages. [cited by applicant]
Jiao et al., “TinyBERT: Distilling BERT for Natural Language Understanding,” arXiv:1909.10351v5, Oct. 2020, 12 pages. [cited by applicant]
Kazemzadeh et al., “ReferItGame: Referring to Objects in Photographs of Natural Scenes,” Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Oct. 2014, 12 pages. [cited by applicant]
Khudia et al., “FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference,” arXiv:2101.05615v1, Jan. 2021, 5 pages. [cited by applicant]
Kim et al., “ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision,” arXiv:2102.03334v1, Feb. 2021, 11 pages. [cited by applicant]
Krishna et al., “Visual Genome, Connecting Language and Vision Using Crowdsourced Dense Image Annotations,” arXiv:1602.07332v1, Feb. 2016, 44 pages. [cited by applicant]
Li et al., “What Does BERT with Vision Look At,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, 11 pages. [cited by applicant]
Li et al., “Referring Transformer: A One-step Approach to Multi-task Visual Grounding,” arXiv:2106.03089v1, Jun. 2021, 13 pages. [cited by applicant]
Liu et al., “Progressive Neural Architecture Search,” arXiv:1712.00559v3, Jul. 2018, 20 pages. [cited by applicant]
Liu et al., “ROBERTa: A Robustly Optimized BERT Pretraining Approach,” arXiv:1907.11692v1, Jul. 2019, 13 pages. [cited by applicant]
Lu et al., “ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks,” arXiv:1908.02265v1, Aug. 2019, 11 pages. [cited by applicant]
Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 12 pages. [cited by applicant]
Plummer et al., “Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models,” arXiv:1505.04870v4, Sep. 2016, 22 pages. [cited by applicant]
Sanh et al., “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” arXiv:1910.01108v4, Mar. 2020, 5 pages. [cited by applicant]
Sun et al., “Patient Knowledge Distillation for BERT Model Compression,” arXiv:1908.09355v1, Aug. 2019, 10 pages. [cited by applicant]
Sun et al., “Rethinking Transformer-based Set Prediction for Object Detection,” arXiv:2011.10881v1, Nov. 2020, 13 pages. [cited by applicant]
Tan et al., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” arXiv:1905.11946v5, Sep. 2020, 11 pages. [cited by applicant]
Wang et al., “MINILM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers,” arXiv:2002.10957v2, Apr. 2020, 15 pages. [cited by applicant]
Wu et al., “PhraseCut: Language-based Image Segmentation in the Wild,” arXiv:2008.01187v1, Aug. 2020, 17 pages. [cited by applicant]
Wu et al., “Lite Transformer with Long-Short Range Attention,” arXiv:2004.11886v1, Apr. 2020, 13 pages. [cited by applicant]
Yang et al., “TAP: Text-Aware Pre-training for Text-VQA and Text-Caption,” arXiv:2012.04638v1, Dec. 2020, 14 pages. [cited by applicant]
Yu et al., “Modeling Context in Referring Expressions,” arXiv:1608.00272v3, Aug. 2016, 19 pages. [cited by applicant]
Zhu et al., “Deformable DETR: Deformable Transformers for End-To-End Object Detection,” arXiv:2010.04159v4, Mar. 2021, 16 pages. [cited by applicant]
Lan et al., “ALBERT: A Lite BERT for Self-supervised Learning of Language Representations,” Eighth International Conference on Learning Representations, Feb. 2020, 17 pages. [cited by applicant]
Wu et al., “Lite Transformer with Long-Short Range Attention,” Eighth International Conference on Learning Representations, Apr. 2020, 13 pages. [cited by applicant]
Baevski et al., “Adaptive Input Representations for Neural Language Modeling,” arXiv:1809.10853v3, Feb. 2019, 13 pages. [cited by applicant]
Behnke et al., “Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Nov. 2020, 11 pages. [cited by applicant]
Brown et al., “Language Models are Few-Shot Learners,” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Jul. 2020, 25 pages. [cited by applicant]
Chen et al., “A Thorough Examination of the CNN/Daily Mail Reading Comprehension Task,” arXiv:1606.02858v2, Aug. 2016, 11 pages. [cited by applicant]
Dehghani et al., “Universal Transformers,” arXiv:1807.03819v3, Mar. 2019, 23 pages. [cited by applicant]
Katharopoulos et al., “Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention,” arXiv:2006.16236v3, Aug. 2020, 17 pages. [cited by applicant]
Ma et al., “A Tensorized Transformer for Language Modeling,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), November 1029, 11 pages. [cited by applicant]
Mehta et al., “DeLight: Deep and Light-Weight Transformer,” arXiv.2008.00623v2, Feb. 2021, 19 pages. [cited by applicant]
Merity et al., “Regularizing and Optimizing LSTM Language Models,” arXiv:1708.02182v1, Aug. 2017, 10 pages. [cited by applicant]