IP Library Granted Patent US 12,657,895
Granted Patent B2
US 12,657,895 · App. 17/829,227 · Granted Jun 16, 2026

Training large-scale vision transformer neural networks

Inventors: Lucas Klaus Beyer (Zurich, CH); Neil Matthew Tinmouth Houlsby (Zurich, CH); Alexander Kolesnikov (Zurich, CH); Xiaohua Zhai (Zurich, CH)
Assignee: Google LLC
G06V10/82G06V10/7747
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,895
App. No.
17/829,227
Granted
Jun 16, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training Vision Transformer (ViT) neural networks.

Claims (68)

1 . A method of training a vision Transformer neural network, the vision Transformer neural network configured to:

obtain a plurality of image patches of an image, wherein each image patch comprises a different subset of the pixels of the image;

process the plurality of image patches to generate an input sequence comprising a respective input element at each of a plurality of positions, wherein the input sequence includes a respective input element corresponding to each of the plurality of image patches;

process the input sequence through a plurality of self-attention neural network blocks to generate an output sequence comprising a respective output element at each of the positions; and

process one or more of the output elements using one or more output layers to generate a classification output for the image, and the method comprising:

obtaining first training data, the first training data comprising a plurality of training images and a respective target classification output for each training image; and

training the vision Transformer neural network on the first training data, the training comprising:

during the training,

updating parameters of the one or more output layers using a first weight decay value, and

updating parameters of the plurality of self-attention neural network blocks using a second weight decay value,

wherein the first weight decay value is higher than the second weight decay value, and wherein updating the parameters of the one or more output layers and the parameters of the one or more self-attention blocks comprises repeatedly performing the following:

computing, using one or more of the training examples, gradients of an objective function; and

applying an optimizer to the gradients to generate a gradient-based update to the parameters, wherein the optimizer makes use of a respective momentum value for each of the parameters, and wherein the respective momentum values are stored with a reduced precision relative to the parameters.

2 . The method of claim 1 , wherein the first weight decay value is greater than or equal to .3 while the second weight decay value is less than .3.

3 . The method of claim 2 , wherein the first weight decay value is greater than or equal to 3.0.

4 . The method of claim 3 , wherein the second weight decay value is less than or equal to . 1.

5 . The method of claim 1 , wherein each input element in the input sequence corresponds to a respective one of the image patches, and wherein the one or more output layers comprise:

an aggregation layer block that is configured to aggregate all of the output elements to generate an aggregated output element; and

one or more final output layers that are configured to generate the classification output from the aggregated output element.

6 . The method of claim 5 , wherein the one or more final output layers are a single linear layer that is configured to map the aggregated output element to the classification output.

7 . The method of claim 5 , wherein the aggregation block is configured to apply multihead attention pooling to the output elements to generate the aggregated output element.

8 . The method of claim 5 , wherein the aggregation block is configured to apply global average pooling to the output elements to generate the aggregated output element.

9 . The method of claim 1 , wherein the training further comprises:

during an initial phase of the training, linearly annealing a learning-rate for the parameters of the output layers and the self-attention blocks away from zero;

during a final phase of the training, linearly annealing the learning-rate toward zero; and

during a main phase of the training that is between the initial phase and the final phase, applying a schedule that prevents the learning-rate from reaching zero.

10 . The method of claim 1 , wherein the respective momentum values are stored with half-precision relative to the parameters.

11 . The method of claim 1 , further comprising:

after training the vision Transformer on the first training data, training the plurality of self-attention neural network blocks jointly with a different set of one or more output layers on second training data to perform a different, downstream task.

12 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one more computers to perform operations for training a vision Transformer neural network, the vision Transformer neural network configured to:

obtain a plurality of image patches of an image, wherein each image patch comprises a different subset of the pixels of the image;

process the plurality of image patches to generate an input sequence comprising a respective input element at each of a plurality of positions, wherein the input sequence includes a respective input element corresponding to each of the plurality of image patches;

process the input sequence through a plurality of self-attention neural network blocks to generate an output sequence comprising a respective output element at each of the positions; and

process one or more of the output elements using one or more output layers to generate a classification output for the image, and the method comprising:

obtaining first training data, the first training data comprising a plurality of training images and a respective target classification output for each training image; and

training the vision Transformer neural network on the first training data, the training comprising:

during the training,

updating parameters of the one or more output layers using a first weight decay value, and

updating parameters of the plurality of self-attention neural network blocks using a second weight decay value,

wherein the first weight decay value is higher than the second weight decay value, and wherein updating the parameters of the one or more output layers and the parameters of the one or more self-attention blocks comprises repeatedly performing the following:

computing, using one or more of the training examples, gradients of an objective function; and

applying an optimizer to the gradients to generate a gradient-based update to the parameters, wherein the optimizer makes use of a respective momentum value for each of the parameters, and wherein the respective momentum values are stored with a reduced precision relative to the parameters.

13 . A system comprising one or more computers and one or more storage devices storing instructions that when executed cause the one more computers to perform operations for training a vision Transformer neural network, the vision Transformer neural network configured to:

obtain a plurality of image patches of an image, wherein each image patch comprises a different subset of the pixels of the image;

process the plurality of image patches to generate an input sequence comprising a respective input element at each of a plurality of positions, wherein the input sequence includes a respective input element corresponding to each of the plurality of image patches;

process the input sequence through a plurality of self-attention neural network blocks to generate an output sequence comprising a respective output element at each of the positions; and

process one or more of the output elements using one or more output layers to generate a classification output for the image, and the method comprising:

obtaining first training data, the first training data comprising a plurality of training images and a respective target classification output for each training image; and

training the vision Transformer neural network on the first training data, the training comprising:

during the training,

updating parameters of the one or more output layers using a first weight decay value, and

updating parameters of the plurality of self-attention neural network blocks using a second weight decay value,

wherein the first weight decay value is higher than the second weight decay value, and where in updating the parameters of the one or more output layers and the parameters of the one or more self-attention blocks comprises repeatedly performing the following:

computing, using one or more of the training examples, gradients of an objective function; and

applying an optimizer to the gradients to generate a gradient-based update to the parameters, wherein the optimizer makes use of a respective momentum value for each of the parameters, and where in the respective momentum values are stored with a reduced precision relative to the parameters.

14 . The system of claim 13 , wherein the first weight decay value is greater than or equal to .3 while the second weight decay value is less than .3.

15 . The system of claim 14 , wherein the first weight decay value is greater than or equal to 3.0.

16 . The system of claim 15 , wherein the second weight decay value is less than or equal to .1.

17 . The system of claim 13 , wherein each input element in the input sequence corresponds to a respective one of the image patches, and wherein the one or more output layers comprise:

an aggregation layer block that is configured to aggregate all of the output elements to generate an aggregated output element; and

one or more final output layers that are configured to generate the classification output from the aggregated output element.

18 . The system of claim 13 , the operations further comprising:

after training the vision Transformer on the first training data, training the plurality of self-attention neural network blocks jointly with a different set of one or more output layers on second training data to perform a different, downstream task.

19 . The system of claim 13 , wherein the respective momentum values are stored with half-precision relative to the parameters.

20 . The system of claim 13 , wherein the training further comprises:

during an initial phase of the training, linearly annealing a learning-rate for the parameters of the output layers and the self-attention blocks away from zero;

during a final phase of the training, linearly annealing the learning-rate toward zero; and

during a main phase of the training that is between the initial phase and the final phase, applying a schedule that prevents the learning-rate from reaching zero.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2022
From: BEYER, LUCAS KLAUS; HOULSBY, NEIL MATTHEW TINMOUTH; KOLESNIKOV, ALEXANDER; ZHAI, XIAOHUA
To: GOOGLE LLC
Reel/Frame 060468/0785 →
Continuity (2)
Provisional Application 63194900 · May 28, 2021
Related Publication 20220383630A1 · Dec 1, 2022
References Cited (102)
US 20180357540A1 · Hwang · 2018 [cited by examiner]
US 20210166131A1 · Pfau · 2021 [cited by examiner]
US 20240086716A1 · Su · 2024 [cited by examiner]
CN 111256906A · 2020 [cited by applicant]
CN 112069883A · 2020 [cited by applicant]
CN 112699991A · 2021 [cited by applicant]
Touvron et al. (“Training data-efficient image transformers &distillation through attention”, Jan. 15, 2021) (Year: 2021). [cited by examiner]
Loshchilov et al. (“Decoupled Weight Decay Regularization”, Jan. 4, 2019) (Year: 2019). [cited by examiner]
Bjorck et al. (“Understanding Decoupled and Early Weight Decay”, 2021) (Year: 2021). [cited by examiner]
Zhang et al. (“Three Mechanisms of Weight Decay Regularization”, Oct. 29, 2018) (Year: 2018). [cited by examiner]
Abnar et al., “Quantifying attention flow in transformers,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, pp. 4190-4197. [cited by applicant]
Bachman et al., “Learning representations by maximizing mutual information across views,” Proceedings of the 33rd International Conference on Neural Information Processing Systems, Dec. 2019, pp. 15535-15545. [cited by applicant]
Baevski et al., “Adaptive input representations for neural language modeling,” CoRR, arXiv:1809.10853, Sep. 28, 2018, 13 pages. [cited by applicant]
Barbu et al., “Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models,” Proceedings of the 33rd International Conference on Neural Information Processing Systems, Dec. 2019,… [cited by applicant]
Bello et al., “Attention augmented convolutional networks,” Proceedings of the IEEE/CVF international conference on computer vision, Oct. 27-Nov. 2, 2019, pp. 3286-3295. [cited by applicant]
Bello et al., “Revisiting ResNets: Improved training and scaling strategies,” Advances in Neural Information Processing Systems, Dec. 6, 2021, pp. 22614-22627. [cited by applicant]
Beyer et al., “Are we done with imagenet?” CoRR, arXiv:2006.07159, Jun. 12, 2020, 15 pages. [cited by applicant]
Brown et al., “Language models are few-shot learners,” CoRR, arXiv:2005.14165v1, May 28, 2020, 72 pages. [cited by applicant]
Brown et al., “Language models are few-shot learners,” Proceedings of the 34th International Conference on Neural Information Processing Systems, Dec. 2020, pp. 1877-1901. [cited by applicant]
Carion et al., “End-to-end object detection with transformers,” European conference on computer vision, Aug. 23, 2020, pp. 213-229. [cited by applicant]
Chen et al., “A simple framework for contrastive learning of visual representations,” International conference on machine learning, Nov. 21, 2020, pp. 1597-1607. [cited by applicant]
Chen et al., “Big self-supervised models are strong semi-supervised learners,” Proceedings of the 34th International Conference on Neural Information Processing Systems, Dec. 2020, pp. 22243-22255. [cited by applicant]
Chen et al., “Generative pretraining from pixels,” International conference on machine learning, Nov. 21, 2020, pp. 1691-1703. [cited by applicant]
Chen et al., “UNITER: Universal image-text representation learning,” European conference on computer vision, Aug. 23, 2020, pp. 104-120. [cited by applicant]
Child et al., “Generating long sequences with sparse transformers,” CoRR, arXiv:1904.10509, Apr. 23, 2019, 10 pages. [cited by applicant]
Cordonnier et al., “On the relationship between self-attention and convolutional layers,” Eighth International Conference on Learning Representations, Apr. 26-30, 2020, 18 pages. [cited by applicant]
Cortes et al., “Support-vector networks,” Machine learning, Sep. 1995, 20:273-297. [cited by applicant]
Deng et al., “ImageNet: A large-scale hierarchical image database,” 2009 IEEE conference on computer vision and pattern recognition, Jun. 20, 2009, pp. 248-255. [cited by applicant]
Devlin et al., “Bert: Pre-training of deep bidirectional transformers for language understanding,” CoRR, arXiv:1810.04805, Oct. 11, 2018, 16 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of deep bidirectional transformers for language understanding,” Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human… [cited by applicant]
Djolonga et al., “On Robustness and Transferability of Convolutional Neural Networks,” CoRR, preprint arXiv:2007.08558, Jul. 16, 2020, 24 pages. [cited by applicant]
Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” In International Conference on Learning Representations, Jun. 2021, 22 pages. [cited by applicant]
Extended European Search Report in European Appln. No. 22176029.1, dated Oct. 31, 2022, 9 pages. [cited by applicant]
Ghosh et al., “Context-aware attentive knowledge tracing,” Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, Jul. 24, 2020, pp. 2330-2339. [cited by applicant]
Grill et al., “Bootstrap your own latent-a new approach to self-supervised learning,” Proceedings of the 34th International Conference on Neural Information Processing Systems Dec. 2020, pp. 21271-21284. [cited by applicant]
He et al., “Deep residual learning for image recognition,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 2016, pp. 770-778. [cited by applicant]
He et al., “Momentum contrast for unsupervised visual representation learning,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13-19, 2020, pp. 9729-9738. [cited by applicant]
Henaff et al., “Data-efficient image recognition with contrastive predictive coding,” International conference on machine learning, Nov. 21, 2020, pp. 4182-4192. [cited by applicant]
Henighan et al., “Scaling laws for autoregressive generative modeling,” CoRR, arXiv:2010.14701, Oct. 28, 2020, 37 pages. [cited by applicant]
Ho et al., “Axial attention in multidimensional transformers,” CoRR, arXiv:1912.12180, Dec. 20, 2019, 11 pages. [cited by applicant]
Hu et al., “Relation networks for object detection,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 2018, pp. 3588-3597. [cited by applicant]
Huang et al., CCNet: Criss-cross attention for semantic segmentation. Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, pp. 603-612. [cited by applicant]
Huang et al., “GPipe: Efficient training of giant neural networks using pipeline parallelism,” Proceedings of the 33rd International Conference on Neural Information Processing Systems, Dec. 2019, pp. 103-112. [cited by applicant]
Ioffe et al., “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” International conference on machine learning, Jun. 1, 2015, pp. 448-456. [cited by applicant]
Jia et al., “Scaling up visual and vision-language representation learning with noisy text supervision,” International conference on machine learning, Jul. 1, 2021, pp. 4904-4916. [cited by applicant]
Kaplan et al., “Scaling laws for neural language models,” CoRR, arXiv:2001.08361, Jan. 23, 2020, 30 pages. [cited by applicant]
Kolesnikov et al., “Big transfer (BiT): General visual representation learning,” Computer Vision—ECCV 2020: 16th European Conference, Aug. 2020, pp. 491-507. [cited by applicant]
Krizhevsky et al., “Imagenet classification with deep convolutional neural networks,” Proceedings of the 25th International Conference on Neural Information Processing Systems, Dec. 2012, pp. 1097-1105. [cited by applicant]
Krizhevsky, “Learning multiple layers of features from tiny images,” Technical report, Apr. 8, 2009, 60 pages. [cited by applicant]
LeCun et al., “Backpropagation applied to handwritten zip code recognition,” Neural computation, Dec. 1989, 1(4):541-551. [cited by applicant]
Lee et al., “Set transformer: A framework for attention-based permutation-invariant neural networks,” Proceedings of the 36th International Conference on Machine Learning, May 24, 2019, pp. 3744-3753. [cited by applicant]
Li et al., “VisualBERT: A simple and performant baseline for vision and language,” CoRR, arXiv:1908.03557, Aug. 9, 2019, 14 pages. [cited by applicant]
Locatello et al., “Object-centric learning with slot attention,” CoRR, arXiv:2006.15055, Oct. 14, 2020, 27 pages. [cited by applicant]
Lu et al., “VILBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” Proceedings of the 33rd International Conference on Neural Information Processing Systems, Dec. 2019, pp. 13… [cited by applicant]
Mahajan et al., “Exploring the limits of weakly supervised pretraining,” Computer Vision—ECCV 2018, Lecture Notes in Computer Science, Springer, Oct. 9, 2018, 11206:185-201. [cited by applicant]
Mahajan et al., “Exploring the limits of weakly supervised pretraining,” Proceedings of the European conference on computer vision, Sep. 2018, pp. 181-196. [cited by applicant]
Nilsback et al., “Automated flower classification over a large No. of classes,” 2008 Sixth Indian conference on computer vision, graphics & image processing, Dec. 16, 2008, pp. 722-729. [cited by applicant]
Parkhi et al., “Cats and dogs,” 2012 IEEE conference on computer vision and pattern recognition, Jun. 16, 2012, pp. 3498-3505. [cited by applicant]
Parmar et al., “Image transformer,” International conference on machine learning, Jul. 3, 2018, pp. 4055-4064. [cited by applicant]
Pham et al., “Meta pseudo labels,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Jun. 20-25, 2021, pp. 11557-11568. [cited by applicant]
Polyak et al., “Acceleration of stochastic approximation by averaging,” SIAM journal on control and optimization, Jul. 1992, 30(4):838-855. [cited by applicant]
Radford et al., “Improving language understanding with unsupervised learning,” Jun. 11, 2018, retrieved on Dec. 14, 2023, retrieved from URL: <https://openai.com/research/language-unsupervised>, 15 pages. [cited by applicant]
Radford et al., “Language models are unsupervised multitask learners,” Technical Report, 2019, 24 pages. [cited by applicant]
Radford et al., “Learning transferable visual models from natural language supervision,” International conference on machine learning, Jul. 1, 2021, pp. 8748-8763. [cited by applicant]
Raffel et al., “Exploring the limits of transfer learning with a unified text-to-text transformer,” CoRR, arXiv:1910.10683, Oct. 2019, 67 pages. [cited by applicant]
Ramachandran et al., “Stand-alone self-attention in vision models,” Proceedings of the 33rd International Conference on Neural Information Processing Systems, Dec. 2019, pp. 68-80. [cited by applicant]
Recht et al., “Do imagenet classifiers generalize to imagenet?” International conference on machine learning, May 24, 2019, pp. 5389-5400. [cited by applicant]
Russakovsky et al., “ImageNet large scale visual recognition challenge,” International journal of computer vision, Dec. 2015, 115:211-252. [cited by applicant]
Salimans et al., “Weight Normalization: A simple reparameterization to accelerate training of deep neural networks,” Proceedings of the 30th International Conference on Neural Information Processing Systems, Dec. 2016, … [cited by applicant]
Sohoni et al., “Low-memory neural network training: A technical report,” CoRR, arXiv:1904.10631, Apr. 24, 2019, 38 pages. [cited by applicant]
Srinivas et al., “Bottleneck transformers for visual recognition,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Jun. 20, 2021, pp. 16519-16529. [cited by applicant]
Sun et al., “Revisiting unreasonable effectiveness of data in deep learning era,” Proceedings of the IEEE international conference on computer vision, Oct. 2017, pp. 843-852. [cited by applicant]
Sun et al., “VideoBERT: A joint model for video and language representation learning,” Proceedings of the IEEE/CVF international conference on computer vision, Apr. 2019, pp. 7464-7473. [cited by applicant]
Tan et al., “EfficientNet: Rethinking model scaling for convolutional neural networks,” International conference on machine learning, May 24, 2019, pp. 6105-6114. [cited by applicant]
Touvron et al., “Fixing the train-test resolution discrepancy: FixEfficientNet,” CoRR, arXiv:2003.08237, Nov. 18, 2020, 5 pages. [cited by applicant]
Touvron et al., “Training data-efficient image transformers & distillation through attention,” International conference on machine learning, Jul. 1, 2021, pp. 10347-10357. [cited by applicant]
Tschannen et al., “Self-supervised learning of video-induced visual invariances,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020, pp. 13806-13815. [cited by applicant]
Vaswani et al., “Attention is all you need,” Proceedings of the 31st International Conference on Neural Information Processing Systems, Dec. 2017, pp. 6000-6010. [cited by applicant]
Vaswani et al., “Scaling local self-attention for parameter efficient visual backbones,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 12894-12904. [cited by applicant]
Wang et al., “Axial-DeepLab: Stand-alone axial-attention for panoptic segmentation,” CoRR, arXiv:2003.07853, Aug. 6, 2020, 26 pages. [cited by applicant]
Wang et al., “Learning deep transformer models for machine translation,” 57th Annual Meeting of the Association for Computational Linguistics, Jul. 28-Aug. 2, 2019, pp. 1810-1822. [cited by applicant]
Wang et al., “Non-local neural networks,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 18-22, 2018, pp. 7794-7803. [cited by applicant]
Wang et al., “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” Proceedings of the IEEE/CVF international conference on computer vision, Oct. 2021, pp. 568-578. [cited by applicant]
Weissenborn et al., “Scaling autoregressive video models,” CoRR, arXiv:1906.02634, Jun. 6, 2019, 24 pages. [cited by applicant]
Welinder et al., “Caltech-UCSD Birds 200,” Technical Report CNS-TR-2010-001, California Institute of Technology, 2010, 15 pages. [cited by applicant]
Wu et al., “Group normalization,” Proceedings of the European Conference on Computer Vision, Sep. 8-14, 2018, pp. 3-19. [cited by applicant]
Wu et al., “Visual transformers: Token-based image representation and processing for computer vision,” CoRR, arXiv:2006.03677, Jun. 5, 2020, 12 pages. [cited by applicant]
Xie et al., “Self-training with noisy student improves imagenet classification,” CoRR, arXiv:1911.04252, Nov. 2019, 18 pages. [cited by applicant]
Xie et al., “Self-training with noisy student improves imagenet classification,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Jun. 2020, pp. 10687-10698. [cited by applicant]
Yuan et al., “Tokens-to-token vit: Training vision transformers from scratch on imagenet,” Proceedings of the IEEE/CVF international conference on computer vision, Oct. 2021, pp. 558-567. [cited by applicant]
Zhai et al., “A large-scale study of representation learning with the visual task adaptation benchmark,” CoRR, arXiv:1910.04867, Oct. 1, 2019, 33 pages. [cited by applicant]
Zhai et al., “S41: Self-supervised semi-supervised learning,” Proceedings of the IEEE/CVF international conference on computer vision, Oct. 27-Nov. 2, 2019, pp. 1476-1485. [cited by applicant]
Zhao et al., “Exploring self-attention for image recognition,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Jun. 2020, pp. 10076-10085. [cited by applicant]
Bjorck et al., “Understanding Decoupled and Early Weight Decay,” Paper, Presented at the Conference on Artificial Intelligence, Virtual Event, Feb. 2-9, 2021; Proceedings of the Thirty-Fifth AAAI Conference on Artificia… [cited by applicant]
Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” CoRR, Submitted on Oct. 22, 2020, arXiv:2010.11929v1, pp. 1-21. [cited by applicant]
Duanmu et al., “Fast Mode and Partition Decision Using Machine Learning for Intra-Frame Coding in HEVC Screen Content Coding Extension,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems, Dec. 2016, 6… [cited by applicant]
Ishii et al., “Layer-Wise Weight Decay for Deep Neural Networks,” Paper, Presented at the Pacific-Rim Symposium on Image and Video Technology, Wuhan, China, Nov. 20-24, 2017; LNIP, 2018, 10749:276-289. [cited by applicant]
Jillani et al., “Low Complexity Intra MB Encoding in AVC/H.264,” IEEE Transactions on Consumer Electronics, Feb. 2009, 55(1):277-285. [cited by applicant]
Loshchilov et al., “Decoupled Weight Decay Regularization,” CoRR, Submitted on Jan. 4, 2019, arXiv:1711.05101v3, pp. 1-11, 19 pages. [cited by applicant]
Office Action in Chinese Appln. No. 202210600264.5, mailed on Dec. 30, 2024, 25 pages (with English translation). [cited by applicant]
Pfaff et al., “Neural network based intra prediction for video coding,” Proceedings of the Applications of Digital Image Processing XLI, Sep. 2018, 10752:7 pages. [cited by applicant]
Wu, “Research on Target Recognition of SAR Image based on Deep Learning,” Chinese Master's Theses Full-Text Database, Science and Technology Series, Jun. 15, 2019, 144 pages (with English translation). [cited by applicant]