IP Library › Granted Patent US 12,322,156
Granted Patent B2
US 12,322,156 · App. 17/918,080 · Granted Jun 3, 2025

Input image size switchable network for adaptive runtime efficient image classification

Inventors: Anbang Yao (Beijing, CN); Yikai Wang (Beijing, CN); Ming Lu (Beijing, CN); Shandong Wang (Beijing, CN); Feng Chen (Shanghai, CN)
Assignee: Intel Corporation
G06V10/764G06V10/32G06V10/454G06V10/774G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,156
App. No.
17/918,080
Granted
Jun 3, 2025
Kind
B2
Abstract

Techniques related to implementing and training image classification networks are discussed. Such techniques include applying shared convolutional layers to input images regardless of resolution and applying normalization selectively based on the input image resolution. Such techniques further include training using mixed image size parallel training and mixed image size ensemble distillation.

Claims (59)

1. A system for image classification, the system comprising:

a non-transitory memory to store a first image at a first resolution and a second image at a second resolution less than the first resolution; and

one or more processors coupled to the memory, the one or more processors to:

apply a convolutional neural network layer to the first image or first feature maps corresponding to the first image using a plurality of convolutional layer parameters to generate one or more second feature maps corresponding to the first image;

perform a first normalization on the one or more second feature maps using a plurality of first normalization parameters to generate one or more third feature maps;

generate a first label for the first image using the one or more third feature maps;

apply the convolutional neural network layer to the second image or fourth feature maps corresponding to the second image using the plurality of convolutional layer parameters to generate one or more fifth feature maps corresponding to the second image;

perform a second normalization on the one or more fifth feature maps using a plurality of second normalization parameters exclusive of the first normalization parameters to generate one or more sixth feature maps; and

generate a second label for the second image using the one or more sixth feature maps.

2. The system of claim 1 , wherein the one or more processors are to generate the first label by applying one or more additional convolutional neural network layers each followed by an additional first normalization and applying a fully connected layer, and wherein the one or more processors are to generate the second label by applying the one or more additional convolutional neural network layers each followed by additional second normalization using parameters exclusive of those used in the additional first normalizations and applying the fully connected layer.

3. The system of claim 2 , wherein the first normalization and additional first normalizations are selected in response to the first image being at the first resolution and the second normalization and additional second normalizations are selected in response to the second image being at the second resolution.

4. The system of claim 1 , wherein the convolutional layer parameters, the first normalization parameters, and the second normalization parameters include pre-trained convolutional neural network parameters trained based on:

generation of a training set including first and second training images at the first and second resolutions, respectively, and a corresponding ground truth label, the first and second training images corresponding to a same image instance; and

parameter adjustment, at a training iteration, of the pre-trained convolutional neural network parameters in parallel using a loss term including a sum of cross entropy losses based on application of the pre-trained convolutional neural network parameters to the first and second training images.

5. The system of claim 4 , wherein the parameter adjustment includes adjustment to fully connected layer parameters of the pre-trained convolutional neural network parameters, the convolutional layer parameters and fully connected layer parameters to be shared across input image sizes and the first and second normalization parameters to be unshared across the input image sizes.

6. The system of claim 1 , wherein the convolutional layer parameters, the first normalization parameters, and the second normalization parameters include pre-trained convolutional neural network parameters trained based on:

generation of a training set including first, second, and third training images at the first resolution, the second resolution, and a third resolution less than the second resolution, and a corresponding ground truth label, the first, second, and third training images corresponding to a same image instance;

generation, at a training iteration, of an ensemble prediction based on first, second, and third predictions made using the first, second, and third training images, respectively; and

comparison of the ensemble prediction to the ground truth label.

7. The system of claim 6 , wherein the ensemble prediction includes a weighted average of logits corresponding to the pre-trained convolutional neural network parameters as applied to each of the first, second, and third training images, the weighted average of the logits weighted using logit importance scores.

8. The system of claim 6 , wherein, at the training iteration, a parameters update is based on minimization of a loss function including a sum of an ensemble loss term based on a classification probability using the ensemble prediction and a distillation loss term based on a divergence of each the first, second, and third predictions from the ensemble prediction.

9. The system of claim 8 , wherein the distillation loss term includes a first divergence of the second prediction from the first prediction, a second divergence of the second prediction from the third prediction, and a third divergence of the third prediction from the first prediction.

10. The system of claim 8 , wherein the loss function includes a sum of cross entropy losses based on application of the pre-trained convolutional neural network parameters to the first, second, and third training images.

11. The system of claim 1 , wherein the convolutional neural network layer includes one of a pruned convolutional neural network layer or a quantized convolutional neural network layer.

12. A method for image classification, the method comprising:

receiving a first image at a first resolution and a second image at a second resolution less than the first resolution;

applying a convolutional neural network layer to the first image or first feature maps corresponding to the first image using a plurality of convolutional layer parameters to generate one or more second feature maps corresponding to the first image;

performing a first normalization on the one or more second feature maps using a plurality of first normalization parameters to generate one or more third feature maps;

generating a first label for the first image using the one or more third feature maps;

applying the convolutional neural network layer to the second image or fourth feature maps corresponding to the second image using the plurality of convolutional layer parameters to generate one or more fifth feature maps corresponding to the second image;

performing a second normalization on the one or more fifth feature maps using a plurality of second normalization parameters exclusive of the first normalization parameters to generate one or more sixth feature maps; and

generating a second label for the second image using the one or more sixth feature maps.

13. The method of claim 12 , wherein generating the first label includes applying one or more additional convolutional neural network layers each followed by an additional first normalization and applying a fully connected layer, and wherein the generating of the second label includes applying the one or more additional convolutional neural network layers each followed by additional second normalization using parameters exclusive of those used in the additional first normalizations and applying the fully connected layer.

14. The method of claim 12 , wherein the convolutional layer parameters, the first normalization parameters, and the second normalization parameters include pre-trained convolutional neural network parameters trained based on:

generation of a training set including first and second training images at the first and second resolutions, respectively, and a corresponding ground truth label, the first and second training images corresponding to a same image instance; and

parameter adjustment, at a training iteration, of the pre-trained convolutional neural network parameters in parallel using a loss term including a sum of cross entropy losses based on application of the pre-trained convolutional neural network parameters to the first and second training images.

15. The method of claim 12 , wherein the convolutional layer parameters, the first normalization parameters, and the second normalization parameters include pre-trained convolutional neural network parameters trained based on:

generation of a training set including first, second, and third training images at the first resolution, the second resolution, and a third resolution less than the second resolution, and a corresponding ground truth label, the first, second, and third training images corresponding to a same image instance;

generation, at a training iteration, of an ensemble prediction based on first, second, and third predictions made using the first, second, and third training images, respectively; and

comparison of the ensemble prediction to the ground truth label.

16. The method of claim 15 , wherein, at the training iteration, a parameters update is based on minimization of a loss function including a sum of an ensemble loss term based on a classification probability using the ensemble prediction and a distillation loss term based on a divergence of each the first, second, and third predictions from the ensemble prediction.

17. At least one non-transitory machine readable medium comprising a plurality of instructions that, in response to being executed on a device, cause the device to at least:

receive a plurality of sequential video images including representations of a human face;

receive a first image at a first resolution and a second image at a second resolution less than the first resolution;

apply a convolutional neural network layer to the first image or first feature maps corresponding to the first image using a plurality of convolutional layer parameters to generate one or more second feature maps corresponding to the first image;

perform a first normalization on the one or more second feature maps using a plurality of first normalization parameters to generate one or more third feature maps;

generate a first label for the first image using the one or more third feature maps;

apply the convolutional neural network layer to the second image or fourth feature maps corresponding to the second image using the plurality of convolutional layer parameters to generate one or more fifth feature maps corresponding to the second image;

perform a second normalization on the one or more fifth feature maps using a plurality of second normalization parameters exclusive of the first normalization parameters to generate one or more sixth feature maps; and

generate a second label for the second image using the one or more sixth feature maps.

18. The machine readable medium of claim 17 , wherein the instructions are to cause the device to generate the first label by applying one or more additional convolutional neural network layers each followed by an additional first normalization and applying a fully connected layer, and wherein the instructions are to cause the device to generate the second label by applying the one or more additional convolutional neural network layers each followed by additional second normalization using parameters exclusive of those used in the additional first normalizations and applying the fully connected layer.

19. The machine readable medium of claim 17 , wherein the convolutional layer parameters, the first normalization parameters, and the second normalization parameters include pre-trained convolutional neural network parameters trained based on:

generation of a training set including first and second training images at the first and second resolutions, respectively, and a corresponding ground truth label, the first and second training images corresponding to a same image instance; and

parameter adjustment, at a training iteration, of the pre-trained convolutional neural network parameters in parallel using a loss term including a sum of cross entropy losses based on application of the pre-trained convolutional neural network parameters to the first and second training images.

20. The machine readable medium of claim 17 , wherein the convolutional layer parameters, the first normalization parameters, and the second normalization parameters include pre-trained convolutional neural network parameters trained based on:

generation of a training set including first, second, and third training images at the first resolution, the second resolution, and a third resolution less than the second resolution, and a corresponding ground truth label, the first, second, and third training images corresponding to a same image instance;

generation, at a training iteration, of an ensemble prediction based on first, second, and third predictions made using the first, second, and third training images, respectively; and

comparison of the ensemble prediction to the ground truth label.

21. The machine readable medium of claim 20 , wherein, at the training iteration, a parameters update is based on minimization of a loss function including a sum of an ensemble loss term based on a classification probability using the ensemble prediction and a distillation loss term based on a divergence of each the first, second, and third predictions from the ensemble prediction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2025
From: YAO, ANBANG; WANG, YIKAI; LU, MING; WANG, SHANDONG; CHEN, FENG
To: INTEL CORPORATION
Reel/Frame 070254/0622 →
Continuity (1)
Related Publication 20230343068A1 · Oct 26, 2023
References Cited (32)
US 20150161170A1 · Yee et al. · 2015 [cited by applicant]
US 20180315399A1 · Kaul · 2018 [cited by examiner]
US 20190318502A1 · He · 2019 [cited by examiner]
US 20200258223A1 · Yip · 2020 [cited by examiner]
US 20200320748A1 · Levinshtein · 2020 [cited by examiner]
US 20200388033A1 · Matlock · 2020 [cited by examiner]
US 20230343068A1 · Yao · 2023 [cited by examiner]
CN 104537393 · 2015 [cited by applicant]
CN 105975931 · 2016 [cited by applicant]
CN 107967484 · 2018 [cited by applicant]
CN 110472681A · 2019 [cited by applicant]
CN 111179212A · 2020 [cited by applicant]
WO 2018109505A1 · 2018 [cited by applicant]
WO 2019233244A1 · 2019 [cited by applicant]
International Preliminary Report on Patentability for PCT Patent Application No. PCT/CN2020/096035, dated Dec. 29, 2022. [cited by applicant]
International Search Report and Written Opinion for PCT Application No. PCT/CN2020/096035, dated Mar. 11, 2021. [cited by applicant]
Elad, H., et al., “Mix & Match: Training Convnets with Mixed Image Sizes for Improved Accuracy, Speed and Scale Resiliency”, https://arxiv.org/abs/1908.08986; Aug. 12, 2019. [cited by applicant]
He, K., et al., “Deep residual learning for image recognition”, CVPR (2016). [cited by applicant]
Howard, Andrew G., et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications”, https://arxiv.org/abs/1704.04861, Apr. 17, 2017. [cited by applicant]
Sandler, M., et al., “Mobilenetv2: Inverted residuals and linear bottlenecks”, CVPR (2018). [cited by applicant]
Yu, J., et al., “Slimmable neural networks”, ICLR (2019). [cited by applicant]
Yu, J., et al., “Universally slimmable networks and improved training techniques”, ICCV (2019). [cited by applicant]
Zhang, D., et al., “LQ-Nets: Learned quantization for highly accurate and compact deep neural networks”, ECCV (2018). [cited by applicant]
Zhao et al., “Pyramid Scene Parsing Network,” Arxiv.org, Cornell university Library, Dec. 4, 2016, 11 pages. [cited by applicant]
Redmon et al. “YOLO9000: Better, Faster, Stronger,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Dec. 25, 2016, 9 pages. [cited by applicant]
Zhu et al., “A 3D Coarse-to-Fine Framework for Volumetric Medical Image Segmentation,” 2018 International Conference on 3D Vision (3DV), 2018, 9 pages. [cited by applicant]
Ge et al., “Low-Resolution Face Recognition in the Wild via Selective Knowledge Distillation,” IEEE Transactions on Image Processing, vol. 28, No. 4, Apr. 2019, 12 pages. [cited by applicant]
Valada et al., “Self-Supervised Model Adaptation for Multimodal Semantic Segmentation,” Arxiv.org, Cornell University Library, Jul. 8, 2019, 44 pages. [cited by applicant]
Li et al., “Learning to Learn Parameterized Classification Networks for Scalable Input Images,” Arxiv.org, Cornell University Library, Jul. 13, 2020, 26 pages. [cited by applicant]
Wang et al., “Resolution Switchable Networks for Runtime Efficient Image Recognition,” Computer Vision and Pattern Recognition, Jul. 29, 2020, 20 pages. [cited by applicant]
Japan Patent Office, “Decision to Grant a Patent,” issued in connection with Japanese Patent Application No. 2022-564542, dated Feb. 20, 2024, 5 pages. [English Translation Included]. [cited by applicant]
European Patent Office, “Extended European Search Report,” issued in connection with European Patent Application No. 20941170.1, dated Feb. 26, 2024, 10 pages. [cited by applicant]