IP Library › Granted Patent US 12,394,190
Granted Patent B2
US 12,394,190 · App. 17/701,209 · Granted Aug 19, 2025

Method and apparatus for classifying images using an artificial intelligence model

Inventors: Burak Uzkent (Mountain View, CA); Vasili Ramanishka (Mountain View, CA); Yilin Shen (Santa Clara, CA); Hongxia Jin (San Jose, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06V10/82G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,190
App. No.
17/701,209
Granted
Aug 19, 2025
Kind
B2
Abstract

An apparatus for performing image processing, may include at least one processor configured to: input an image to a vision transformer comprising a plurality of encoders that correspond to at least one fixed encoder and a plurality of adaptive encoders; process the image via the at least one fixed encoder to obtain image representations; determine one or more layers of the plurality of adaptive encoders to drop, by inputting the image representations to a policy network configured to determine layer dropout actions for the plurality of adaptive encoders; and obtain a class of the input image using remaining layers of the plurality of adaptive encoders other than the dropped one or more layers.

Claims (43)

1. An apparatus for performing image processing, the apparatus comprising:

a memory storing instructions; and

at least one processor configured to execute the instructions to:

input an image to a vision transformer comprising a plurality of encoders that correspond to at least one fixed encoder and a plurality of adaptive encoders;

process the image via the at least one fixed encoder to obtain image representations;

determine one or more layers of the plurality of adaptive encoders to drop, by inputting the image representations to a policy network configured to determine layer dropout actions for the plurality of adaptive encoders; and

obtain a class of the input image using remaining layers of the plurality of adaptive encoders other than the dropped one or more layers.

2. The apparatus of claim 1 , wherein each of the plurality of encoders comprises a multi-head self-attention (MSA) layer and a multilayer perceptron (MLP) layer.

3. The apparatus of claim 1 , wherein the layer dropout actions indicate whether each multi-head self-attention (MSA) layer and each multilayer perceptron (MLP) layer included in the plurality of adaptive encoders is dropped or not.

4. The apparatus of claim 1 , wherein the policy network comprises a first policy network configured to determine whether to drop one or more multi-head self-attention (MSA) layers, and a second policy network configured to determine whether to drop one or more multilayer perceptron (MLP) layers.

5. The apparatus of claim 4 , wherein the first policy network is configured to receive, as input, the image representations that are output from the at least one fixed encoder of the vision transformer, and output the layer dropout actions for each MSA layer of the plurality of adaptive encoders.

6. The apparatus of claim 5 , wherein the second policy network is further configured to receive, as input, the image representations and the layer dropout actions for each MSA layer, and output the layer dropout actions for each MLP layer of the plurality of adaptive encoders.

7. The apparatus of claim 6 , wherein the second policy network comprises a dense layer configured to receive, as input, a concatenation of the image representations and the layer dropout actions for each MSA layer.

8. The apparatus of claim 1 , wherein the policy network is configured to receive a reward that is calculated based on a number of the dropped one or more layers, and image classification prediction accuracy of the vision transformer.

9. The apparatus of claim 8 , wherein the at least one processor is configured to execute the instructions to:

calculate the reward using a reward function that increases the reward as the number of the dropped one or more layers increases and the image classification prediction accuracy increase.

10. A method of performing image processing, the method being performed by at least one processor, and the method comprising:

inputting an image to a vision transformer comprising a plurality of encoders that correspond to at least one fixed encoder and a plurality of adaptive encoders;

processing the image via the at least one fixed encoder to obtain image representations;

determining one or more layers of the plurality of adaptive encoders to drop, by inputting the image representations to a policy network configured to determine layer dropout actions for the plurality of adaptive encoders; and

obtaining a class of the input image using remaining layers of the plurality of adaptive encoders other than the dropped one or more layers.

11. The method of claim 10 , wherein each of the plurality of encoders comprises a multi-head self-attention (MSA) layer and a multilayer perceptron (MLP) layer.

12. The method of claim 10 , wherein the layer dropout actions indicate whether each multi-head self-attention (MSA) layer and each multilayer perceptron (MLP) layer included in the plurality of adaptive encoders is dropped or not.

13. The method of claim 10 , wherein the determining the one or more layers of the plurality of adaptive encoders to drop, comprises:

determining whether to drop one or more multi-head self-attention (MSA) layers, via a first policy network; and

determining whether to drop one or more multilayer perceptron (MLP) layers, via a second policy network.

14. The method of claim 13 , wherein the determining whether to drop the one or more multi-head self-attention (MSA) layers, comprises:

inputting the image representations that are output from the at least one fixed encoder of the vision transformer, to the first policy network; and

outputting the layer dropout actions for each MSA layer of the plurality of adaptive encoders, from the first policy network.

15. The method of claim 14 , wherein the determining whether to drop the one or more multilayer perceptron (MLP) layers, comprises:

inputting, to the second policy network, the image representations and the layer dropout actions for each MSA layer; and

outputting the layer dropout actions for each MLP layer of the plurality of adaptive encoders, from the second policy network.

16. The method of claim 15 , further comprising:

concatenating the image representations and the layer dropout actions for each MSA layer; and

inputting a concatenation of the image representations and the layer dropout actions for each MSA layer, to a dense layer of the second policy network.

17. The method of claim 10 , wherein the policy network is trained using a reward function that calculates a reward based on a number of the dropped one or more layers, and image classification prediction accuracy of the vision transformer.

18. The method of claim 17 , wherein the reward function increases the reward as the number of the dropped one or more layers increases and the image classification prediction accuracy increase.

19. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to:

input an image to a vision transformer comprising a plurality of encoders that correspond to at least one fixed encoder and a plurality of adaptive encoders;

process the image via the at least one fixed encoder to obtain image representations;

determine one or more of multi-head self-attention (MSA) layers and multilayer perceptron (MLP) layers of the plurality of adaptive encoders to drop, by inputting the image representations to a policy network configured to determine layer dropout actions for the plurality of adaptive encoders; and

obtain a class of the input image using remaining layers of the plurality of adaptive encoders other than the dropped one or more layers.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the policy network is trained using a reward function that increases a reward in direct proportion to a number of the dropped one or more layers and image classification prediction accuracy of the vision transformer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2022
From: UZKENT, BURAK; RAMANISHKA, VASILI; SHEN, YILIN; JIN, HONGXIA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 059343/0660 →
Continuity (2)
Provisional Application 63165500 · Mar 24, 2021
Related Publication 20220309774A1 · Sep 29, 2022
References Cited (25)
US 10789427B2 · Shazeer et al. · 2020 [cited by applicant]
US 10936907B2 · Suresh · 2021 [cited by examiner]
US 11093819B1 · Li et al. · 2021 [cited by applicant]
US 11138392B2 · Chen et al. · 2021 [cited by applicant]
US 11494660B2 · Chidlovskii · 2022 [cited by applicant]
US 11600087B2 · Chukka · 2023 [cited by examiner]
US 20200202168A1 · Mao et al. · 2020 [cited by applicant]
US 20200320402A1 · Yoon et al. · 2020 [cited by applicant]
US 20210150252A1 · Sarlin et al. · 2021 [cited by applicant]
US 20210255862A1 · Volkovs et al. · 2021 [cited by applicant]
US 20210294834A1 · Mai · 2021 [cited by examiner]
US 20210334475A1 · He et al. · 2021 [cited by applicant]
US 20220036564A1 · Ye et al. · 2022 [cited by applicant]
US 20220292341A1 · Mehta · 2022 [cited by examiner]
US 20230140474A1 · Ji · 2023 [cited by examiner]
US 20230169746A1 · Dwivedi · 2023 [cited by examiner]
CN 113658322A · 2021 [cited by applicant]
CN 113688813A · 2021 [cited by applicant]
CN 112861917B · 2021 [cited by applicant]
KR 1020160034814A · 2016 [cited by applicant]
Alexey Doscovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” arXiv:2010.11929v2, pp. 1-22, Jun. 3, 2021. [cited by applicant]
International Search Report and Written Opinion(PCT/ISA/210, PCT/ISA/220, and PCT/ISA/237), dated Dec. 20, 2022, issued by the International Searching Authority, Application No. PCT/KR2022/012888. [cited by applicant]
Extended European Search Report issued Apr. 10, 2025 in European Patent Application No. 22933766.2. [cited by applicant]
Lingchen Meng et al., “AdaViT: Adaptive Vision Transformers for Efficient Image Recognition”, 2021, XP093263160, pp. 1-12. [cited by applicant]
Zuxuan Wu et al., “BlockDrop: Dynamic Inference Paths in Residual Networks”, IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, XP033473806, pp. 8817-8826. [cited by applicant]