IP Library › Granted Patent US 12,367,707
Granted Patent B2
US 12,367,707 · App. 17/800,481 · Granted Jul 22, 2025

Method of training an image classification model

Inventors: Jiankang Deng (London, GB); Stefanos Zafeiriou (London, GB)
Assignee: Huawei Technologies Co., Ltd.
G06V40/172G06N3/08G06V10/761G06V10/764G06V10/82G06V30/18057G06V30/18105G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,707
App. No.
17/800,481
Granted
Jul 22, 2025
Kind
B2
Abstract

A method of training a neural network for classifying an image into one of a plurality of classes, the method comprising: extracting, from the neural network, a plurality of subclass center vectors for each class; inputting an image into the neural network, wherein the image is associated with a predetermined class; generating, using the neural network, an embedding vector corresponding to the input image; determining a similarity score between the embedding vector and each of the plurality of subclass center vectors; updating parameters of the neural network in dependence on a plurality of the similarity scores using an objective function; extracting a plurality of updated parameters from the neural network; and updating each subclass center vector in dependence on the extracted updated parameters.

Claims (66)

1. A method of training a neural network for classifying an image into one of a plurality of classes, wherein the method is applied to a computer device comprising at least one processor, and the method comprises:

extracting, from the neural network, a plurality of subclass center vectors for each class;

inputting an image into the neural network, wherein the image is associated with a predetermined class;

generating, using the neural network, an embedding vector corresponding to the input image;

determining a similarity score between the embedding vector and each of the plurality of subclass center vectors;

updating parameters of the neural network based on a plurality of the similarity scores using an objective function;

extracting a plurality of updated parameters from the neural network;

updating each subclass center vector based on the extracted updated parameters; and

determining a closest subclass center vector for each class using the similarity scores,

wherein the objective function comprises a multi-center loss term comparing the similarity score between the embedding vector and the closest subclass center vector from the predetermined class to the similarity scores between the embedding vector and the closest subclass center vectors from each of the other classes.

2. The method of claim 1 , further comprising: prior to updating the parameters of the neural network:

inputting a further image into the neural network, wherein the further image is associated with a predetermined class;

generating, using the neural network, a further embedding vector corresponding to the further input image;

determining a further similarity score between the further embedding vector and each of the plurality of subclass center vectors;

wherein updating the parameters of the neural network is further based on the further similarity scores.

3. The method of claim 1 , wherein the multi-center loss term is a margin-based softmax loss function.

4. The method of claim 1 , wherein the embedding vector and each subclass center vector are normalized, and wherein the similarity score is an angle between the embedding vector and one of the subclass center vectors.

5. The method of claim 1 , wherein each class comprises a dominant subclass, and the method further comprises:

for each class, determining an intra-class similarity score between a dominant subclass center vector and each of the other subclass center vectors in the class,

wherein the objective function comprises an intra-class compactness term that uses the intra-class similarity scores.

6. The method of claim 5 , wherein each subclass center vector is normalized, and the intra-class similarity score is an angle between the dominant subclass center vector and another subclass center vector in the class.

7. The method of claim 1 , wherein the neural network comprises a plurality of connected layers, and wherein each subclass center vector is updated using updated parameters extracted from a last fully connected layer of the neural network.

8. The method of claim 1 , wherein each class comprises a dominant subclass, and the method further comprises:

discarding a non-dominant subclass from a class based on the subclass center vector of the non-dominant subclass being above a threshold distance from the dominant subclass center vector of the class.

9. The method of claim 8 , wherein the discarding the non-dominant subclass from the class is performed based on a threshold condition being satisfied.

10. The method of claim 9 , wherein the threshold condition is a first threshold number of training epochs being exceeded.

11. The method of claim 1 , further comprising:

discarding all non-dominant subclasses based on a further threshold condition being satisfied.

12. The method of claim 11 , wherein the further threshold condition is a second threshold number of training epochs being exceeded.

13. The method of claim 1 , wherein the image is a face image.

14. The method of claim 1 , wherein the class corresponds to a classification condition of the image, and

wherein the image is from a batch that contains label noise such that the batch comprises at least one image labelled to relate to a class that does not correspond to the classification condition of the at least one image.

15. An apparatus comprising:

one or more processors; and

a memory, the memory comprising computer readable instructions that are executed by the one or more processors, and cause the apparatus to perform a method of training a neural network for classifying an image into one of a plurality of classes, the method including:

extracting, from the neural network, a plurality of subclass center vectors for each class;

inputting an image into the neural network, wherein the image is associated with a predetermined class;

generating, using the neural network, an embedding vector corresponding to the input image;

determining a similarity score between the embedding vector and each of the plurality of subclass center vectors;

updating parameters of the neural network based on a plurality of the similarity scores using an objective function;

extracting a plurality of updated parameters from the neural network;

updating each subclass center vector based on the extracted updated parameters; and

determining a closest subclass center vector for each class using the similarity scores,

wherein the objective function comprises a multi-center loss term comparing the similarity score between the embedding vector and the closest subclass center vector from the predetermined class to the similarity scores between the embedding vector and the closest subclass center vectors from each of the other classes.

16. A non-transitory computer-readable medium comprising computer readable instructions that are executed by a processor of a computing device, and cause the computing device to perform a method of training a neural network for classifying an image into one of a plurality of classes, the method including:

extracting, from the neural network, a plurality of subclass center vectors for each class;

inputting an image into the neural network, wherein the image is associated with a predetermined class;

generating, using the neural network, an embedding vector corresponding to the input image;

determining a similarity score between the embedding vector and each of the plurality of subclass center vectors;

updating parameters of the neural network based on a plurality of the similarity scores using an objective function;

extracting a plurality of updated parameters from the neural network;

updating each subclass center vector based on the extracted updated parameters; and

determining a closest subclass center vector for each class using the similarity scores,

wherein the objective function comprises a multi-center loss term comparing the similarity score between the embedding vector and the closest subclass center vector from the predetermined class to the similarity scores between the embedding vector and the closest subclass center vectors from each of the other classes.

17. The apparatus of claim 15 , wherein the method further comprises: prior to updating the parameters of the neural network,

inputting a further image into the neural network, wherein the further image is associated with a predetermined class;

generating, using the neural network, a further embedding vector corresponding to the further input image;

determining a further similarity score between the further embedding vector and each of the plurality of subclass center vectors;

wherein updating the parameters of the neural network is further based on the further similarity scores.

18. The non-transitory computer-readable medium of claim 16 , wherein the method further comprises: prior to updating the parameters of the neural network,

inputting a further image into the neural network, wherein the further image is associated with a predetermined class;

generating, using the neural network, a further embedding vector corresponding to the further input image;

determining a further similarity score between the further embedding vector and each of the plurality of subclass center vectors;

wherein updating the parameters of the neural network is further based on the further similarity scores.

19. The apparatus of claim 15 , wherein the multi-center loss term is a margin-based softmax loss function.

20. The non-transitory computer-readable medium of claim 16 , wherein the multi-center loss term is a margin-based softmax loss function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2025
From: DENG, JIANKANG; ZAFEIRIOU, STEFANOS
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 071425/0569 →
Priority Claims (1)
GB 2002157 · Feb 17, 2020 · national
Continuity (1)
Related Publication 20230085401A1 · Mar 16, 2023
References Cited (30)
US 20180165809A1 · Stanitsas · 2018 [cited by examiner]
US 20180365565A1 · Panciatici · 2018 [cited by examiner]
US 20190114544A1 · Sundaram et al. · 2019 [cited by applicant]
US 20190220699A1 · Gao · 2019 [cited by examiner]
US 20190220700A1 · Gao · 2019 [cited by examiner]
US 20190348062A1 · Gao · 2019 [cited by examiner]
US 20200210814A1 · Mehr · 2020 [cited by examiner]
US 20200226748A1 · Kaufman · 2020 [cited by examiner]
US 20200265304A1 · Nagaraj Chandrashekar · 2020 [cited by examiner]
US 20210224977A1 · Jia · 2021 [cited by examiner]
US 20220005332A1 · Metzler · 2022 [cited by examiner]
CN 107194464A · 2017 [cited by applicant]
CN 108062574A · 2018 [cited by applicant]
CN 109753589A · 2019 [cited by applicant]
CN 110147882A · 2019 [cited by applicant]
CN 110782044A · 2020 [cited by applicant]
WO 2019197022A1 · 2019 [cited by applicant]
Tao Chen et al., “SS-HCNN: Semi-Supervised Hierarchical Convolutional Neural Network for Image Classification,” Jan. 30, 2019, IEEE Transactions on Image Processing, vol. 28, No. 5, May 2019, pp. 2389-2396. [cited by examiner]
Yunchao Wei et al., “HCP: A Flexible CNN Framework for Multi-Label Image Classification,” Aug. 11, 2016, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 38, No. 9, Sep. 2016, pp. 1901-1905. [cited by examiner]
Jiankang Deng et al., “ArcFace: Additive Angular Margin Loss for Deep Face Recognition ,” Jun. 2019, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4690-4696. [cited by examiner]
Huan Wan et al., “Separability-Oriented Subclass Discriminant Analysis,” Jan. 5, 2018, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, No. 2, Feb. 2018, pp. 409-420. [cited by examiner]
Jiankang Deng et al., “Sub-center ArcFace: Boosting Face Recognition by Large-Scale Noisy Web Faces,” Nov. 27, 2020, Computer Vision—ECCV 2020. ECCV 2020. Lecture Notes in Computer Science(), vol. 12356. Springer, Cham.… [cited by examiner]
Ying-Nong Chen et al., “Face Recognition Using Nearest Feature Space Embedding,” Nov. 9, 2010, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 33, No. 6, Jun. 2011, pp. 1073-1084. [cited by examiner]
Muhammad Imran Razzak et al., “Face Recognition using Layered Linear Discriminant Analysis and Small Subspace,” Sep. 16, 2010,2010 10th IEEE International Conference on Computer and Information Technology, pp. 1-3. [cited by examiner]
Björn Barz et al., “Hierarchy-based Image Embeddings for Semantic Image Retrieval,” Mar. 7, 2019,2019 IEEE Winter Conference on Applications of Computer Vision, pp. 638-645. [cited by examiner]
Tianqi Chen et al., “MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems,” Dec. 3, 2015, arXiv:1512.01274v1 [cs. DC] Dec. 3, 2015, pp. 1-4. [cited by examiner]
Guiguang Ding et al., “DECODE: Deep Confidence Network for Robust Image Classification,” Jun. 13, 2019, IEEE Transactions on Image Processing, vol. 28, No. 8, Aug. 2019, pp. 3752-3758. [cited by examiner]
Wei Hu et al., “Noise-Tolerant Paradigm for Training Face Recognition CNNs,” Jun. 2019, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 11887-11893. [cited by examiner]
Deng et al., “ArcFace: Additive Angular Margin Loss for Deep Face Recognition,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, pp. 4685-4694, Institute of Electrical and Electronics En… [cited by applicant]
Wen et al., “A Discriminative Feature Learning Approach for Deep Face Recognition,” Computer Vision—ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, Springer International Publishing, XP093253685, total … [cited by applicant]