IP Library Granted Patent US 12,499,660
Granted Patent B2
US 12,499,660 · App. 18/047,780 · Granted Dec 16, 2025

Method and apparatus for training a neural network, image recognition method and storage medium

Inventors: Meng Zhang (Beijing, CN); Rujie Liu (Beijing, CN)
Assignee: Fujitsu Limited
G06V10/774G06V10/776G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,660
App. No.
18/047,780
Granted
Dec 16, 2025
Kind
B2
Abstract

A method and an apparatus for training a neural network, an image recognition method and a computer readable storage medium are disclosed. The neural network includes a first model and a second model. The method for training a neural network includes: acquiring a second image from a first image, wherein a quality of the second image is lower than that of the first image; inputting the first image into the first model of the neural network, and inputting the second image into the second model of the neural network; calculating an attention map and a gradient map of the first model and an attention map and a gradient map of the second model; constructing a loss function based on a matrix of a dot product of the gradient map and the attention map of the first model and a matrix of a dot product of the gradient map and the attention map of the second model; and training the neural network by minimizing the loss function.

Claims (29)

1 . A method for training a neural network comprising a first model and a second model, the method comprising:

acquiring a second image from a first image, wherein a quality of the second image is lower than that of the first image;

inputting the first image into the first model of the neural network, and inputting the second image into the second model of the neural network;

calculating an attention map and a gradient map of the first model and an attention map and a gradient map of the second model;

constructing a loss function based on a square of a difference between a matrix of a dot product of the gradient map and the attention map of the first model and a matrix of a dot product of the gradient map and the attention map of the second model; and

training the neural network by minimizing the loss function.

2 . The method according to claim 1 , further comprising: after calculating of attention map, softening the attention map of the first model and the attention map of the second model.

3 . The method according to claim 2 , wherein the loss function is constructed as a square of a difference between a matrix of a dot product of the gradient map and a softened attention map of the first model and a matrix of a dot product of the gradient map and a softened attention map of the second model.

4 . The method according to claim 1 , wherein the first model and the second model are two symmetrical branches of the neural network.

5 . The method according to claim 4 , wherein the first model and the second model each comprise one or more convolutional layers and one or more fully connected layers.

6 . The method according to claim 1 , wherein the matrix is a Gram matrix.

7 . The method according to claim 1 , further comprising: training the neural network by using the loss function, a knowledge distillation loss function, and a classification loss function.

8 . The method according to claim 1 , wherein the first image and the second image include a face.

9 . An image recognition method, comprising:

inputting an image to be recognized into the second model of the neural network trained by the method according to claim 1 for recognition.

10 . An apparatus for training a neural network comprising a first model and a second model, comprising:

an acquisition means configured to acquire a second image from a first image, wherein a quality of the second image is lower than that of the first image;

an input means configured to input the first image into the first model of the neural network, and input the second image into the second model of the neural network;

a calculation means configured to calculate an attention map and a gradient map of the first model and an attention map and a gradient map of the second model; and

a construction means configured to construct a loss function based on a square of a difference between a matrix of a dot product of the gradient map and the attention map of the first model and a matrix of a dot product of the gradient map and the attention map of the second model,

wherein the neural network is trained by minimizing the loss function.

11 . The apparatus according to claim 10 , further including: a softening means configured to soften, after calculating of attention map, the attention map of the first model and the attention map of the second model.

12 . The apparatus according to claim 11 , wherein the loss function is constructed as a square of a difference between a matrix of a dot product of the gradient map and a softened attention map of the first model and a matrix of a dot product of the gradient map and a softened attention map of the second model.

13 . The apparatus according to claim 10 , wherein the first model and the second model are two symmetrical branches of the neural network.

14 . The apparatus according to claim 13 , wherein the first model and the second model each include one or more convolution layers and one or more fully connected layers.

15 . The apparatus according to claim 10 , wherein the matrix is a Gram matrix.

16 . The apparatus according to claim 10 , wherein the neural network is trained by using the loss function, a knowledge distillation loss function, and a classification loss function.

17 . The apparatus according to claim 10 , wherein the first image and the second image include a face.

18 . A non-transitory computer readable medium storing a program which can be executed by a processor to perform the method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2022
From: ZHANG, MENG; LIU, RUJIE
To: FUJITSU LIMITED
Reel/Frame 061486/0102 →
Priority Claims (1)
CN 202111581419.7 · Dec 22, 2021 · national
Continuity (1)
Related Publication 20230196735A1 · Jun 22, 2023
References Cited (13)
US 11373274B1 · Yoon · 2022 [cited by examiner]
US 11586925B2 · Chang · 2023 [cited by examiner]
CN 107977932A · 2018 [cited by applicant]
CN 109543548A · 2019 [cited by applicant]
CN 111429436A · 2020 [cited by applicant]
CN 112052945A · 2020 [cited by applicant]
CN 113076980A · 2021 [cited by applicant]
CN 113723174A · 2021 [cited by applicant]
English translation of CN107977932 by Li et al. (Year: 2021). [cited by examiner]
Chen, C., et al., “Progressive Semantic-Aware Style Transformation for Blind Face Restoration”, arxiv.org, Cornell University Library, XP081897692, (21 Pages Total), (Mar. 21, 2021). [cited by applicant]
Zangeneh, E. et al., “Low Resolution Face Recognition Using a Two-Branch Deep Convolutional Neural Network Architecture”, arxiv.org, Cornell University Library, XP080771113, (11 Pages Total), (Jun. 20, 2017). [cited by applicant]
Communication from the European Patent Office in European Application No. 22206074.1, dated May 9, 2023. [cited by applicant]
Chinese Office Action mailed May 19, 2025, for corresponding Chinese Patent Application No. 202111581419.7, with Machine Translation (CNOA). [cited by applicant]