IP Library Granted Patent US 12,026,933
Granted Patent B2
US 12,026,933 · App. 18/032,794 · Granted Jul 2, 2024

Image recognition method and apparatus, and device and readable storage medium

Inventors: Baoyu Fan (Shandong, CN); Li Wang (Shandong, CN)
Assignee: INSPUR SUZHOU INTELLIGENT TECHNOLOGY CO., LTD.
G06V10/7715G06V10/751G06V10/774G06V10/776G06V10/82G06V20/70G06V40/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,026,933
App. No.
18/032,794
Granted
Jul 2, 2024
Kind
B2
Abstract

The present application discloses an image recognition method and apparatus, and a device and a readable storage medium. The method includes: obtaining a target image to be recognized; inputting the target image into a trained feature extraction model for feature extraction to obtain image features; and using the image features to recognize the target image to obtain a recognition result. In the present application, under the condition that the parameter quantity and the calculation amount of the feature extraction model do not need to be increased, the feature extraction model has relatively good feature extraction performance by means of isomorphic branch extraction and feature mutual mining, namely knowledge collaboration assistive training, and more accurate image recognition may be completed on the basis of the image features extracted by the feature extraction model.

Claims (183)

1. A method for picture identification, comprising:

obtaining a target picture to be identified;

inputting the target picture into a trained feature extraction model for feature extraction to obtain an image feature; and

performing identification on the target picture by using the image feature to obtain an identification result;

wherein training the feature extraction model comprises:

deriving homogeneous branches from a main network of a model to obtain homogeneous auxiliary training model;

inputting training samples to the homogeneous auxiliary training model in batches to obtain a sample image feature set corresponding to each of the homogeneous branches, training samples in each batch including a plurality of samples, and training samples in each batch correspond to a plurality of classes;

calculating a maximum inter-class distance and a minimum inter-class distance respectively corresponding to each of sample image features between every two sample image feature sets;

calculating a knowledge synergy loss value by using the maximum inter-class distance and the minimum inter-class distance;

adjusting parameters of the homogeneous auxiliary training model by using the knowledge synergy loss value until the homogeneous auxiliary training model converges; and

removing the homogeneous branches in the converged homogeneous auxiliary training model to obtain the feature extraction model including only a main branch.

2. The method for picture identification according to claim 1 , wherein deriving the homogeneous branches from the main network of the model to obtain the homogeneous auxiliary training model comprises:

auxiliarily derivating the homogeneous branches from the main network of the model; and/or

hierarchically derivating the homogeneous branches from the main network of the model.

3. The method for picture identification according to claim 1 , wherein calculating the knowledge synergy loss value by using the maximum inter-class distance and the minimum inter-class distance comprises:

calculating difference between a maximum inter-class distance and minimum inter-class distance, and accumulating the difference to obtain the knowledge synergy loss value.

4. The method for picture identification according to claim 1 , wherein adjusting parameters of the homogeneous auxiliary training model by using the knowledge synergy loss value until the homogeneous auxiliary training model converges comprises:

calculating a triplet loss value of the homogeneous auxiliary training model by using a triplet loss function;

determining a sum of the triple loss value and the knowledge synergy loss value as a total loss value of the homogeneous auxiliary training model; and

adjusting the parameters of the homogeneous auxiliary training model by using the total loss value until the homogeneous auxiliary training model converges.

5. The method for picture identification according to claim 4 , wherein calculating the triplet loss value of the homogeneous auxiliary training model by using the triplet loss function comprises:

traversing all training samples of each batch to calculate an absolute distance of intra-class difference of each sample in each batch; and

calculating the triplet loss value of the homogeneous auxiliary training model by using the absolute distance.

6. The method for picture identification according to claim 4 , wherein the triplet loss function is expressed as:

L

TriHard

b

=

-

1

N

a

=

1

N

[

max

y

p

=

y

a

d

(

f

e

a

,

f

e

p

)

-

min

y

n

y

a

d

(

f

e

a

,

f

e

n

)

+

m

]

+

wherein [⋅] + represents max d (⋅, 0) d (⋅, ⋅) represents calculating a distance between vectors; f e a =f (x a , θ b ), a represents an anchor, that is, an anchor sample; f e (⋅) represents obtaining a feature of an image in an Embedding layer of a network; f p represents an image feature of a same class as the anchor sample; f n represents an image feature of a different class from the anchor sample.

7. The method for picture identification according to claim 4 , wherein the knowledge synergy loss value is calculated by a knowledge synergy for hard sample mutual mining loss function, which is:

L

ksh

=

1

N

(

m

,

n

)

m

n

N

N

(

u

,

v

)

u

v

B

B

-

1

[

max

(

d

pos

(

f

em

u

,

f

en

v

)

)

-

min

(

d

neg

(

f

em

u

,

f

en

v

)

)

+

α

]

+

wherein f em u represents an embedding-layer feature of an m-th sample in a u-th branch, f en v represents an embedding-layer feature of an n-th sample in a v-th branch; d pos (⋅,⋅) represents calculating a distance between samples of a same class, and d neg (⋅,⋅) represents calculating a distance between samples of different classes; a is a hyperparameter and is a constant; [⋅] + represents max(⋅, 0).

8. The method for picture identification according to claim 1 , wherein performing identification on the target picture by using the image feature to obtain an identification result comprises:

calculating vector distances between the image feature and labeled image features in a query data set;

comparing the vector distances to obtain a minimum vector distance; and

determining a label corresponding to the labeled image feature corresponding to the minimum vector distance as the identification result.

9. The method for picture identification according to claim 8 , wherein obtaining the target picture to be identified comprises:

obtaining a pedestrian picture to be identified, and determining the pedestrian picture as the target picture;

accordingly, the identification result corresponding to pedestrian identity information.

10. An electronic device, comprising:

a memory for storing a computer program;

a processor for implementing steps of the method for picture identification according to claim 1 when the computer program is executed.

11. The electronic device according to claim 10 , wherein deriving the homogeneous branches from the main network of the model to obtain the homogeneous auxiliary training model comprises:

auxiliarily derivating the homogeneous branches from the main network of the model; and/or

hierarchically derivating the homogeneous branches from the main network of the model.

12. The electronic device according to claim 10 , wherein calculating the knowledge synergy loss value by using the maximum inter-class distance and the minimum inter-class distance comprises:

calculating difference between a respective maximum inter-class distance and minimum inter-class distance, and accumulating the difference to obtain the knowledge synergy loss value.

13. The electronic device according to claim 10 , wherein adjusting parameters of the homogeneous auxiliary training model by using the knowledge synergy loss value until the homogeneous auxiliary training model converges comprises:

calculating a triplet loss value of the homogeneous auxiliary training model by using a triplet loss function;

determining a sum of the triple loss value and the knowledge synergy loss value as a total loss value of the homogeneous auxiliary training model; and

adjusting the parameters of the homogeneous auxiliary training model by using the total loss value until the homogeneous auxiliary training model converges.

14. The electronic device according to claim 13 , wherein calculating the triplet loss value of the homogeneous auxiliary training model by using the triplet loss function comprises:

traversing all training samples of each batch to calculate an absolute distance of intra-class difference of each sample in each batch; and

calculating the triplet loss value of the homogeneous auxiliary training model by using the absolute distance.

15. The electronic device according to claim 10 , wherein performing identification on the target picture by using the image feature to obtain an identification result comprises:

calculating vector distances between the image feature and labeled image features in a query data set;

comparing the vector distances to obtain a minimum vector distance; and

determining a label corresponding to the labeled image feature corresponding to the minimum vector distance as the identification result.

16. The electronic device according to claim 15 , wherein obtaining the target picture to be identified comprises:

obtaining a pedestrian picture to be identified, and determining the pedestrian picture as the target picture;

accordingly, the identification result corresponding to pedestrian identity information.

17. A non-transitory readable storage medium, having a computer program stored thereon and the computer program, when executed by a processor, implementing steps of the method for picture identification according to claim 1 .

18. The non-transitory readable storage medium according to claim 17 , wherein deriving the homogeneous branches from the main network of the model to obtain the homogeneous auxiliary training model comprises:

auxiliarily derivating the homogeneous branches from the main network of the model; and/or

hierarchically derivating the homogeneous branches from the main network of the model.

19. The method for picture identification according to claim 1 , wherein the target picture comprises a human image, an image of an object, or an image of a monitored scenario.

20. The method for picture identification according to claim 1 , wherein the homogeneous auxiliary training model converges refers to a case where a loss value of the homogeneous auxiliary training model tends to be stable, or where the loss value is less than a preset threshold.

Assignments (2)
LICENSE Recorded Jun 30, 2026
From: IEIT SYSTEMS CO., LTD
To: AIVRES SYSTEMS INC.
Reel/Frame 075857/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2023
From: FAN, BAOYU; WANG, LI
To: INSPUR SUZHOU INTELLIGENT TECHNOLOGY CO., LTD.
Reel/Frame 063383/0262 →
Priority Claims (1)
CN 202110727794.1 · Jun 29, 2021 · national
Continuity (1)
Related Publication 20230316722A1 · Oct 5, 2023