IP Library › Granted Patent US 11,361,192
Granted Patent B2
US 11,361,192 · App. 16/853,733 · Granted Jun 14, 2022

Image classification method, computer device, and computer-readable storage medium

Inventors: Pai Peng (Shenzhen, CN); Xiaowei Guo (Shenzhen, CN); Kailin Wu (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06K9/6263G06K9/6257G06N3/04G06N3/08G06V10/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,361,192
App. No.
16/853,733
Granted
Jun 14, 2022
Kind
B2
Abstract

Embodiments of the present disclosure provide an image classification method for a computer device. The method includes obtaining an original image and a category of an object included in the original image; adjusting a display parameter of the original image to satisfy a value condition to obtain an adjusted original image; and transforming the display parameter of the original image according to a distribution condition that distribution of the display parameter needs to satisfy, to obtain a transformed image. The method also includes training a neural network model based on the category of the object and a training set constructed by the adjusted original image and the transformed image; and determining a category of an object included in a to-be-predicted image based on the trained neural network model.

Claims (75)

1. An image classification method for a computer device, comprising:

obtaining a plurality of original images, each original image containing at least two objects, the at least two objects being objects that are same or symmetrical;

for each original image,

obtaining a category of each object contained in the original image;

adjusting a display parameter of the original image to satisfy a value condition to obtain an adjusted original image for the each object;

generating, for the each object, a transformed image from the original image by transforming a second display parameter of the original image according to a preset distribution condition that distribution of the display parameter needs to satisfy;

training a neural network model based on the categories of the objects in the original images and a training set constructed by the adjusted original images and the transformed images of the at least two objects, comprising:

inputting the training set into an initial neural network model for training;

respectively obtaining image features of the at least two objects output from a last average pooling layer of the initial neural network model;

concatenating the image features of the at least two objects at a cascade layer of a combined neural network model;

mapping the concatenated image features to a sample label space at a fully-connected layer of the combined neural network model, to obtain model weights;

connecting at least two classification layers of the combined neural network model to the fully-connected layer, each classification layer corresponding to one of the at least two objects and outputting probabilities of the corresponding object belonging to different categories;

iteratively training the combined neural network model until a loss function satisfies a convergence condition, to obtain the trained neural network model; and

determining a category of an object contained in a to-be-predicted image based on the trained neural network model.

2. The image classification method according to claim 1 , further comprising:

obtaining the loss function of the neural network model by combining a K loss function and a multi-categorization logarithmic loss function at a preset ratio.

3. The image classification method according to claim 1 , wherein the determining a category of an object contained in a to-be-predicted image based on the trained neural network model comprises:

correspondingly extracting image features of at least two objects included in the to-be-predicted image using the combined neural network model;

concatenating the extracted image features of the at least two objects and performing down-sampling processing; and

respectively mapping the image features from the down-sampling process to a probability that each of the at least two objects belong to different categories.

4. The image classification method according to claim 1 , wherein the adjusting a display parameter of the original image to satisfy a value condition comprises:

detecting an imaging region of the object contained in the original image; and

adjusting dimensions of the original image to be consistent with dimensions of the imaging region of the object contained in the original image.

5. The image classification method according to claim 1 , wherein the adjusting a display parameter of the original image to satisfy a value condition comprises:

performing image enhancement processing on each color channel of the original image based on a recognition degree that the original image needs to satisfy.

6. The image classification method according to claim 1 , wherein the adjusting a display parameter of the original image to satisfy a value condition comprises:

cropping a non-imaging region of the object in the original image; and

adjusting the cropped image to preset dimensions.

7. The image classification method according to claim 1 , wherein the transforming the display parameter of the original image according to a distribution condition that distribution of the display parameter needs to satisfy, to obtain a transformed image comprises:

determining, according to a value space in which a display parameter of at least one category of the original image is located and a distribution condition that the value space satisfies, a missing display parameter according to the display parameter of the original image compared with the distribution condition; and

transforming the display parameter of the original image into the missing display parameter, to obtain the transformed image.

8. A computer device, comprising:

a memory storing computer-readable instructions; and

a processor coupled to the memory for executing the computer-readable instructions to perform:

obtaining a plurality of original images, each original image containing at least two objects, the at least two objects being objects that are same or symmetrical;

for each original image,

obtaining a category of each object contained in the original image;

adjusting a display parameter of the original image to satisfy a value condition to obtain an adjusted original image for the each object;

generating, for the each object, a transformed image from the original image by transforming a second display parameter of the original image according to a preset distribution condition that distribution of the display parameter needs to satisfy;

training a neural network model based on the categories of the objects in the original images and a training set constructed by the adjusted original images and the transformed images of the at least two objects, comprising:

inputting the training set into an initial neural network model for training;

respectively obtaining image features of the at least two objects output from a last average pooling layer of the initial neural network model;

concatenating the image features of the at least two objects at a cascade layer of a combined neural network model;

mapping the concatenated image features to a sample label space at a fully-connected layer of the combined neural network model, to obtain model weights;

connecting at least two classification layers of the combined neural network model to the fully-connected layer, each classification layer corresponding to one of the at least two objects and outputting probabilities of the corresponding object belonging to different categories;

iteratively training the combined neural network model until a loss function satisfies a convergence condition, to obtain the trained neural network model; and

determining a category of an object contained in a to-be-predicted image based on the trained neural network model.

9. The computer device according to claim 8 , wherein the processor further performs:

obtaining the loss function of the neural network model by combining a K loss function and a multi-categorization logarithmic loss function at a preset ratio.

10. The computer device according to claim 8 , wherein the determining a category of an object contained in a to-be-predicted image based on the trained neural network model comprises:

in the combined neural network model, correspondingly extracting image features of at least two objects included in the to-be-predicted image;

concatenating the extracted image features of the at least two objects and performing down-sampling processing on the extracted image feature; and

respectively mapping the image features from the down-sampling process to a probability that each of the at least two objects belong to different categories.

11. A non-transitory computer-readable storage medium storing computer-readable instructions executable by at least one processor to perform:

obtaining a plurality of original images, each original image containing at least two objects, the at least two objects being objects that are same or symmetrical;

for each original image,

obtaining a category of each object contained in the original image;

adjusting a display parameter of the original image to satisfy a value condition to obtain an adjusted original image for the each object;

generating, for the each object, a transformed image from the original image by transforming a second display parameter of the original image according to a preset distribution condition that distribution of the display parameter needs to satisfy;

training a neural network model based on the categories of the objects in the original images and a training set constructed by the adjusted original images and the transformed images of the at least two objects, comprising:

inputting the training set into an initial neural network model for training;

respectively obtaining image features of the at least two objects output from a last average pooling layer of the initial neural network model;

concatenating the image features of the at least two objects at a cascade layer of a combined neural network model;

mapping the concatenated image features to a sample label space at a fully-connected layer of the combined neural network model, to obtain model weights;

connecting at least two classification layers of the combined neural network model to the fully-connected layer, each classification layer corresponding to one of the at least two objects and outputting probabilities of the corresponding object belonging to different categories;

iteratively training the combined neural network model until a loss function satisfies a convergence condition, to obtain the trained neural network model; and

determining a category of an object contained in a to-be-predicted image based on the trained neural network model.

12. The non-transitory computer-readable storage medium according to claim 11 , wherein the computer-readable instructions are executable by the at least one processor to perform:

obtaining the loss function of the neural network model by combining a K loss function and a multi-categorization logarithmic loss function at a preset ratio.

13. The non-transitory computer-readable storage medium according to claim 11 , wherein the determining a category of an object contained in a to-be-predicted image based on the trained neural network model comprises:

in the combined neural network model, correspondingly extracting image features of at least two objects included in the to-be-predicted image;

concatenating the extracted image features of the at least two objects and performing down-sampling processing on the extracted image feature; and

respectively mapping the image features from the down-sampling process to a probability that each of the at least two objects belong to different categories.

14. The method according to claim 1 , wherein the preset distribution condition comprises: a preset value range that limits a value of a transformed display parameter of the transformed image and a statistical distribution type that the transformed display parameter needs to satisfy with respect to the second display parameter of the original image, each transformed image containing one of the at least two objects.

15. The method according to claim 14 , wherein the statistical distribution type is one of average distribution, random distribution, and Gaussian distribution; and the second display parameter includes at least one of a direction, dimensions, brightness, contrast, an aspect ratio, resolution, or a color.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2020
From: PENG, PAI; GUO, XIAOWEI; WU, KAILIN
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 052447/0799 →
Priority Claims (1)
CN 201711060265.0 · Nov 1, 2017 · national
Continuity (2)
Continuation PCTCN2018111491 · Oct 23, 2018
Related Publication 20200250491A1 · Aug 6, 2020
Cited By (2)
US 12,242,964 US 12,731,391