Processing apparatus and method and storage medium
A processing apparatus includes a collection module and a training module, the training module includes a backbone network and a region proposal network (RPN) layer, the backbone network is connected to the RPN layer, and the RPN layer includes a class activation map (CAM) unit. The collection module is configured to obtain an image, where the image includes an image with an instance-level label and an image with an image-level label. The backbone network is used to output a feature map of the image based on the image obtained by the collection module.
1 . A processing apparatus comprising:
a collection system comprising a processor, wherein the processor is configured to obtain a first image comprising:
a second image comprising:
an instance-level label;
a class label; and
a bounding box of an object; and
a third image comprising:
an image-level label; and
the class label; and
a training system comprising:
a convolutional neural network configured to output, based on the first image, a feature map of the first image;
a region proposal network (RPN) layer coupled to the convolutional neural network, comprising a plurality of convolutional layers and a class activation map (CAM) component, and configured to:
determine, based on the feature map, a proposal region of the first image;
determine, using the CAM component, a CAM corresponding to the proposal region, wherein the CAM indicates a CAM response intensity of the proposal region;
determine, based on a weighted sum of a confidence of the proposal region and the CAM response intensity, a probability that the proposal region belongs to a foreground; and
determine, based on the probability, a feature of the proposal region; and
a plurality of heads, wherein each head is coupled to the RPN layer and configured to compute, based on the feature of the proposal region, multiple instance detection (MID) losses of the second image with the instance-level label and MID losses of the third image with the image-level label to train an object detection model.
2 . The processing apparatus of claim 1 , wherein the processor is further configured to:
receive the second image; and
obtain, based on the class label, the third image.
3 . The processing apparatus of claim 1 , wherein the CAM response intensity is directly proportional to the probability.
4 . The processing apparatus of claim 1 , wherein the CAM response intensity is based on a CAM heat map.
5 . The processing apparatus of claim 1 , wherein the MID loss determines a parameter of the training system.
6 . The processing apparatus of claim 5 , wherein the heads are coupled in series.
7 . The processing apparatus of claim 5 , wherein the RPN layer is further configured to calculate CAM losses of the second image and the third image, and wherein each of the CAM losses determines the parameter.
8 . The processing apparatus of claim 7 , wherein the RPN layer is further configured to calculate an RPN loss of the second image, wherein each of the N heads is further configured to calculate a proposal loss of the second image, a classification loss of the second image, and a regression loss of the second image, and wherein the RPN loss, the proposal loss, the classification loss, and the regression loss determine the parameter.
9 . The processing apparatus of claim 1 , wherein the processing apparatus is a cloud server.
10 . An object detection apparatus comprising:
a convolutional neural network configured to:
receive a to-be-detected image; and
output a feature map of the to-be-detected image;
a region proposal network (RPN) layer coupled to the convolutional neural network, comprising a plurality of convolutional layers and a class activation map (CAM) component, and configured to:
determine, based on the feature map, a proposal region of the to-be-detected image;
determine, using the CAM component, a CAM corresponding to the proposal region, wherein the CAM indicates a CAM response intensity of the proposal region;
determine, based on a weighted sum of a confidence of the proposal region and the CAM response intensity, a probability that the proposal region belongs to a foreground; and
determine, based on the probability, a feature of the proposal region; and
a plurality of heads, wherein each head is coupled to the RPN layer and configured to compute, based on the feature of the proposal region, multiple instance detection (MID) losses associated with the second image having the instance-level label and MID losses of the third image having the image-level label, to train an object detection model.
11 . The object detection apparatus of claim 10 , wherein the CAM response intensity is directly proportional to the probability.
12 . The object detection apparatus of claim 10 , wherein the CAM response intensity is based on a CAM heat map.
13 . The object detection apparatus of claim 10 , wherein any one of the heads is further configured to output, based on the probability, a detection result comprising a class of an object and a bounding box of the object.
14 . The object detection apparatus of claim 13 , wherein the heads are coupled in series.
15 . A processing method, comprising:
obtaining, by a collection system of a processing apparatus, a first image, wherein the first image comprises a second image and a third image, wherein the second image comprises an instance-level label, a class label, and a bounding box of an object, and wherein the third image comprises an image-level label and the class label;
outputting, by a convolutional neural network (CNN) of a training system of the processing apparatus and based on the first image, a feature map of the first image;
determining, by a region proposal network (RPN) layer of the training system and based on the feature map, a proposal region of the first image;
determining a class activation map (CAM) corresponding to the proposal region, wherein the CAM indicates a CAM response intensity of the proposal region;
determining, based on a weighted sum of a confidence of the proposal region and the CAM response intensity, a probability that the proposal region belongs to a foreground;
determining, based on the probability, a feature of the proposal region; and
computing, by a plurality of heads coupled to the RPN layer and based on the feature of the proposal region, multiple instance detection (MID) losses associated with the second image having the instance-level label and MID losses of the third image having the image-level label to train an object detection model.
16 . The processing method of claim 15 , further comprising:
receiving the second image; and
obtaining, based on the class label, the third image.
17 . The processing method of claim 15 , wherein the CAM response intensity is directly proportional to the probability.
18 . The processing method of claim 15 , wherein the CAM response intensity is based on a CAM heat map of the proposal region.
19 . The processing method of claim 15 , wherein each of the MID losses determines a parameter of a training system.
20 . The processing method of claim 19 , further comprising calculating CAM losses of the second image and the third image, wherein each of the CAM losses determines the parameter.
21 . The processing method of claim 20 , further comprising:
calculating a region proposal network (RPN) loss of the second image; and
calculating a proposal loss of the second image, a classification loss of the second image, and a regression loss of the second image, wherein the RPN loss, the proposal loss, the classification loss, and the regression loss determine the parameter.
22 . An object detection method, comprising:
receiving, by a convolutional neural network (CNN) of a training system of a processing apparatus, a to-be-detected image from a collection system of the processing apparatus;
outputting a feature map of the to-be-detected image;
determining, by a region proposal network (RPN) layer of the training system and based on the feature map, a proposal region of the to-be-detected image;
determining a class activation map (CAM) corresponding to the proposal region, wherein the CAM indicates a CAM response intensity of the proposal region; and
determining, based on a weighted sum of a confidence of the proposal region and the CAM response intensity, a probability that the proposal region belongs to a foreground;
determining, based on the probability, a feature of the proposal region; and
computing, by a plurality of heads coupled to the RPN layer and based on the feature of the proposal region, multiple instance detection (MID) losses associated with the second image having the instance-level label and MID losses of the third image having the image-level label to train an object detection model.
23 . The object detection method of claim 22 , wherein the CAM response intensity is directly proportional to the probability.
24 . The object detection method of claim 22 , wherein the CAM response intensity is based on a CAM heat map.
25 . The object detection method of claim 22 , further comprising outputting, based on the probability, a detection result comprising a class of an object and a bounding box of the object.