Information processing apparatus which reduces computation operations, information processing method, and computer readable medium
An object is to provide an information processing apparatus capable of reducing redundant computation in CNN. An information processing apparatus according to the present disclosure includes at least one memory configured to store an instruction, and at least one processor configured to execute the instruction to use mask channel in input feature maps to mask pixels of feature channels in the input feature maps and to generate masked feature channels, and perform a convolution operation between the masked feature channels and convolution kernel to generate output feature maps.
1 . An information processing apparatus comprising at least one memory configured to store an instruction, and at least one processor configured to execute the instruction to:
use at least one mask channel derived from one or more input feature maps of a set to mask pixels of at least one feature channel derived from the one or more input feature maps of the set and to generate at least one masked feature channel;
perform a convolution operation between the at least one masked feature channel and convolution kernels to generate output feature maps;
calculate task loss from a prediction and groundtruth data of an image;
calculate a mask loss from mask channels of the output feature maps and groundtruth mask of the image;
calculate a total loss from the task loss and the mask loss; and
train a convolutional neural network based on the total loss to obtain an updated convolutional neural network.
2 . The information processing apparatus according to claim 1 , wherein
the at least one processor is further configured to split the input feature maps into the mask channels and the at least one feature channel.
3 . The information processing apparatus according to claim 1 , wherein
the at least one processor is further configured to process the output feature maps.
4 . The information processing apparatus according to claim 1 , wherein
the at least one processor is further configured to generate the input feature maps using an image data.
5 . The information processing apparatus according to claim 1 ,
wherein the at least one processor is further configured to:
store the convolution kernels in convolution kernel storage, the convolution kernels including one or a plurality of kernels of mask channels for generating mask channels of the output feature maps and one or a plurality of kernels of feature channels for generating feature channels of the output feature maps; and
perform convolution with the kernels in the convolution kernel storage across the masked feature channels.
6 . The information processing apparatus according to claim 1 , wherein
the output feature maps are predictions of the image.
7 . The information processing apparatus according to claim 1 ,
wherein the at least one processor is further configured to:
generate groundtruth mask from groundtruth BBox data; and
calculate the mask loss from the generated groundtruth mask and the mask channels of the output feature maps.
8 . An information processing method comprising:
using at least one mask channel derived from one or more input feature maps of a set to mask pixels of at least one feature channel derived from the one or more input feature maps of the set and to generate at least one masked feature channel;
performing a convolution operation between the at least one masked feature channel and convolution kernels to generate output feature maps;
calculating task loss from a prediction and groundtruth data of an image;
calculating a mask loss from mask channels of the output feature maps and groundtruth mask of the image;
calculating a total loss from the task loss and the mask loss; and
training a convolutional neural network based on the total loss to obtain an updated convolutional neural network.
9 . A non-transitory computer readable medium storing a program for causing a computer to execute:
using at least one mask channel derived from one or more input feature maps of a set to mask pixels of at least one feature channel derived from the one or more input feature maps of the set and to generate at least one masked feature channel;
performing a convolution operation between the at least one masked feature channel and convolution kernels to generate output feature maps;
calculating task loss from a prediction and groundtruth data of an image;
calculating a mask loss from mask channels of the output feature maps and groundtruth mask of the image;
calculating a total loss from the task loss and the mask loss; and
training a convolutional neural network based on the total loss to obtain an updated convolutional neural network.