IP Library › Granted Patent US 12,488,568
Granted Patent B2
US 12,488,568 · App. 18/007,784 · Granted Dec 2, 2025

Information processing apparatus which reduces computation operations, information processing method, and computer readable medium

Inventor: Salita Sombatsiri (Tokyo, JP)
Assignee: NEC CORPORATION
G06V10/771G06T5/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,568
App. No.
18/007,784
Granted
Dec 2, 2025
Kind
B2
Abstract

An object is to provide an information processing apparatus capable of reducing redundant computation in CNN. An information processing apparatus according to the present disclosure includes at least one memory configured to store an instruction, and at least one processor configured to execute the instruction to use mask channel in input feature maps to mask pixels of feature channels in the input feature maps and to generate masked feature channels, and perform a convolution operation between the masked feature channels and convolution kernel to generate output feature maps.

Claims (37)

1 . An information processing apparatus comprising at least one memory configured to store an instruction, and at least one processor configured to execute the instruction to:

use at least one mask channel derived from one or more input feature maps of a set to mask pixels of at least one feature channel derived from the one or more input feature maps of the set and to generate at least one masked feature channel;

perform a convolution operation between the at least one masked feature channel and convolution kernels to generate output feature maps;

calculate task loss from a prediction and groundtruth data of an image;

calculate a mask loss from mask channels of the output feature maps and groundtruth mask of the image;

calculate a total loss from the task loss and the mask loss; and

train a convolutional neural network based on the total loss to obtain an updated convolutional neural network.

2 . The information processing apparatus according to claim 1 , wherein

the at least one processor is further configured to split the input feature maps into the mask channels and the at least one feature channel.

3 . The information processing apparatus according to claim 1 , wherein

the at least one processor is further configured to process the output feature maps.

4 . The information processing apparatus according to claim 1 , wherein

the at least one processor is further configured to generate the input feature maps using an image data.

5 . The information processing apparatus according to claim 1 ,

wherein the at least one processor is further configured to:

store the convolution kernels in convolution kernel storage, the convolution kernels including one or a plurality of kernels of mask channels for generating mask channels of the output feature maps and one or a plurality of kernels of feature channels for generating feature channels of the output feature maps; and

perform convolution with the kernels in the convolution kernel storage across the masked feature channels.

6 . The information processing apparatus according to claim 1 , wherein

the output feature maps are predictions of the image.

7 . The information processing apparatus according to claim 1 ,

wherein the at least one processor is further configured to:

generate groundtruth mask from groundtruth BBox data; and

calculate the mask loss from the generated groundtruth mask and the mask channels of the output feature maps.

8 . An information processing method comprising:

using at least one mask channel derived from one or more input feature maps of a set to mask pixels of at least one feature channel derived from the one or more input feature maps of the set and to generate at least one masked feature channel;

performing a convolution operation between the at least one masked feature channel and convolution kernels to generate output feature maps;

calculating task loss from a prediction and groundtruth data of an image;

calculating a mask loss from mask channels of the output feature maps and groundtruth mask of the image;

calculating a total loss from the task loss and the mask loss; and

training a convolutional neural network based on the total loss to obtain an updated convolutional neural network.

9 . A non-transitory computer readable medium storing a program for causing a computer to execute:

using at least one mask channel derived from one or more input feature maps of a set to mask pixels of at least one feature channel derived from the one or more input feature maps of the set and to generate at least one masked feature channel;

performing a convolution operation between the at least one masked feature channel and convolution kernels to generate output feature maps;

calculating task loss from a prediction and groundtruth data of an image;

calculating a mask loss from mask channels of the output feature maps and groundtruth mask of the image;

calculating a total loss from the task loss and the mask loss; and

training a convolutional neural network based on the total loss to obtain an updated convolutional neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2023
From: SOMBATSIRI, SALITA
To: NEC CORPORATION
Reel/Frame 065292/0721 →
Continuity (1)
Related Publication 20230237770A1 · Jul 27, 2023
References Cited (34)
US 10325179B1 · Kim · 2019 [cited by examiner]
US 11551027B2 · Fu · 2023 [cited by examiner]
US 11593587B2 · Lee · 2023 [cited by examiner]
US 20180365794A1 · Lee · 2018 [cited by examiner]
US 20190304102A1 · Chen et al. · 2019 [cited by applicant]
US 20190311202A1 · Lee et al. · 2019 [cited by applicant]
US 20190379589A1 · Ryan et al. · 2019 [cited by applicant]
US 20200012904A1 · Zhao et al. · 2020 [cited by applicant]
US 20200143204A1 · Nakano · 2020 [cited by examiner]
US 20210264557A1 · Mao · 2021 [cited by examiner]
US 20220351333A1 · Navarrete Michelini · 2022 [cited by examiner]
US 20230206456A1 · Lee · 2023 [cited by examiner]
US 20230410481A1 · Guo · 2023 [cited by examiner]
US 20240013504A1 · Yu · 2024 [cited by examiner]
US 20250046063A1 · Badowski · 2025 [cited by examiner]
CN 115497004A · 2022 [cited by examiner]
CN 115546474A · 2022 [cited by examiner]
JP 2020064333A · 2020 [cited by applicant]
Anvarov, & Kim, & Song,. (2020). Action Recognition Using Deep 3D CNNs with Sequential Feature Aggregation and Attention. Electronics. 9. 147. 10.3390/electronics9010147. (Year: 2020). [cited by examiner]
He, K., Gkioxari, G., Dollár, P., & Girshick, R. (2017). Mask r-cnn. In Proceedings of the IEEE international conference on computer vision (pp. 2961-2969). (Year: 2017). [cited by examiner]
Sombatsiri, S., Shibata, S., Kobayashi, Y., Inoue, H., Takenaka, T., Hosomi, T., . . . & Takeuchi, Y. (2019). Parallelism-flexible convolution core for sparse convolutional neural networks on FPGA. IPSJ Transactions on … [cited by examiner]
Ren, S., He, K., Girshick, R., & Sun, J. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28. (Year: 2015). [cited by examiner]
C. -F. Liou, P. Kuo and J. -I. Guo, “Residual Knowledge Retention for Edge Devices,” 2021 IEEE 30th International Symposium on Industrial Electronics (ISIE), Kyoto, Japan, 2021, pp. 1-6, doi: 10.1109/ISIE45552.2021.9576… [cited by examiner]
Dai, J., He, K., & Sun, J. (2015). Convolutional feature masking for joint object and stuff segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3992-4000). (Year: 2015). [cited by examiner]
Zhang, K., Li, T., Liu, B., & Liu, Q. (2019). Co-saliency detection via mask-guided fully convolutional networks with multi-scale label smoothing. In Proceedings of the IEEE/CVF conference on computer vision and pattern… [cited by examiner]
Anonymous. Stack Overflow. Asked Oct. 13, 2016. Accessed Aug. 7, 2025. Available at <https://stackoverflow.com/questions/40012749/can-cnn-learn-to-weigh-certain-feature-channels-much-much-more-than-others> (Year: 2016). [cited by examiner]
International Search Report for PCT Application No. PCT/JP2020/022405, mailed on Sep. 8, 2020. [cited by applicant]
Written opinion for PCT Application No. PCT/JP2020/022405, mailed on Sep. 8, 2020. [cited by applicant]
Figurnov et al., “Spatially Adaptive Computation Time for Residual Networks”, CVPR2017, 2017. [cited by applicant]
Yu et al., “Combining Background Subtraction and Convolutional Neural Network for Anomaly Detection in Pumping-Unit Surveillance”, Algorithms 2019, May 29, 2019. [cited by applicant]
Wang et al., “SkipNet: Learning Dynamic Routing in Convolutional Networks”, ECCV 2018, 2018. [cited by applicant]
Wu et al., “BlockDrop: Dynamic Inference Paths in Residual Networks”, CVPR 2018, 2018. [cited by applicant]
Lin et al., “Focal Loss for Dense Object Detection”, ICCV 2017, 2017. [cited by applicant]
Liu et al., “SSD: Single Shot MultiBox Detector”, ECCV 2016, Dec. 29, 2016. [cited by applicant]