IP Library › Granted Patent US 12,671,824
Granted Patent B2
US 12,671,824 · App. 18/277,723 · Granted Jun 30, 2026

Region importance estimation

Inventors: Hayato Itsumi (Tokyo, JP); Koichi Nihei (Tokyo, JP); Florian Beye (Tokyo, JP); Charvi Vitthal (Tokyo, JP); Takanori Iwai (Tokyo, JP); Yusuke Shinohara (Tokyo, JP)
Assignee: NEC CORPORATION
H04N19/167G06T7/10G06T7/50G06V10/25G06V10/82G06V20/70H04N19/119H04N19/124H04N19/127H04N19/17G06T2207/20081G06V20/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,671,824
App. No.
18/277,723
Filed
Aug 17, 2023
Granted
Jun 30, 2026
Kind
B2
Art Unit
2669
USPC
382/106
Abstract

To provide a technique which makes it possible to suitably estimate an important region and a non-important region in an image, an information processing apparatus includes: an obtaining means for obtaining input data which includes at least one of image data and point cloud data; an estimating means for estimating levels of importance with respect to a respective plurality of regions which are included in a frame indicated by the input data; a replacing means for generating replaced data by replacing at least one of the plurality of regions, which are included in the input data, with alternative data in accordance with the levels of importance; an evaluating means for deriving an evaluation value by referring to the replaced data; and a training means for training the estimating means with reference to the evaluation value.

Claims (56)

1 . An information processing apparatus comprising

at least one processor,

the at least one processor carrying out:

an obtaining process of obtaining input data which includes at least one of image data and point cloud data;

an estimating process of estimating, with use of an estimation model, levels of importance with respect to a respective plurality of regions which are included in a frame indicated by the input data;

a replacing process of generating replaced data by replacing at least one of the plurality of regions, which are included in the input data, with alternative data in accordance with the levels of importance, wherein the alternative data has a data size smaller than a data size of the input data for the at least one of the plurality of regions;

an evaluating process of deriving an evaluation value by referring to the replaced data; and

a training process of training the estimation model with reference to the evaluation value,

wherein in the evaluating process, the at least one processor derives the evaluation value by referring to

an output obtained from a given controller in a case where the input data is inputted into the given controller and

an output obtained from the given controller in a case where the replaced data is inputted into the given controller,

wherein the given controller is an autonomous operation controller of a movable body, and the output is a control command for the movable body.

2 . The information processing apparatus as set forth in claim 1 , wherein in the evaluating process, the at least one processor derives the evaluation value by further referring to the input data.

3 . The information processing apparatus as set forth in claim 1 , wherein:

in the evaluating process, the at least one processor derives the evaluation value as a difference between (i) the output obtained from the given controller in the case where the input data is inputted into the given controller and (ii) the output obtained from the given controller in the case where the replaced data is inputted into the given controller; and

in the training process, the at least one processor trains the estimation model so that the evaluation value becomes low.

4 . The information processing apparatus as set forth in claim 1 , wherein in the replacing process, the at least one processor replaces, with the alternative data, one or more of the plurality of regions which one or more have been selected in ascending order of the levels of importance and have a given proportion in the frame.

5 . The information processing apparatus as set forth in claim 4 , wherein:

in the replacing process, the at least one processor generates the replaced data for each of a plurality of given proportions which differ from each other; and

in the evaluating process, the at least one processor

derives a preliminary evaluation value with respect to each replaced data generated in the replacing, and

derives the evaluation value by averaging preliminary evaluation values.

6 . The information processing apparatus as set forth in claim 1 , wherein in the evaluating process, the at least one processor derives the evaluation value further with reference to a data size of the replaced data.

7 . The information processing apparatus as set forth in claim 6 , wherein in the training process, the at least one processor trains the estimation model so that the data size of the replaced data becomes small.

8 . The information processing apparatus as set forth in claim 1 , wherein the alternative data used in the replacing process is data which includes at least one of noise and image data that has a large quantization error than that in the input data.

9 . The information processing apparatus as set forth in claim 1 , wherein in the estimating process, the at least one processor estimates the levels of importance with respect to the respective plurality of regions which are included in the frame indicated by the input data, with use of reference data relating to the frame.

10 . The information processing apparatus as set forth in claim 9 , wherein the reference data includes a segmentation image which is obtained by applying a segmentation process to the frame.

11 . The information processing apparatus as set forth in claim 9 , wherein the reference data includes a depth map which corresponds to the frame.

12 . The information processing apparatus as set forth in claim 9 , wherein the reference data includes an object detecting result which is obtained by applying, to the frame, a process of detecting an object.

13 . The information processing apparatus as set forth in claim 1 , wherein in the estimating process, the at least one processor estimates the levels of importance with use of a self-attention module.

14 . The information processing apparatus as set forth in claim 1 , wherein:

the evaluation value includes a reward value derived from the output; and

in the training process, the at least one processor trains the estimation model so that the reward value becomes high.

15 . The information processing apparatus as set forth in claim 1 , wherein:

the evaluation value is a loss value derived from the output; and

in the training process, the at least one processor trains the estimation model so that the loss value becomes low.

16 . An information processing method comprising:

obtaining input data which includes at least one of image data and point cloud data;

estimating, with use of an estimation model, levels of importance with respect to a respective plurality of regions which are included in a frame indicated by the input data;

generating replaced data by replacing at least one of the plurality of regions, which are included in the input data, with alternative data in accordance with the levels of importance, wherein the alternative data has a data size smaller than a data size of the input data for the at least one of the plurality of regions;

deriving an evaluation value by referring to the replaced data; and

training the estimation model with reference to the evaluation value,

wherein the evaluation value is derived by referring to

an output obtained from a given controller in a case where the input data is inputted into the given controller and

an output obtained from the given controller in a case where the replaced data is inputted into the given controller,

wherein the given controller is an autonomous operation controller of a movable body, and the output is a control command for the movable body.

17 . A computer-readable non-transitory recording medium in which a program is recorded, the program being for causing a computer to carry out:

an obtaining process of obtaining input data which includes at least one of image data and point cloud data;

an estimating process of estimating, with use of an estimation model, levels of importance with respect to a respective plurality of regions which are included in a frame indicated by the input data;

a replacing process of generating replaced data by replacing at least one of the plurality of regions, which are included in the input data, with alternative data in accordance with the levels of importance, wherein the alternative data has a data size smaller than a data size of the input data for the at least one of the plurality of regions;

an evaluating process of deriving an evaluation value by referring to the replaced data; and

a training process of training the estimation model with reference to the evaluation value,

wherein in the evaluating process, the computer derives the evaluation value by referring to

an output obtained from a given controller in a case where the input data is inputted into the given controller and

an output obtained from the given controller in a case where the replaced data is inputted into the given controller,

wherein the given controller is an autonomous operation controller of a movable body, and the output is a control command for the movable body.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2024
From: ITSUMI, HAYATO; NIHEI, KOICHI; BEYE, FLORIAN; IWAI, TAKANORI; SHINOHARA, YUSUKE; VITTHAL, CHARVI
To: NEC CORPORATION
Reel/Frame 066734/0101 →
Priority Claims (1)
WO PCT/JP2021/006866 · Feb 24, 2021 · international
Continuity (2)
Related Publication 20240135560A1 · Apr 25, 2024
Related Publication 20240236317A9 · Jul 11, 2024
References Cited (54)
US 11048973B1 · Ramanathan et al. · 2021 [cited by applicant]
US 20110216977A1 · Yu et al. · 2011 [cited by applicant]
US 20130294699A1 · Yu et al. · 2013 [cited by applicant]
US 20160057432A1 · Shibayama et al. · 2016 [cited by applicant]
US 20160086342A1 · Yamaji · 2016 [cited by applicant]
US 20160212405A1 · Zhang et al. · 2016 [cited by applicant]
US 20190075302A1 · Huang et al. · 2019 [cited by applicant]
US 20190198540A1 · Shibata et al. · 2019 [cited by applicant]
US 20200174472A1 · Zhang et al. · 2020 [cited by applicant]
US 20200298847A1 · Tawari · 2020 [cited by examiner]
US 20200326526A1 · Yeh · 2020 [cited by applicant]
US 20200327350A1 · Anand · 2020 [cited by examiner]
US 20210174482A1 · Ji et al. · 2021 [cited by applicant]
US 20210368186A1 · Sugio et al. · 2021 [cited by applicant]
US 20220019821A1 · Brickwedde · 2022 [cited by examiner]
US 20220021887A1 · Banerjee · 2022 [cited by examiner]
US 20220107497A1 · Murata et al. · 2022 [cited by applicant]
US 20220277548A1 · Nakao · 2022 [cited by examiner]
US 20230119685A1 · Kim · 2023 [cited by examiner]
US 20230300333A1 · Kubota · 2023 [cited by examiner]
US 20240155142A1 · Sugio et al. · 2024 [cited by applicant]
US 20260039846A1 · Sugio et al. · 2026 [cited by applicant]
CN 102819023A · 2012 [cited by applicant]
CN 112231516A · 2021 [cited by applicant]
CN 112257647A · 2021 [cited by applicant]
JP 2003274393A · 2003 [cited by applicant]
JP 2012103850A · 2012 [cited by applicant]
JP 2015521820A · 2015 [cited by applicant]
JP 2016046707A · 2016 [cited by applicant]
JP 2018056838A · 2018 [cited by applicant]
JP 2019083491A · 2019 [cited by applicant]
JP 2020083309A · 2020 [cited by applicant]
WO 2010032297A1 · 2010 [cited by applicant]
WO 2018121690A1 · 2018 [cited by applicant]
WO 2019128971A1 · 2019 [cited by applicant]
WO 2020110580A1 · 2020 [cited by applicant]
WO 2020162495A1 · 2020 [cited by applicant]
Li, Mu, et al. “Learning Convolutional Networks for Content-Weighted Image Compression.” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 2018. (Year: 2018). [cited by examiner]
Ramachandran, Prajit, et al. “Stand-alone self-attention in vision models.” Advances in neural information processing systems 32 (2019). (Year: 2019). [cited by examiner]
Sallab, Ahmad EL, et al. “Deep reinforcement learning framework for autonomous driving.” arXiv preprint arXiv:1704.02532v1 (2017). (Year: 2017). [cited by examiner]
Ohn-Bar, Eshed, and Mohan M. Trivedi. “What makes an on-road object important?.” 2016 23rd International Conference on Pattern Recognition (ICPR). IEEE, 2016. (Year: 2016). [cited by examiner]
Rahimpour, Alireza, et al. “Context aware road-user importance estimation (iCARE).” 2019 IEEE Intelligent Vehicles Symposium (IV). IEEE, 2019. (Year: 2019). [cited by examiner]
Sandri, Gustavo, et al. “Point cloud compression incorporating region of interest coding.” 2019 IEEE International Conference on Image Processing (ICIP). IEEE, 2019. (Year: 2019). [cited by examiner]
International Search Report for PCT Application No. PCT/JP2022/005530, mailed on Apr. 5, 2022. [cited by applicant]
Yusuke Shinohara et al., “Video Compression Estimating Recognition Accuracy for Remote Site Object Detection”, 2020 International Wireless Communications and Mobile Computing (IWCMC), Jul. 27, 2020, pp. 285-290, <URL: h… [cited by applicant]
Tomonori Kubota et al., “A high-compression video coding method for video analysis using Deep Learning”, IEICE Technical Report, May 11, 2020 (received date), pp. 121-126. [cited by applicant]
Takanori Nakao et al., “Moving image coding method for AI analysis”, IPSJ Technical Report, Feb. 20, 2020 (received date), pp. 1-6, <URL: https://ipsj.ixsq.nii.ac.jp/ej/?action=repository_uri&item_id=203339&file_id=1&fi… [cited by applicant]
Leonardo Galteri et al., “Video Compression for Object Detection Algorithms”, 2018 24th International Conference on Pattern Recognition (ICPR), Nov. 29, 2018, pp. 3007-3012, <URL: https://ieeexplore.ieee.org/document/85… [cited by applicant]
International Search Report for PCT Application No. PCT/JP2021/006866, mailed on May 11, 2021. [cited by applicant]
US Office Action for U.S. Appl. No. 18/277,719, mailed on Mar. 4, 2026. [cited by applicant]
Wong et al., “Characterization of perceptual importance for object-based image segmentation,” Proceedings 2000 international Conference on Image Processing (Cat. No.00CH37101), Vancouver, BC, Canada, 2000, pp. 54-57, vo… [cited by applicant]
Shibata et al., “Unified Image Fusion Framework With Learning-Based Application-Adaptive Importance Measure,” in IEEE Transactions on Computational Imaging, vol. 5, No. 1, pp. 82-96, Mar. 2019, doi: 10.1109/TC !.2018.28… [cited by applicant]
Pan et al., “Label and Sample: Efficient Training of Vehicle Object Detector from Sparsely Labeled Data,” arXiv: 1808.08603, Aug. 26, 2018. (Year: 2018). [cited by applicant]
Jyothi et al., “Research study of neural networks for image categorization and retrieval,” 2010 The 2nd international Conference on Computer and Automation Engineering (ICCAE), Singapore, 2010, pp. 686-690, doi: 10.1109… [cited by applicant]