IP Library › Granted Patent US 12,602,908
Granted Patent B2
US 12,602,908 · App. 18/016,602 · Granted Apr 14, 2026

Information processing device, information processing method, and a non-transitory computer readable recording medium including selecting a base image including a target region with an object subject to machine learning and for combination with another target region in another image to include in a dataset for training a machine learning model

Inventor: Yoshikazu Watanabe (Tokyo, JP)
Assignee: NEC CORPORATION
G06V10/774G06V10/22G06V10/267G06V10/7715G06V10/776
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,908
App. No.
18/016,602
Granted
Apr 14, 2026
Kind
B2
Abstract

An information processing device according to the present invention performs operations including: selecting a base image from a base dataset including a target region including an object to be subjected to machine learning and a background region not including an object to be subjected to machine learning, and generating a processing target image; selecting the target region included in another image included in the base dataset; combining an image of the selected target region and information on an object to be subjected to machine learning included in the image of the target region with the processing target image; generating a dataset of the processing target images obtained by combining a predetermined number of the target regions; calculating a feature of an image included in the dataset; generating a learned model using first machine learning using the feature and the dataset; and outputting the generated learned model generated.

Claims (56)

1 . An information processing device comprising:

a memory; and

at least one processor coupled to the memory,

the processor being configured to perform operations, the operations comprising:

selecting a base image from a base dataset that is a set of images including a target region including an object to be subjected to machine learning and a background region not including an object to be subjected to machine learning, and generating a processing target image that is a copy of the selected base image;

selecting the target region included in another image included in the base dataset;

combining an image of the selected target region and information on an object to be subjected to machine learning included in the image of the target region with the processing target image;

generating a dataset that is a set of the processing target images obtained by combining a predetermined number of the target regions;

calculating a feature of an image included in the dataset;

generating a learned model using first machine learning that is machine learning using the feature and the dataset;

outputting the generated learned model;

executing third machine learning that is machine learning using the dataset but does not reuse the feature;

determining the feature to be reused using a result of profiling of an execution status of the third machine learning;

executing the first machine learning using the determined feature and the dataset; and

determining which output of a layer in a neural network is used as the feature to be reused using a result of profiling.

2 . The information processing device according to claim 1 , wherein the operations further comprise:

changing a shape of the selected target region.

3 . The information processing device according to claim 2 , wherein the operations further comprise:

executing at least one of a change of a width, a change of a height, a change of a size, a change of an aspect ratio, a rotation of an image, and a change of an inclination of an image as a change of the shape of the target region.

4 . The information processing device according to claim 1 , wherein the operations further comprise:

cutting out a foreground, which is a region of an object to be subjected to machine learning, from the target region, and combining the cut-out foreground with the processing target image.

5 . The information processing device according to claim 1 , wherein the operations further comprise:

executing the first machine learning using the base dataset or second machine learning different from the first machine learning, and evaluating a result of the first machine learning using the base dataset or the second machine learning using the base dataset; and

based on the evaluation, generating the dataset related to the evaluation.

6 . The information processing device according to claim 5 , wherein the operations further comprise:

evaluating accuracy of recognition of an object to be subjected to machine learning as a result of evaluation of the first machine learning using the base dataset or the second machine learning using the base dataset; and

generating the dataset in such a way that the dataset includes many objects with low recognition accuracy.

7 . The information processing device according to claim 1 , wherein the operations further comprise:

redetermining the feature to be reused using a result of profiling of the first machine learning using the feature determined to be reused and the dataset, and executing the first machine learning using the redetermined feature and the dataset.

8 . An information processing method comprising:

selecting a base image from a base dataset that is a set of images including a target region including an object to be subjected to machine learning and a background region not including an object to be subjected to machine learning, and generating a processing target image that is a copy of the selected base image;

selecting the target region included in another image included in the base dataset;

combining an image of the selected target region and information on an object to be subjected to machine learning included in the image of the target region with the processing target image;

generating a dataset that is a set of the processing target images obtained by combining a predetermined number of the target regions;

calculating a feature of an image included in the dataset;

generating a learned model using first machine learning that is machine learning using the feature and the dataset;

outputting the generated learned model;

executing third machine learning that is machine learning using the dataset but does not reuse the feature;

determining the feature to be reused using a result of profiling of an execution status of the third machine learning;

executing the first machine learning using the determined feature and the dataset; and

determining which output of a layer in a neural network is used as the feature to be reused using a result of profiling.

9 . A non-transitory computer-readable recording medium embodying a program, the program causing a computer to perform a method, the method comprising:

selecting a base image from a base dataset that is a set of images including a target region including an object to be subjected to machine learning and a background region not including an object to be subjected to machine learning, and generating a processing target image that is a copy of the selected base image;

selecting the target region included in another image included in the base dataset;

combining an image of the selected target region and information on an object to be subjected to machine learning included in the image of the target region with the processing target image;

generating a dataset that is a set of the processing target images obtained by combining a predetermined number of the target regions;

calculating a feature of an image included in the dataset;

generating a learned model using first machine learning that is machine learning using the feature and the dataset;

outputting the generated learned model;

executing third machine learning that is machine learning using the dataset but does not reuse the feature;

determining the feature to be reused using a result of profiling of an execution status of the third machine learning;

executing the first machine learning using the determined feature and the dataset; and

determining which output of a layer in a neural network is used as the feature to be reused using a result of profiling.

10 . The information processing device according to claim 1 , wherein the operations further comprise:

calculating, for each layer, a decrease in a load of entire machine processing when the output of the layer is used; and

determining the output of the layer so that the load of the entire processing decreases.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2023
From: WATANABE, YOSHIKAZU
To: NEC CORPORATION
Reel/Frame 062397/0665 →
Continuity (1)
Related Publication 20230281967A1 · Sep 7, 2023
References Cited (13)
US 20190197356A1 · Kurita · 2019 [cited by examiner]
US 20200117991A1 · Suzuki · 2020 [cited by examiner]
US 20210264314A1 · Oura · 2021 [cited by examiner]
JP 2015148895A · 2015 [cited by applicant]
JP 2019114116A · 2019 [cited by applicant]
JP 2019185751A · 2019 [cited by applicant]
JP 2020027424A · 2020 [cited by applicant]
WO 2019021855A1 · 2019 [cited by applicant]
International Search Report for PCT Application No. PCT/JP2020/028622, mailed on Oct. 20, 2020. [cited by applicant]
English translation of Written opinion for PCT Application No. PCT/JP2020/028622, mailed on Oct. 20, 2020. [cited by applicant]
Shaoqing Ren, Kaiming He, Ross Girshick, Jian Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, [online], Jan. 6, 2016, Cornel University, [Searched on Oct. 16, 2019], Internet, <URL… [cited by applicant]
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, Alexander C. Berg, “SSD: Single Shot MultiBox Detector”, [online], Dec. 29, 2016, Cornel University, [Searched on Oct. 16, 2019], … [cited by applicant]
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, Piotr Dollar, “Focal Loss for Dense Object Detection”, [online], Feb. 2, 2018, Cornel University, [Searched on Oct. 16, 2019], Internet, <URL:https//arxiv.org/abs/17… [cited by applicant]