IP Library › Granted Patent US 11,681,910
Granted Patent B2
US 11,681,910 · App. 16/621,796 · Granted Jun 20, 2023

Training apparatus, recognition apparatus, training method, recognition method, and program

Inventors: Tsutomu Horikawa (Kanagawa, JP); Daichi Ono (Kanagawa, JP)
Assignee: Sony Interactive Entertainment Inc.
G06N3/08G06F18/2148G06N20/00G06T7/0002G06T7/50G06T7/73G06V10/7747G06V20/64G06T2200/04G06T2207/10004
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,681,910
App. No.
16/621,796
Granted
Jun 20, 2023
Kind
B2
Abstract

Provided are a training apparatus, a recognition apparatus, a training method, a recognition method, and a program that can accurately recognize what an object represented in an image associated with depth information is. An object data acquiring section acquires three-dimensional data representing an object. A training data generating section generates a plurality of training data each representing a mutually different part of the object on the basis of the three-dimensional data. A training section trains a machine learning model using the generated training data as the training data for the object.

Claims (39)

1. A training apparatus for training a machine learning model used to identify an object represented in an image, the training apparatus comprising:

a three-dimensional data acquiring circuit operating to acquire three-dimensional data representing a complete virtual three-dimensional shape of an object;

a training data generating circuit operating to generate a plurality of training data sets from the complete virtual three-dimensional shape of the object, each training data set including a two-dimensional data representation for only a portion of the object viewed from a respective viewpoint in a virtual space, including all features of the appearance of the object from the respective viewpoint, and associated depth information, such that each training data set represents a mutually different part of the object; and

a training circuit operating to train the machine learning model to identify the object using the plurality of training data sets.

2. The training apparatus according to claim 1 , wherein

the training circuit operates to train the machine learning model by inputting recognition target data comprising the plurality of training data sets, including the two-dimensional data, as three-dimensional data generated on a basis of image data associated with the depth information, and

the training data generating circuit operates to generate the plurality of training data sets, each including the three-dimensional data.

3. The training apparatus according to claim 1 , wherein

the training operates to train the machine learning model by inputting recognition target data comprising the plurality of training data sets, including the two-dimensional, as an image associated with the depth information, and

the training data generating circuit operates to generate the plurality of training data sets, each including the image associated with the depth information.

4. A recognition apparatus for executing a process of identifying an object represented in an image, the recognition apparatus comprising:

a machine learning model; and

a recognition circuit operating to identify the object represented in the image on a basis of an output from the machine learning model in response to recognition target data corresponding to the image,

wherein the machine learning model has been trained via a process comprising:

acquiring three-dimensional data representing a complete virtual three-dimensional shape of an object;

generating a plurality of training data sets from the complete virtual three-dimensional shape of the object, each training data set including a two-dimensional data representation for only a portion of the object viewed from a respective viewpoint in a virtual space, including all features of the appearance of the object from the respective viewpoint, and associated depth information, such that each training data set represents a mutually different part of the object; and

training the machine learning model to identify the object using the plurality of training data sets.

5. A training method for training a machine learning model used for a process of identifying an object represented in an image, the training method comprising:

acquiring three-dimensional data representing a complete virtual three-dimensional shape of an object;

generating a plurality of training data sets from the complete virtual three-dimensional shape of the object, each training data set including a two-dimensional data representation for only a portion of the object viewed from a respective viewpoint in a virtual space, including all features of the appearance of the object from the respective viewpoint, and associated depth information, such that each training data set represents a mutually different part of the object; and

training the machine learning model to identify the object using the plurality of training data sets.

6. A recognition method for executing a process of identifying an object represented in an image, the recognition method comprising:

inputting recognition target data corresponding to the image into a machine learning model; and

identifying the object represented in the image on a basis of an output from the machine learning model in response to the recognition target data,

wherein the machine learning model has been trained via a process comprising:

acquiring three-dimensional data representing a complete virtual three-dimensional shape of an object;

generating a plurality of training data sets from the complete virtual three-dimensional shape of the object, each training data set including a two-dimensional data representation for only a portion of the object viewed from a respective viewpoint in a virtual space, including all features of the appearance of the object from the respective viewpoint, and associated depth information, such that each training data set represents a mutually different part of the object; and

training the machine learning model to identify the object using the plurality of training data sets.

7. A non-transitory, computer-readable storage medium containing a program, which when executed by a computer, causes the computer to carry out a process for training a machine learning model used for a process of identifying an object represented in an image, by carrying out actions, comprising:

acquiring three-dimensional data representing a complete virtual three-dimensional shape of an object;

generating a plurality of training data sets from the complete virtual three-dimensional shape of the object, each training data set including a two-dimensional data representation for only a portion of the object viewed from a respective viewpoint in a virtual space, including all features of the appearance of the object from the respective viewpoint, and associated depth information, such that each training data set represents a mutually different part of the object; and

training the machine learning model to identify the object using the plurality of training data sets.

8. A non-transitory, computer-readable storage medium containing a program, which when executed by a computer, causes the computer to carry out a process of identifying an object represented in an image, by carrying out actions, comprising:

inputting recognition target data corresponding to the image into a machine learning model; and

identifying the object represented in the image on a basis of an output from the machine learning model in response to the recognition target data,

wherein the machine learning model has been trained via a process comprising:

acquiring three-dimensional data representing a complete virtual three-dimensional shape of an object;

generating a plurality of training data sets from the complete virtual three-dimensional shape of the object, each training data set including a two-dimensional data representation for only a portion of the object viewed from a respective viewpoint in a virtual space, including all features of the appearance of the object from the respective viewpoint, and associated depth information, such that each training data set represents a mutually different part of the object; and

training the machine learning model to identify the object using the plurality of training data sets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2019
From: HORIKAWA, TSUTOMU; ONO, DAICHI
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 051261/0692 →
Continuity (1)
Related Publication 20200193632A1 · Jun 18, 2020
Cited By (1)
US 12,587,623