IP Library › Granted Patent US 11,170,246
Granted Patent B2
US 11,170,246 · App. 16/621,544 · Granted Nov 9, 2021

Recognition processing device, recognition processing method, and program

Inventors: Daichi Ono (Kanagawa, JP); Tsutomu Horikawa (Kanagawa, JP)
Assignee: Sony Interactive Entertainment Inc.
G06K9/3233G06K9/00624G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,170,246
App. No.
16/621,544
Granted
Nov 9, 2021
Kind
B2
Abstract

Provided are a recognition processing device, a recognition processing method, and a program capable of efficiently narrowing down a three-dimensional region on which recognition processing using a three-dimensional convolutional neural network is to be executed. A first recognition process executing section executes a first recognition process on a captured image obtained by capturing an image of a real space and used to generate voxel data. A target two-dimensional region determining section determines a two-dimensional region occupying part of the captured image on the basis of a result of the first recognition process. A target three-dimensional region determining section determines a three-dimensional region in the real space on the basis of the two-dimensional region and a position of a camera when the camera obtains the captured image. A second recognition process executing section executes a second recognition process using a three-dimensional convolutional neural network on the voxel data associated with a position in the three-dimensional region.

Claims (29)

1. A recognition processing device executing recognition processing using a three-dimensional convolutional neural network on voxel data in which a position in a real space and a voxel value are associated with each other, the recognition processing device comprising:

a first recognition process executing section executing a first recognition process on a captured image which is obtained by capturing an image of the real space and which is used to generate the voxel data;

a two-dimensional region determining section determining a two-dimensional region occupying part of the captured image on a basis of a result of the first recognition process;

a three-dimensional region determining section determining a three-dimensional region in the real space on a basis of the two-dimensional region and a position of a camera when the camera obtains the captured image; and

a second recognition process executing section executing a second recognition process using a three-dimensional convolutional neural network on the voxel data associated with a position in the three-dimensional region, wherein:

the first recognition process executing section executes the first recognition process on each of a first captured image and a second captured image obtained by capturing images of the real space from positions different from each other, and

the two-dimensional region determining section determines a first two-dimensional region occupying part of the first captured image and a second two-dimensional region occupying part of the second captured image on a basis of a result of the first recognition process.

2. The recognition processing device according to claim 1 , wherein:

the three-dimensional region determining section determines a first three-dimensional region determined on a basis of the first two-dimensional region and a position of a camera when the camera obtains the first captured image, and a second three-dimensional region determined on a basis of the second two-dimensional region and a position of a camera when the camera obtains the second captured image, and

the second recognition process executing section executes the second recognition process using the three-dimensional convolutional neural network on the voxel data associated with a position in the three-dimensional region in the real space according to the first three-dimensional region and the second three-dimensional region.

3. The recognition processing device according to claim 1 , wherein

the first recognition process executing section executes the first recognition process on a captured image associated with depth information, and

the three-dimensional region determining section determines the three-dimensional region in the real space on a basis of the depth information associated with a position in the two-dimensional region.

4. The recognition processing device according to claim 1 , wherein the first recognition process executing section executes the first recognition process using a two-dimensional convolutional neural network on the captured image.

5. The recognition processing device according to claim 2 , wherein the second recognition process executing section executes the second recognition process using the three-dimensional convolutional neural network on the voxel data associated with a position in the three-dimensional region where the first three-dimensional region and the second three-dimensional region intersect.

6. A recognition processing method for executing recognition processing using a three-dimensional convolutional neural network on voxel data in which a position in a real space and a voxel value are associated with each other, the recognition processing method comprising:

executing a first recognition process on a captured image which is obtained by capturing an image of the real space and which is used to generate the voxel data;

determining a two-dimensional region occupying part of the captured image on a basis of a result of the first recognition process;

determining a three-dimensional region in the real space on a basis of the two-dimensional region and a position of a camera when the camera obtains the captured image; and

executing a second recognition process using a three-dimensional convolutional neural network on the voxel data associated with a position in the three-dimensional region, wherein:

the executing includes executing the first recognition process on each of a first captured image and a second captured image obtained by capturing images of the real space from positions different from each other, and

the determining a two-dimensional region includes determining a first two-dimensional region occupying part of the first captured image and a second two-dimensional region occupying part of the second captured image on a basis of a result of the first recognition process.

7. A non-transitory, computer-readable storage medium containing a program, which when executed by a computer, causes the computer to execute recognition processing using a three-dimensional convolutional neural network on voxel data in which a position in a real space and a voxel value are associated with each other by carrying out actions, comprising:

executing a first recognition process on a captured image which is obtained by capturing an image of the real space and which is used to generate the voxel data;

determining a two-dimensional region occupying part of the captured image on a basis of a result of the first recognition process;

determining a three-dimensional region in the real space on a basis of the two-dimensional region and a position of a camera when the camera obtains the captured image; and

executing a second recognition process using a three-dimensional convolutional neural network on the voxel data associated with a position in the three-dimensional region, wherein:

the executing includes executing the first recognition process on each of a first captured image and a second captured image obtained by capturing images of the real space from positions different from each other, and

the determining a two-dimensional region includes determining a first two-dimensional region occupying part of the first captured image and a second two-dimensional region occupying part of the second captured image on a basis of a result of the first recognition process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2019
From: ONO, DAICHI; HORIKAWA, TSUTOMU
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 051250/0185 →
Continuity (1)
Related Publication 20210056337A1 · Feb 25, 2021
Cited By (1)
US 12,675,950