IP Library Granted Patent US 12705789
Granted Patent B2
US 12705789 · App. 18/615,824 · Granted Aug 11, 2026

Information processing apparatus, orientation estimation method, and storage medium

Inventors: Hiroshi Tojo (Tokyo, JP); Kosuke Saito (Tokyo, JP)
Assignee: Canon Kabushiki Kaisha
G06T7/73G06T7/194G06T7/50G06V10/44G06T2207/20044G06T2207/20084G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705789
App. No.
18/615,824
Granted
Aug 11, 2026
Kind
B2
Abstract

An information processing apparatus includes at least one processor, and at least one memory storing executable instructions which, when executed by the at least one processor, cause the at least one processor to perform operations including acquiring an image, detecting an entire human body from the acquired image, estimating a skeleton of the detected entire human body and generating skeleton information about the skeleton of the entire human body, extracting a first feature quantity based on the generated skeleton information, extracting a second feature quantity based on a clipped image including the detected entire human body, and estimating an orientation of the detected entire human body based on a third feature quantity in which the first and second feature quantities are connected.

Claims (42)

1 . An information processing apparatus comprising:

at least one processor; and

at least one memory storing executable instructions which, when executed by the at least one processor, cause the at least one processor to perform operations including:

acquiring an image;

detecting an entire human body from the acquired image;

estimating a skeleton of the detected entire human body and generating skeleton information about the skeleton of the entire human body;

extracting a first feature quantity based on the generated skeleton information;

extracting a second feature quantity based on a clipped image including the detected entire human body; and

estimating an orientation of the detected entire human body based on a third feature quantity in which the first and second feature quantities are connected,

wherein the operations further include estimating, based on the third feature quantity, relation information about a relation between the entire human body and a background in which the entire human body is excluded from the clipped image, and

wherein the relation information includes information for dividing the clipped image into the entire human body and a shadow of the human body.

2 . The information processing apparatus according to claim 1 , wherein the operations further include estimating whether the orientation of the detected entire human body is an orientation in which at least three points of the human body are in contact with a ground.

3 . The information processing apparatus according to claim 1 , wherein the relation information includes information for dividing the clipped image into the entire human body and the background.

4 . The information processing apparatus according to claim 1 , wherein the relation information includes information for dividing the clipped image into the entire human body and a floor surface.

5 . The information processing apparatus according to claim 1 , wherein the operations further include dividing the clipped image into an upper part of the human body, a lower part of the human body, a floor surface, and a wall surface.

6 . The information processing apparatus according to claim 1 , wherein the relation information includes an imaging angle of the clipped image.

7 . The information processing apparatus according to claim 1 , wherein the relation information includes depth information for the clipped image.

8 . The information processing apparatus according to claim 1 , wherein the extraction of the first feature quantity, the extraction of the second feature quantity, and the estimation of the human body are performed by using a trained neural network.

9 . The information processing apparatus according to claim 1 , wherein the estimation of the relation information is performed by using a trained neural network.

10 . A method for estimating an orientation, the method comprising:

acquiring an image;

detecting an entire human body from the acquired image;

estimating a skeleton of the detected entire human body and generating skeleton information about the skeleton of the entire human body;

extracting a first feature quantity based on the generated skeleton information;

extracting a second feature quantity based on a clipped image including the detected entire human body; and

estimating an orientation of the detected entire human body based on a third feature quantity in which the first and second feature quantities are connected,

wherein the method further comprises estimating, based on the third feature quantity, relation information about a relation between the entire human body and a background in which the entire human body is excluded from the clipped image, and

wherein the relation information includes information for dividing the clipped image into the entire human body and a shadow of the human body.

11 . The method according to claim 10 , wherein the method further comprises estimating whether the orientation of the detected entire human body is an orientation in which at least three points of the human body are in contact with a ground.

12 . The method according to claim 10 , wherein the relation information includes information for dividing the clipped image into the entire human body and the background.

13 . The method according to claim 10 , wherein the relation information includes information for dividing the clipped image into the entire human body and a floor surface.

14 . The method according to claim 10 , wherein the relation information includes an imaging angle of the clipped image.

15 . The method according to claim 10 , wherein the relation information includes depth information for the clipped image.

16 . A non-transitory computer-readable storage medium storing a program for causing a computer to execute a method comprising:

acquiring an image;

detecting an entire human body from the acquired image;

estimating a skeleton of the detected entire human body and generating skeleton information about the skeleton of the entire human body;

extracting a first feature quantity based on the generated skeleton information;

extracting a second feature quantity based on a clipped image including the detected entire human body; and

estimating an orientation of the detected entire human body based on a third feature quantity in which the first and second feature quantities are connected,

wherein the method further comprises estimating, based on the third feature quantity, relation information about a relation between the entire human body and a background in which the entire human body is excluded from the clipped image, and

wherein the relation information includes information for dividing the clipped image into the entire human body and a shadow of the human body.