IP Library Patent Application 18569996
Patent Application
App. No. 18/569,996

ENHANCED TECHNIQUES FOR REAL-TIME MULTI-PERSON THREE-DIMENSIONAL POSE TRACKING USING A SINGLE CAMERA

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/569,996
Abstract

This disclosure describes systems, methods, and devices related to real-time multi-person three-dimensional pose tracking using a single camera. A method may include receiving, by a device, two-dimensional image data from a camera, the two-dimensional image data representing a first person and a second person; generating, based on the two-dimensional image data, two-dimensional positions of body parts represented by the first person; generating, using a deep neural network, based on the two-dimensional positions, a three-dimensional pose regression of the body parts represented by the first person; identifying, based on the two-dimensional positions and the three-dimensional pose regression, contact between a ground plane and a foot of the first person; generating an absolute three-dimensional position of the contact between the ground plane and the foot of the first person; generating, based on the absolute three-dimensional position, a three-dimensional pose of the body parts represented by the first person.

Claims (125)

1 . A method for real-time three-dimensional human pose tracking using two-dimensional image data, the method comprising:

receiving, by at least one processor of a device, two-dimensional image data from a camera, the two-dimensional image data representing a first person and a second person;

generating, by the at least one processor, based on the two-dimensional image data, first two-dimensional positions of body parts represented by the first person;

generating, by the at least one processor, based on the two-dimensional image data, second two-dimensional positions of body parts represented by the second person;

generating, by the at least one processor, using a deep neural network, based on the first two-dimensional positions, a first root-relative three-dimensional pose regression of the body parts represented by the first person;

generating, by the at least one processor, using the deep neural network, based on the second two-dimensional positions, a second root-relative three-dimensional pose regression of the body parts represented by the second person;

identifying, by the at least one processor, based on the first two-dimensional positions and the first root-relative three-dimensional pose regression, contact between a ground plane and a foot of the first person;

identifying, by the at least one processor, based on the second two-dimensional positions and the second root-relative three-dimensional pose regression, contact between the ground plane and a foot of the second person;

generating, by the at least one processor, a first absolute three-dimensional position of the contact between the ground plane and the foot of the first person;

generating, by the at least one processor, a second absolute three-dimensional position of the contact between the ground plane and the foot of the second person;

generating, by the at least one processor, based on the first absolute three-dimensional position, a first three-dimensional pose of the body parts represented by the first person; and

generating, by the at least one processor, based on the second absolute three-dimensional position, a second three-dimensional pose of the body parts represented by the second person.

2 . The method of claim 1 , further comprising:

identifying two-dimensional positions of the ground plane based on the two-dimensional image data; and

generating a homographic matrix associated with mapping the two-dimensional positions of the ground plane to three-dimensional positions of the ground plane,

wherein generating the first absolute three-dimensional position is further based on the homographic matrix, and

wherein generating the second absolute three-dimensional position is further based on the homographic matrix.

3 . The method of claim 2 , further comprising:

generating extrinsic parameters of the camera based on the homographic matrix and a focal length of the camera,

wherein generating the first three-dimensional pose is further based on the extrinsic parameters,

wherein generating the second three-dimensional pose is further based on the extrinsic parameters, and

wherein the extrinsic parameters are indicative of a rotation matrix and a translation vector for the camera.

4 . The method of claim 1 , further comprising:

identifying a first two-dimensional position of the contact between the ground plane and the foot of the first person;

identifying a second two-dimensional position of the contact between the ground plane and the foot of the second person;

determining, based on the first two-dimensional position, that a first velocity of the foot of the first person is below a threshold velocity; and

determining, based on the second two-dimensional position, that a second velocity of the foot of the second person is below the threshold velocity,

wherein identifying the contact between the ground plane and the foot of the first person is further based on the first velocity of the foot of the first person being below the threshold velocity, and

wherein identifying the contact between the ground plane and the foot of the second person is further based on the second velocity of the foot of the second person being below the threshold velocity.

5 . The method of claim 1 , further comprising:

identifying a two-dimensional position of the contact between the ground plane and the foot of the first person; and

identifying two-dimensional positions of the ground plane based on the two-dimensional image data; and

generating a homographic matrix associated with mapping the two-dimensional positions of the ground plane to three-dimensional positions of the ground plane,

wherein generating the first three-dimensional pose is further based on mapping the two-dimensional position of the contact between the ground plane and the foot of the first person to the first absolute three-dimensional position using the homographic matrix.

6 . The method of claim 5 , further comprising:

determining a difference between the first three-dimensional pose and the first root-relative three-dimensional pose regression; and

generating, based on the difference, a third three-dimensional pose of the body parts represented by the first person.

7 . The method of claim 5 , further comprising:

determining a difference between a first image comprising the two-dimensional image data and a second image comprising second two-dimensional image data representing the first person and the second person; and

generating, based on the difference, a third three-dimensional pose of the body parts represented by the first person.

8 . The method of claim 1 , further comprising:

identifying a two-dimensional position of the contact between the ground plane and the foot of the first person;

identifying two-dimensional positions of the ground plane based on the two-dimensional image data;

generating a homographic matrix associated with mapping the two-dimensional positions of the ground plane to three-dimensional positions of the ground plane;

generating a three-dimensional root position of the foot of the first person based on the homographic matrix; and

determining a difference between the first absolute three-dimensional position and the three-dimensional root position of the foot of the first person,

wherein generating the first absolute three-dimensional position is based on the difference.

9 . The method of claim 1 , wherein a first image and a second image comprise the two-dimensional image data, the method comprising:

identifying two-dimensional positions of the ground plane based on the two-dimensional image data;

generating a homographic matrix associated with mapping the two-dimensional positions of the ground plane to three-dimensional positions of the ground plane;

generating extrinsic parameters of the camera based on the homographic matrix and a focal length of the camera; and

determining, based on the extrinsic parameters, a difference between the first absolute three-dimensional position and the first two-dimensional positions,

wherein generating the first absolute three-dimensional position is based on the difference, and

wherein the extrinsic parameters are indicative of a rotation matrix and a translation vector for the camera.

10 . A system for real-time three-dimensional human pose tracking using two-dimensional image data, the system comprising at least one processor coupled to memory, the at least one processor configured to:

receive two-dimensional image data from a camera, the two-dimensional image data representing a first person and a second person;

generate, based on the two-dimensional image data, first two-dimensional positions of body parts represented by the first person;

generate, based on the two-dimensional image data, second two-dimensional positions of body parts represented by the second person;

generate, using a deep neural network, based on the first two-dimensional positions, a first root-relative three-dimensional pose regression of the body parts represented by the first person;

generate, using the deep neural network, based on the second two-dimensional positions, a second root-relative three-dimensional pose regression of the body parts represented by the second person;

identify, based on the first two-dimensional positions and the first root-relative three-dimensional pose regression, contact between a ground plane and a foot of the first person;

identify, based on the second two-dimensional positions and the second root-relative three-dimensional pose regression, contact between the ground plane and a foot of the second person;

generate a first absolute three-dimensional position of the contact between the ground plane and the foot of the first person;

generate a second absolute three-dimensional position of the contact between the ground plane and the foot of the second person;

generate, based on the first absolute three-dimensional position, a first three-dimensional pose of the body parts represented by the first person; and

generate, based on the second absolute three-dimensional position, a second three-dimensional pose of the body parts represented by the second person.

11 . The system of claim 10 , wherein the at least one processor is further configured to:

identify two-dimensional positions of the ground plane based on the two-dimensional image data; and

generate a homographic matrix associated with mapping the two-dimensional positions of the ground plane to three-dimensional positions of the ground plane,

wherein to generate the first absolute three-dimensional position is further based on the homographic matrix, and

wherein to generate the second absolute three-dimensional position is further based on the homographic matrix.

12 . The system of claim 10 , wherein the at least one processor is further configured to:

generate a homographic matrix associated with mapping the first two-dimensional positions to three-dimensional positions,

wherein to generate the first absolute three-dimensional position is further based on the homographic matrix, and

wherein to generate the second absolute three-dimensional position is further based on the homographic matrix.

13 . The system of claim 10 , wherein the at least one processor is further configured to:

identify a first two-dimensional position of the contact between the ground plane and the foot of the first person;

identify a second two-dimensional position of the contact between the ground plane and the foot of the second person;

determine, based on the first two-dimensional position, that a first velocity of the foot of the first person is below a threshold velocity; and

determine, based on the second two-dimensional position, that a second velocity of the foot of the second person is below the threshold velocity,

wherein to identify the contact between the ground plane and the foot of the first person is further based on the first velocity of the foot of the first person being below the threshold velocity, and

wherein to identify the contact between the ground plane and the foot of the second person is further based on the second velocity of the foot of the second person being below the threshold velocity.

14 . The system of claim 10 , wherein the at least one processor is further configured to:

identify a two-dimensional position of the contact between the ground plane and the foot of the first person; and

identify two-dimensional positions of the ground plane based on the two-dimensional image data; and

generate a homographic matrix associated with mapping the two-dimensional positions of the ground plane to three-dimensional positions of the ground plane,

wherein to generate the first three-dimensional pose is further based on mapping the two-dimensional position of the contact between the ground plane and the foot of the first person to the first absolute three-dimensional position using the homographic matrix.

15 . The system of claim 10 , wherein the at least one processor is further configured to:

determine a difference between the first three-dimensional pose and the first root-relative three-dimensional pose regression; and

generate, based on the difference, a third three-dimensional pose of the body parts represented by the first person.

16 . The system of claim 10 , wherein the at least one processor is further configured to:

determine a difference between a first image comprising the two-dimensional image data and a second image comprising second two-dimensional image data representing the first person and the second person; and

generate, based on the difference, a third three-dimensional pose of the body parts represented by the first person.

17 . The system of claim 10 , wherein the at least one processor is further configured to:

identify a two-dimensional position of the contact between the ground plane and the foot of the first person;

identify two-dimensional positions of the ground plane based on the two-dimensional image data;

generate a homographic matrix associated with mapping the two-dimensional positions of the ground plane to three-dimensional positions of the ground plane;

generate a three-dimensional root position of the foot of the first person based on the homographic matrix; and

determine a difference between the first absolute three-dimensional position and the three-dimensional root position of the foot of the first person,

wherein to generate the first absolute three-dimensional position is based on the difference.

18 . The system of claim 10 , wherein a first image and a second image comprise the two-dimensional image data, and wherein the at least one processor is further configured to:

identify two-dimensional positions of the ground plane based on the two-dimensional image data;

generate a homographic matrix associated with mapping the two-dimensional positions of the ground plane to three-dimensional positions of the ground plane;

generate extrinsic parameters of the camera based on the homographic matrix and a focal length of the camera; and

determine, based on the extrinsic parameters, a difference between the first absolute three-dimensional position and the first two-dimensional positions,

wherein to generate the first absolute three-dimensional position is based on the difference, and

wherein the extrinsic parameters are indicative of a rotation matrix and a translation vector for the camera.

19 . An apparatus for real-time three-dimensional human pose tracking using two-dimensional image data, the apparatus comprising:

means for receiving two-dimensional image data from a camera, the two-dimensional image data representing a first person and a second person;

means for generating, based on the two-dimensional image data, first two-dimensional positions of body parts represented by the first person;

means for generating, based on the two-dimensional image data, second two-dimensional positions of body parts represented by the second person;

means for generating, using a deep neural network, based on the first two-dimensional positions, a first root-relative three-dimensional pose regression of the body parts represented by the first person;

means for generating, using the deep neural network, based on the second two-dimensional positions, a second root-relative three-dimensional pose regression of the body parts represented by the second person;

means for identifying, based on the first two-dimensional positions and the first root-relative three-dimensional pose regression, contact between a ground plane and a foot of the first person;

means for identifying, based on the second two-dimensional positions and the second root-relative three-dimensional pose regression, contact between the ground plane and a foot of the second person;

means for generating a first absolute three-dimensional position of the contact between the ground plane and the foot of the first person;

means for generating a second absolute three-dimensional position of the contact between the ground plane and the foot of the second person;

means for generating, based on the first absolute three-dimensional position, a first three-dimensional pose of the body parts represented by the first person; and

means for generating, based on the second absolute three-dimensional position, a second three-dimensional pose of the body parts represented by the second person.

20 . The apparatus of claim 19 , further comprising:

means for identifying two-dimensional positions of the ground plane based on the two-dimensional image data; and

means for generating a homographic matrix associated with mapping the two-dimensional positions of the ground plane to three-dimensional positions of the ground plane,

wherein generating the first absolute three-dimensional position is further based on the homographic matrix, and

wherein generating the second absolute three-dimensional position is further based on the homographic matrix.

21 - 25 . (canceled)

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2026
From: INTEL CORPORATION
To: INTEL PRODUCTS IP LLC
Reel/Frame 075991/0662 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2026
From: WANG, SHANDONG; CHEN, YURONG; LU, MING; XU, LI; YAO, ANBANG
To: INTEL CORPORATION
Reel/Frame 074871/0186 →