IP Library Granted Patent US 10,984,575
Granted Patent B2
US 10,984,575 · App. 16/269,312 · Granted Apr 20, 2021

Body pose estimation

Inventors: Avihay Assouline (Tel Aviv, IL); Itamar Berger (Hod Hasharon, IL); Yuncheng Li (Los Angeles, CA)
Assignee: Snap Inc.
G06T13/40G06K9/00375
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,984,575
App. No.
16/269,312
Filed
Feb 6, 2019
Granted
Apr 20, 2021
Kind
B2
Art Unit
2619
USPC
345/474
Abstract

Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and a method for detecting a pose of a user. The program and method include receiving a monocular image that includes a depiction of a body of a user; detecting a plurality of skeletal joints of the body depicted in the monocular image; and determining a pose represented by the body depicted in the monocular image based on the detected plurality of skeletal joints of the body. A pose of an avatar is modified to match the pose represented by the body depicted in the monocular image by adjusting a set of skeletal joints of a rig of an avatar based on the detected plurality of skeletal joints of the body; and the avatar having the modified pose that matches the pose represented by the body depicted in the monocular image is generated for display.

Claims (51)

1. A method comprising:

receiving, by one or more processors, a monocular image that includes a depiction of a body of a user;

detecting, by the one or more processors, a plurality of skeletal joints of the body depicted in the monocular image;

determining, by the one or more processors, a pose represented by the body depicted in the monocular image based on the detected plurality of skeletal joints of the body;

in response to determining the pose represented by the body of the user, modifying, by the one or more processors, poses of a plurality of avatars from being in different neutral poses to having identical poses that match the pose represented by the body of the user depicted in the monocular image by adjusting a set of skeletal joints of rigs of the avatars based on the detected plurality of skeletal joints of the body; and

generating, for display by the one or more processors, the plurality of avatars having the modified poses that match the pose represented by the body of the user depicted in the monocular image.

2. The method of claim 1 , wherein the monocular image is a first frame of a video, further comprising:

identifying a plurality of skeletal joint features of the monocular image using a first machine learning technique, wherein positions of the plurality of skeletal joints are detected based on the identified plurality of skeletal joint features; and

estimating a position of the user and a scale of an image of the user in a second frame of the video that is adjacent to the first frame by processing the first frame of the video using a second machine learning technique.

3. The method of claim 1 further comprising:

predicting a position of the user in a subsequent image that will be received after the monocular image; and

identifying a plurality of skeletal joint positions corresponding to the detected plurality of skeletal joints, wherein the pose represented by the body is determined based on a pose associated with the plurality of skeletal joint positions.

4. The method of claim 1 further comprising generating for display the avatars having the modified poses together with the depiction of the body of the user, wherein a first of the avatars is in a first neutral pose different from a second neutral pose of a second of the avatars, and wherein the first and second avatars transition to having identical poses that mimic the pose represented by the body of the user in response to determining the pose represented by the body of the user.

5. The method of claim 4 , wherein the user is a first user and wherein the avatars include a first avatar, further comprising:

capturing a first image that includes the display of the first avatar having the modified pose together with the depiction of the body of the first user; and

sending, from a first user device of the first user, the captured first image to a second user device of a second user.

6. The method of claim 5 further comprising receiving a second image from the second user device, the second image including a simultaneous display of a second avatar and a depiction of a body of the second user, wherein a pose of the second avatar in the second image matches a pose depicted by the body of the second user.

7. The method of claim 6 further comprising generating a simultaneous display of the first and second images.

8. The method of claim 1 further comprising:

receiving a video comprising a plurality of monocular images that include the depiction of the body of the user;

tracking changes in the plurality of skeletal joints across the plurality of monocular images;

detecting changes to the pose represented by the body based on tracking the changes in the plurality of skeletal joints; and

continuously or periodically modifying poses of the avatars to match the changes to the pose represented by the body.

9. The method of claim 8 further comprising:

presenting a virtual object in the video; and

adjusting one or more of a rate at which the virtual object moves, a position of the object in the video, and a visual attribute of the object based on the detected changes to the pose represented by the body.

10. The method of claim 1 , wherein the plurality of avatars are a plurality of identical avatars, and wherein the avatars comprise a first avatar, and the method further comprises:

generating, for display by the one or more processors; the plurality of identical avatars having a first set of different poses; and

detecting that the pose represented by the body depicted in the monocular image corresponds to a specified pose.

11. The method of claim 10 further comprising, in response to detecting that the pose represented by the body depicted in the monocular image corresponds to the specified pose:

modifying the first set of different poses of the plurality of avatars to have identical poses that match the pose represented by the body depicted in the monocular image; and

generating, for display by the one or more processors; the plurality of avatars having the identical poses.

12. The method of claim 11 further comprising animating the modification of the first set of different poses of the plurality of avatars in the generated display.

13. The method of claim 1 further comprising causing a given one of the avatars to interact with a virtual object depicted in an image based on the modified pose.

14. The method of claim 1 , wherein the detecting and the determining steps are performed without accessing depth information from a depth sensor.

15. The method of claim 1 , wherein detecting the plurality of skeletal joints of the body comprises identifying points respectively associated with a right wrist, a right elbow, a right shoulder; a nose on a face of the user, a left shoulder, a left elbow, and a left wrist.

16. The method of claim 1 , wherein a rate at which the plurality of skeletal joints is detected is adjusted based on a position of the user relative to an image capture device.

17. A system comprising:

a processor configured to perform operations comprising:

receiving a monocular image that includes a depiction of a body of a user;

detecting a plurality of skeletal joints of the body depicted in the monocular image;

determining a pose represented by the body depicted in the monocular image based on the detected plurality of skeletal joints of the body;

in response to determining the pose represented by the body of the user, modifying poses of a plurality of avatars from being in different neutral poses to having identical poses that match the pose represented by the body of the user depicted in the monocular image by adjusting a set of skeletal joints of rigs of the avatars based on the detected plurality of skeletal joints of the body; and

generating, for display, the plurality of avatars having the modified poses that match the pose represented by the body of the user depicted in the monocular image.

18. The system of claim 17 , wherein the operations further comprise identifying a plurality of skeletal joint features of the monocular image using a machine learning technique; wherein the plurality of skeletal joints are detected based on the identified plurality of skeletal joint features.

19. A non-transitory machine-readable storage medium that includes instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving a monocular image that includes a depiction of a body of a user;

detecting a plurality of skeletal joints of the body depicted in the monocular image;

determining a pose represented by the body depicted in the monocular image based on the detected plurality of skeletal joints of the body;

in response to determining the pose represented by the body of the user, modifying poses of a plurality of avatars from being in different neutral poses to having identical poses that match the pose represented by the body of the user depicted in the monocular image by adjusting a set of skeletal joints of rigs of the avatars based on the detected plurality of skeletal joints of the body; and

generating, for display, the plurality of avatars having the modified poses that match the pose represented by the body of the user depicted in the monocular image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2021
From: ASSOULINE, AVIHAY; BERGER, ITAMAR; LI, YUNCHENG
To: SNAP INC.
Reel/Frame 055718/0070 →
Continuity (1)
Related Publication 20200250874A1 · Aug 6, 2020
Cited By (73)
US 1,089,291 US 12,198,398 US 12,223,672 US 12,229,860 US 12,229,901 US 12,235,991 US 12,236,512 US 12,243,173 US 12,254,577 US 12,271,536 US 12,277,632 US 12,284,146 US 12,284,698 US 12,288,273 US 12,293,433 US 12,299,775 US 12,307,564 US 12,315,495 US 12,321,577 US 12,340,453 US 12,361,934 US 12,387,444 US 12,394,154 US 12,395,456 US 12,400,356 US 12,412,347 US 12,417,562 US 12,418,504 US 12,429,953 US 12,436,598 US 12,469,273 US 12,472,435 US 12,475,621 US 12,475,658 US 12,499,483 US 12,499,638 US 12,504,866 US 12,513,098 US 12,517,626 US 12,518,437 US 12,518,738 US 12,530,847 US 12,530,852 US 12,536,751 US 12,541,930 US 12,548,267 US 12,555,274 US 12,555,310 US 12,567,102 US 12,579,204 US 12,580,784 US 12,580,980 US 12,586,562 US 12,602,842 US 12,614,354 US 12,614,359 US 12,620,188 US 12,620,216 US 12,632,890 US 12,633,073 US 12,646,266 US 12,646,268 US 12,651,292 US 12,651,396 US 12,651,409 US 12,670,674 US 12,682,508 US 12,695,929 US 12,699,802 US 12,700,187 US 12,705,875 US 12,711,538 US 12,725,332