IP Library Granted Patent US 11,343,532
Granted Patent B2
US 11,343,532 · App. 17/162,580 · Granted May 24, 2022

System and method for vision-based joint action and pose motion forecasting

Inventors: Yanjun Zhu (Buffalo, NY); Yanxia Zhang (Cupertino, CA); Qiong Liu (Cupertino, CA); Andreas Girgensohn (Palo Alto, CA); Daniel Avrahami (Mountain View, CA); Francine Chen (Menlo Park, CA); Hao Hu (Sunnyvale, CA)
Assignee: FUJIFILM Business Innovation Corp.
H04N19/521G06T9/002G06V20/41G06V20/46G06V40/23H04N19/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,343,532
App. No.
17/162,580
Granted
May 24, 2022
Kind
B2
Abstract

A computer-implemented method, comprising, extracting a video sequence, the video sequence including a series of video frames, estimating current poses of one or more subjects within each video frame and determining joint locations for a joint associated with the one or more subjects within each video frame, computing optical flows from the video sequence, extracting motion features from the video sequence based on the optical flows, and predicting at least one future pose.

Claims (47)

1. A computer-implemented method, comprising:

extracting a video sequence, the video sequence including a series of video frames;

estimating current poses of one or more subjects within each video frame and determining joint locations for a joint associated with the one or more subjects within each video frame;

computing optical flows from the video sequence;

extracting motion features from the video sequence based on the optical flows; and

predicting at least one future pose.

2. The method of claim 1 further comprising:

determining a current action label for each motion feature contained in state information encoded by an encoder based on the current poses and the motion features, for a first frame of the series of video frames, the series of video frames being in red green blue (RGB) format;

predicting, by a decoder, future action labels for each motion feature in a second frame of the series of video frames subsequent to the first frame, based on the current pose, action label and the state information, wherein the predicting the at least one future pose comprises predicting, by a decoder, future poses for each motion feature in the second frame based on the current poses and the state information; and

refining the current action label, the future action labels, and the future poses based on a loss function.

3. The method of claim 1 , wherein the predicting the at least one future pose comprises predicting the at least one future pose and future motion for the second frame based on the model and the video sequence.

4. The method of claim 1 , wherein the predicting the at least one future pose comprises predicting the at least one future pose based on the current poses and the motion features.

5. The method of claim 2 , wherein the encoder and the decoder are implemented using a recurrent neural network, implementing one or more gated recurrent network.

6. The method of claim 1 wherein the computing the optical flows forms two-channel flow frames where each channel contains displacements at x and y axis, respectively.

7. The method of claim 1 wherein a frame rate is of the video sequence is between 24 fps and 60 fps.

8. The method of claim 1 wherein each joint location of the one or more joint locations comprises a pair of two dimensional coordinates (e.g., X,Y) or 3D coordinates (e.g., X,Y,Z).

9. A non-transitory computer readable medium including instructions executable on a processor, the instructions comprising:

extracting a video sequence, the video sequence including a series of video frames;

estimating current poses of one or more subject within each video frame and determining joint locations for a joint associated with the one or more subjects within each video frame;

computing optical flows from the video sequence;

extracting motion features of the video sequence based on the optical flows; and

predicting at least one future pose.

10. The non-transitory computer readable medium of claim 9 further comprising:

determining a current action label for each motion feature contained in the state information for a first frame of the series of video frames;

predicting, by a decoder, future action labels for each motion feature in a second frame of the series of video frames subsequent to the first frame, based on the current pose, action label and the state information;

predicting, by a decoder, future poses for each motion feature in the second frame based on the current poses and the state information; and

refining the current action label, the future action labels, and the future poses based on a loss function.

11. The non-transitory computer readable medium of claim 10 further comprising predicting at least one future pose and future motion for the second frame based on the model and the video sequence.

12. The non-transitory computer readable medium of claim 10 , wherein the encoder and the decoder are implemented using a recurrent neural network, implementing one or more gated recurrent network.

13. The non-transitory computer readable medium of claim 9 wherein the computing the optical flows forms two-channel flow frames where each channel contains displacements at x and y axis, respectively.

14. The non-transitory computer readable medium of claim 9 wherein the frame rate is between 24 fps and 60 fps.

15. The non-transitory computer readable medium of claim 9 wherein each joint location of the one or more joint locations comprises a pair of two dimensional coordinates (e.g., X,Y) or 3D coordinates (e.g., X,Y,Z).

16. A system for vision-based joint action and pose motion forecasting including a processor and a storage, the system comprising:

an encoder configured to extract a video sequence, the video sequence including a series of video frames;

the processor estimating current poses of one or more subject within each video frame and determining joint locations for one or more joints associated with one or more subjects within each video frame;

the processor computing optical flows from the video sequence;

the processor extracting motion features from the video sequence based on the optical flows; and

the processor predicting at least one future pose.

17. The system of claim 16 , where the processor is configured to perform:

determining a current action label for each motion feature contained in the state information for a first frame of the series of video frames;

predicting, by a decoder, future action labels for each motion feature in a second frame of the series of video frames subsequent to the first frame, based on the current pose, action label and the state information;

predicting, by a decoder, future poses for each motion feature in the second frame based on the current poses and the state information;

refining the current action label, the future action labels, and the future poses based on a loss function; and

predicting at least one future pose and future motion for the second frame based on the model and the video sequence.

18. The system of claim 17 wherein the encoder and the decoder are implemented using a recurrent neural network, implementing one or more gated recurrent network.

19. The system of claim 16 wherein the processor computing the optical flows forms two-channel flow frames where each channel contains displacements at x and y axis, respectively.

20. The system of claim 16 wherein each joint location of the one or more joint locations comprises a pair of two dimensional coordinates (e.g., X,Y) or 3D coordinates (e.g., X,Y,Z).

Assignments (2)
CHANGE OF NAME Recorded May 25, 2021
From: FUJI XEROX CO., LTD.
To: FUJIFILM BUSINESS INNOVATION CORP.
Reel/Frame 056392/0541 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2021
From: ZHU, YANJUN; ZHANG, YANXIA; LIU, QIONG; GIRGENSOHN, ANDREAS; AVRAHAMI, DANIEL; CHEN, FRANCINE; HU, HAO
To: FUJI XEROX CO., LTD.
Reel/Frame 055083/0936 →
Continuity (2)
Continuation 16816130 · Mar 11, 2020
Related Publication 20210289227A1 · Sep 16, 2021