IP Library Granted Patent US 11,074,438
Granted Patent B2
US 11,074,438 · App. 16/590,275 · Granted Jul 27, 2021

Disentangling human dynamics for pedestrian locomotion forecasting with noisy supervision

Inventors: Karttikeya Mangalam (Stanford, CA); Ehsan Adeli-Mosabbeb (Mountain View, CA); Kuan-Hui Lee (San Jose, CA); Adrien Gaidon (Mountain View, CA); Juan Carlos Niebles Duque (Mountain View, CA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
G06K9/00342G06K9/00744G06K9/00805G06N3/08G06T7/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,074,438
App. No.
16/590,275
Granted
Jul 27, 2021
Kind
B2
Abstract

A method for predicting spatial positions of several key points on a human body in the near future in an egocentric setting is described. The method includes generating a frame-level supervision for human poses. The method also includes suppressing noise and filling missing joints of the human body using a pose completion module. The method further includes splitting the poses into a global stream and a local stream. Furthermore, the method includes combining the global stream and the local stream to forecast future human locomotion.

Claims (31)

1. A method for predicting spatial positions of several key points on a human body in the near future in an egocentric setting, comprising:

generating a frame-level supervision for human poses of a pedestrian;

suppressing noise and filling missing joints of the human body of the pedestrian using a pose completion module;

splitting the concurrent motion of the joints of the pedestrian into a global stream and a local stream, in which the global motion stream models large scale movements of a position of the pedestrian with respect to a camera of a vehicle;

encoding, by the local stream, a motion of the human body of the pedestrian with respect to the global stream;

capturing a depth change of an overall pose size and movement of different joints of the human body of the pedestrian; and

combining the global stream and the local stream to forecast future human locomotion of the pedestrian based on the depth change of the overall pose size and movement of different joints of the human body of the pedestrian.

2. The method of claim 1 , further comprising forecasting the global stream based on one joint of the human body of the pedestrian.

3. The method of claim 1 , in which an architecture of a forecaster of the local stream is different from an architecture of a forecaster of the global stream.

4. The method of claim 1 , further comprising estimating motion between consecutive frames due to a motion of an ego vehicle.

5. A non-transitory computer-readable medium having program code recorded thereon for predicting spatial positions of several key points on a human body in the near future in an egocentric setting, the program code being executed by a processor and comprising:

program code to generate a frame-level supervision for human poses of a pedestrian;

program code to suppress noise and filling missing joints of the human body of the pedestrian using a pose completion module;

program code to split the concurrent motion of the joints of the pedestrian into a global stream and a local stream, in which the global motion stream models large scale movements of a position of the pedestrian with respect to a camera of a vehicle;

program code to encode, by the local stream, a motion of the human body of the pedestrian with respect to the global stream;

program code to capture a depth change of an overall pose size and movement of different joints of the human body of the pedestrian; and

program code to combine the global stream and the local stream to forecast future human locomotion of the pedestrian based on the depth change of the overall pose size and movement of different joints of the human body of the pedestrian.

6. The non-transitory computer-readable medium of claim 5 , further comprising program code to forecast the global stream based on one joint of the human body of the pedestrian.

7. The non-transitory computer-readable medium of claim 5 , in which an architecture of a forecaster of the local stream is different from an architecture of a forecaster of the global stream.

8. The non-transitory computer-readable medium of claim 5 , further comprising program code to estimate motion between consecutive frames due to a motion of an ego vehicle.

9. A system for predicting spatial positions of several key points on a human body in the near future in an egocentric setting, the system comprising:

a memory; and

at least one processor, the at least one processor configured:

to generate a frame-level supervision for human poses of a pedestrian;

to suppress noise and filling missing joints of the human body of the pedestrian using a pose completion module;

to split the concurrent motion of the joints of the pedestrian into a global stream and a local stream, in which the global motion stream models large scale movements of a position of the pedestrian with respect to a camera of a vehicle;

to encode, by the local stream, a motion of the human body of the pedestrian with respect to the global stream;

to capture a depth change of an overall pose size and movement of different joints of the human body of the pedestrian; and

to combine the global stream and the local stream to forecast future human locomotion of the pedestrian based on the depth change of the overall pose size and movement of different joints of the human body of the pedestrian.

10. The system of claim 9 , in which the at least one processor is further configured to forecast the global stream based on one joint of the human body of the pedestrian.

11. The system of claim 9 , in which an architecture of a forecaster of the local stream is different from an architecture of a forecaster of the global stream.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2021
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 058460/0299 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2019
From: LEE, KUAN-HUI; GAIDON, ADRIEN
To: TOYOTA RESEARCH INSTITUTE, INC.
Reel/Frame 050608/0650 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2019
From: MANGALAM, KARTTIKEYA; ADELI-MOSABBEB, EHSAN; NIEBLES DUQUE, JUAN CARLOS
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
Reel/Frame 050608/0800 →
Continuity (1)
Related Publication 20210097266A1 · Apr 1, 2021
Cited By (2)
US 12,233,907 US 12,236,705