IP Library Granted Patent US 12,258,047
Granted Patent B2
US 12,258,047 · App. 17/352,540 · Granted Mar 25, 2025

System and method for providing long term and key intentions for trajectory prediction

Inventors: Harshayu Girase (Union City, CA); Haiming Gang (San Jose, CA); Srikanth Malla (Sunnyvale, CA); Jiachen Li (Albany, CA); Akira Kanehara (Utsunomiya, JP); Chiho Choi (San Jose, CA)
Assignee: Honda Motor Co., Ltd.
B60W60/0027G01S17/86G01S17/894G06T7/20B60W2420/403B60W2420/408B60W2554/4029B60W2554/4045B60W2556/10G06T2207/10024G06T2207/10028G06T2207/20072G06T2207/30196G06T2207/30236G06T2207/30241G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,258,047
App. No.
17/352,540
Granted
Mar 25, 2025
Kind
B2
Abstract

A system and method for providing long term and key intentions for trajectory prediction that include receiving image data and LiDAR data associated with RGB images and LiDAR point clouds that are associated with a surrounding environment of an ego agent and processing a long term and key intentions for trajectory prediction dataset (LOKI dataset) that is utilized to complete joint trajectory and intention prediction for heterogeneous traffic agents. The system and method also include encoding a past observation history of each of the heterogeneous traffic agents and sampling a respective goal. The system and method further include decoding and predicting future trajectories associated with each of the heterogeneous traffic agents based on data included within the LOKI dataset, the encoded past observation history, and the respective goal.

Claims (43)

1. A computer-implemented method for providing long term and key intentions for trajectory prediction, comprising:

receiving and aggregating image data and LiDAR data associated with RGB images and LiDAR point clouds that are associated with a surrounding environment of an ego agent, wherein

the aggregated image data and LiDAR data is aggregated environment data;

processing a long term and key intentions for trajectory prediction dataset (LOKI dataset) that is utilized to complete joint trajectory and intention prediction for heterogeneous traffic agents, wherein the LOKI dataset is populated with annotations that include the aggregated environment data and annotated labels that pertain to attributes that influence agent intent for each of the heterogeneous traffic agents;

encoding a past observation history of each of the heterogeneous traffic agents and sampling a respective goal; and

decoding and predicting future trajectories associated with each of the heterogeneous traffic agents based on data included within the LOKI dataset, the encoded past observation history, and the respective goal, wherein

the annotated labels include contextual labels that are associated with factors that affect future behavior of each of the heterogeneous traffic agents including at least one of weather and road conditions,

a scene graph derived from the LOKI dataset is constructed that includes nodes including agent trajectory information, agent intent information, and goal information of the heterogenous traffic agents located within the surrounding environment of the ego agent, and nodes denoting road entrance and road exit information which provide the heterogeneous traffic agents with map topology information pertaining to the surrounding environment, and

decoding and predicting future trajectories includes a decoder being configured to access the LOKI dataset to analyze the LOKI dataset in addition to the scene graph to determine the predicted intentions, goals, and past motion of all of the heterogenous traffic agents and a scene of the surrounding environment of the ego agent to predict the trajectories of each of the heterogenous traffic agents that are located within the surrounding environment of the ego agent.

2. The computer-implemented method of claim 1 , wherein the RGB images and LiDAR point clouds capture the heterogeneous traffic agents that are located within the surrounding environment, wherein the heterogeneous traffic agents include pedestrians and vehicles.

3. The computer-implemented method of claim 2 , wherein the annotated labels include intention labels that are associated with intentions of the pedestrians and the vehicles.

4. The computer-implemented method of claim 2 , wherein the annotated labels include environment labels that pertain to the surrounding environment of the ego agent.

5. The computer-implemented method of claim 1 , wherein the annotated labels include the contextual labels that are associated with the factors that affect the future behavior of each of the heterogeneous traffic agents including at least one of the weather, the road conditions, gender, and age.

6. The computer-implemented method of claim 1 , wherein encoding the past observation history includes inputting actor states and actor trajectories of each of the heterogenous agents within an observation encoder to process a long-term goal proposal that proposes a long-term goal distribution over potential final destination of each heterogeneous traffic agent independently.

7. The computer-implemented method of claim 1 , further including controlling at least one system of the ego agent to operate the ego agent within the surrounding environment of the ego agent based on the predicted future trajectories associated with each of the heterogeneous traffic agents.

8. The computer-implemented method of claim 1 , wherein the annotated labels include descriptions indicating a relative position of static and dynamic objects in the surrounding environment with respect to the ego agent.

9. A system for providing long term and key intentions for trajectory prediction, comprising:

a memory storing instructions when executed by a processor cause the processor to:

receive and aggregate image data and LiDAR data associated with RGB images and LiDAR point clouds that are associated with a surrounding environment of an ego agent, wherein

the aggregated image data and LiDAR data is aggregated environment data;

process a long term and key intentions for trajectory prediction dataset (LOKI dataset) that is utilized to complete joint trajectory and intention prediction for heterogeneous traffic agents, wherein the LOKI dataset is populated with annotations that include the aggregated environment data and annotated labels that pertain to attributes that influence agent intent for each of the heterogeneous traffic agents;

encode a past observation history of each of the heterogeneous traffic agents and sampling a respective goal; and

decode and predict future trajectories associated with each of the heterogeneous traffic agents based on data included within the LOKI dataset, the encoded past observation history, and the respective goal, wherein

the annotated labels include contextual labels that are associated with factors that affect future behavior of each of the heterogeneous traffic agents including at least one of weather and road conditions,

a scene graph derived from the LOKI dataset is constructed that includes nodes including agent trajectory information, agent intent information, and goal information of the heterogenous traffic agents located within the surrounding environment of the ego agent, and nodes denoting entrance and road exit information which provide the heterogeneous traffic agents with map topology information pertaining to the surrounding environment, and

decoding and predicting future trajectories includes a decoder being configured to access the LOKI dataset to analyze the LOKI dataset in addition to the scene graph to determine the predicted intentions, goals, and past motion of all of the heterogenous traffic agents and a scene of the surrounding environment of the ego agent to predict the trajectories of each of the heterogenous traffic agents that are located within the surrounding environment of the ego agent.

10. The system of claim 9 , wherein the RGB images and LiDAR point clouds capture the heterogeneous traffic agents that are located within the surrounding environment, wherein the heterogeneous traffic agents include pedestrians and vehicles.

11. The system of claim 10 , wherein the annotated labels include intention labels that are associated with intentions of the pedestrians and the vehicles.

12. The system of claim 10 , wherein the annotated labels include environment labels that pertain to the surrounding environment of the ego agent.

13. The system of claim 9 , wherein the annotated labels include the contextual labels that are associated with the factors that affect the future behavior of each of the heterogeneous traffic agents including at least one of the weather, the road conditions, gender, and age.

14. The system of claim 9 , wherein encoding the past observation history includes inputting actor states and actor trajectories of each of the heterogenous agents within an observation encoder to process a long-term goal proposal that proposes a long-term goal distribution over potential final destination of each heterogeneous traffic agent independently.

15. The system of claim 9 , further including controlling at least one system of the ego agent to operate the ego agent within the surrounding environment of the ego agent based on the predicted future trajectories associated with each of the heterogeneous traffic agents.

16. The system of claim 9 , wherein the annotated labels include descriptions indicating a relative position of static and dynamic objects in the surrounding environment with respect to the ego agent.

17. A non-transitory computer readable storage medium storing instruction that when executed by a computer, which includes a processor perform a method, the method comprising:

receiving and aggregating image data and LiDAR data associated with RGB images and LiDAR point clouds that are associated with a surrounding environment of an ego agent, wherein

the aggregated image data and LiDAR data is aggregated environment data;

processing a long term and key intentions for trajectory prediction dataset (LOKI dataset) that is utilized to complete joint trajectory and intention prediction for heterogeneous traffic agents, wherein the LOKI dataset is populated with annotations that include the aggregated environment data and annotated labels that pertain to attributes that influence agent intent for each of the heterogeneous traffic agents;

encoding a past observation history of each of the heterogeneous traffic agents and sampling a respective goal; and

decoding and predicting future trajectories associated with each of the heterogeneous traffic agents based on data included within the LOKI dataset, the encoded past observation history, and the respective goal, wherein

the annotated labels include contextual labels that are associated with factors that affect future behavior of each of the heterogeneous traffic agents including at least one of weather and road conditions,

a scene graph derived from the LOKI dataset is constructed that includes nodes including agent trajectory information, agent intent information, and goal information of the heterogenous traffic agents located within the surrounding environment of the ego agent, and nodes denoting road entrance and road exit information which provide the heterogeneous traffic agents with map topology information pertaining to the surrounding environment, and

decoding and predicting future trajectories includes a decoder being configured to access the LOKI dataset to analyze the LOKI dataset in addition to the scene graph to determine the predicted intentions, goals, and past motion of all of the heterogenous traffic agents and a scene of the surrounding environment of the ego agent to predict the trajectories of each of the heterogenous traffic agents that are located within the surrounding environment of the ego agent.

18. The non-transitory computer readable storage medium of claim 17 , further including controlling at least one system of the ego agent to operate the ego agent within the surrounding environment of the ego agent based on the predicted future trajectories associated with each of the heterogeneous traffic agents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2021
From: GIRASE, HARSHAYU; GANG, HAIMING; MALLA, SRIKANTH; LI, JIACHEN; KANEHARA, AKIRA; CHOI, CHIHO
To: HONDA MOTOR CO., LTD.
Reel/Frame 056599/0911 →
Continuity (2)
Provisional Application 63166195 · Mar 25, 2021
Related Publication 20220306160A1 · Sep 29, 2022
References Cited (45)
US 11640174B2 · Tran · 2023 [cited by examiner]
US 20190382007A1 · Casas · 2019 [cited by examiner]
US 20200039520A1 · Misu et al. · 2020 [cited by applicant]
US 20200039521A1 · Misu et al. · 2020 [cited by applicant]
US 20200264609A1 · Hammond · 2020 [cited by examiner]
US 20200401135A1 · Chen et al. · 2020 [cited by applicant]
US 20210146949A1 · Martinez Covarrubias · 2021 [cited by examiner]
US 20210347377A1 · Siebert · 2021 [cited by examiner]
CN 110895674 · 2020 [cited by applicant]
Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. Social lstm: Human trajectory prediction in crowded spaces. In Proceedings of the IEEE conference on computer vis… [cited by applicant]
Alexandre Alahi, Vignesh Ramanathan, and Li Fei-Fei. Socially-aware large-scale crowd forecasting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2203-2210, 2014. [cited by applicant]
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of… [cited by applicant]
Susan Carey and Elizabeth Spelke. Domain-specific knowledge and conceptual change. Mapping the mind: Domain specificity in cognition and culture, 169:200, 1994. [cited by applicant]
Sergio Casas, Wenjie Luo, and Raquel Urtasun. Intentnet: Learning to predict intention from raw sensor data. In Conference on Robot Learning, pp. 947-956. PMLR, 2018. [cited by applicant]
Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, DeWang, Peter Carr, Simon Lucey, Deva Ramanan, et al. Argoverse: 3d tracking and forecasting with rich maps. In Proceedings … [cited by applicant]
Chiho Choi, Srikanth Malla, Abhishek Patil, and Joon Hee Choi. Drogon: A trajectory prediction model based on Intention-conditioned behavior reasoning. In Proceedings of the Conference on Robot Learning, 2020. [cited by applicant]
Patrick Dendorfer, Aljosa Osep, and Laura Leal-Taixe. Goalgan: Multimodal trajectory prediction based on goal position estimation. In Proceedings of the Asian Conference on Computer Vision, 2020. [cited by applicant]
Nachiket Deo and Mohan M Trivedi. Trajectory forecasts in unknown environments conditioned on grid-based plans. arXiv preprint arXiv:2001.00735, 2020. [cited by applicant]
Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. Social gan: Socially acceptable trajectories with generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision … [cited by applicant]
Dirk Helbing and Peter Molnar. Social force model for pedestrian dynamics. Physical review E, 51(5):4282, 1995. [cited by applicant]
Vineet Kosaraju, Amir Sadeghian, Roberto Martin-Martin, Ian Reid, S Hamid Rezatofighi, and Silvio Savarese. Socialbigat: Multimodal trajectory forecasting using bicyclegan and graph attention networks. arXiv preprint ar… [cited by applicant]
Parth Kothari, Sven Kreiss, and Alexandre Alahi. Human trajectory forecasting in crowds: A deep learning perspective. arXiv preprint arXiv:2007.03639, 2020. [cited by applicant]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25:1097-1105, 2012. [cited by applicant]
Sumit Kumar, Yiming Gu, Jerrick Hoang, Galen Clark Haynes, and Micol Marchetti-Bowick. Interaction-based trajectory prediction over a hybrid traffic graph. arXiv preprint arXiv:2009.12916, 2020. Yann LeCun. The mnist da… [cited by applicant]
Namhoon Lee, Wongun Choi, Paul Vernaza, Christopher B Choy, Philip HS Torr, and Manmohan Chandraker. Desire: Distant future prediction in dynamic scenes with interacting agents. In Proceedings of the IEEE Conference on … [cited by applicant]
Jiachen Li, Hengbo Ma, and Masayoshi Tomizuka. Conditional generative neural system for probabilistic trajectory prediction. arXiv preprint arXiv:1905.01631, 2019. [cited by applicant]
Matteo Lisotto, Pasquale Coscia, and Lamberto Ballan. Social and scene-aware trajectory prediction in crowded spaces. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pp. 0-0, 2019. [cited by applicant]
Bingbin Liu, Ehsan Adeli, Zhangjie Cao, Kuan-Hui Lee, Abhijeet Shenoi, Adrien Gaidon, and Juan Carlos Niebles. Spatiotemporal relationship reasoning for pedestrian intent prediction, 2020. [cited by applicant]
Srikanth Malla, Behzad Dariush, and Chiho Choi. Titan: Future forecast using action priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11186-11196, 2020. [cited by applicant]
Karttikeya Mangalam, Yang An, Harshayu Girase, and Jitendra Malik. From goals, waypoints & paths to long term human trajectory forecasting. arXiv preprint arXiv:2012.01526, 2020. [cited by applicant]
Karttikeya Mangalam, Harshayu Girase, Shreyas Agarwal, Kuan-Hui Lee, Ehsan Adeli, Jitendra Malik, and Adrien Gaidon. It is not the journey but the destination: Endpoint conditioned trajectory prediction. In European Con… [cited by applicant]
Abduallah Mohamed, Kun Qian, Mohamed Elhoseiny, and Christian Claudel. Social-stgcnn: A social spatio-temporal graph convolutional neural network for human trajectory prediction. In Proceedings of the IEEE/CVF Conferenc… [cited by applicant]
Amir Rasouli. Pedestrian simulation: A review. arXiv preprint arXiv:2102.03289, 2021. [cited by applicant]
Amir Rasouli, Iuliia Kotseruba, Toni Kunic, and John K Tsotsos. Pie: A large-scale dataset and models for pedestrian Intention estimation and trajectory prediction. In Proceedings of the IEEE/CVF International Conferenc… [cited by applicant]
Amir Rasouli, Iuliia Kotseruba, and John K Tsotsos. Are they going to cross? a benchmark dataset and baseline for pedestrian crosswalk behavior. In Proceedings of the IEEE International Conference on Computer Vision Wor… [cited by applicant]
Nicholas Rhinehart, Rowan McAllister, Kris Kitani, and Sergey Levine. Precog: Prediction conditioned on goals in visual multi-agent settings. arXiv preprint arXiv:1905.01296, 2019. [cited by applicant]
A Robicquet, A Sadeghian, A Alahi, and S Savarese. Learning social etiquette: Human trajectory prediction in crowded scenes. In European Conference on Computer Vision (ECCV), 2020. [cited by applicant]
Andrey Rudenko, Luigi Palmieri, Michael Herman, Kris M Kitani, Dariu M Gavrila, and Kai O Arras. Human motion trajectory prediction: A survey. The International Journal of Robotics Research, 39(8):895-935, 2020. [cited by applicant]
Amir Sadeghian, Vineet Kosaraju, Ali Sadeghian, Noriaki Hirose, Hamid Rezatofighi, and Silvio Savarese. Sophie: An attentive gan for predicting paths compliant to social and physical constraints. In Proceedings of the I… [cited by applicant]
Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, and Marco Pavone. Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data. arXiv preprint arXiv:2001.03093, 2020. [cited by applicant]
Yunsheng Shi, Zhengjie Huang, Shikun Feng, and Yu Sun. Masked label prediction: Unified massage passing model for semi-supervised classification. arXiv preprint arXiv:2009.03509, 2020. [cited by applicant]
Vivian V Valentin, Anthony Dickinson, and John P O'Doherty. Determining the neural substrates of goaldirected learning in the human brain. Journal of Neuroscience, 27(15):4019-4026, 2007. [cited by applicant]
Takuma Yagi, Karttikeya Mangalam, Ryo Yonetani, and Yoichi Sato. Future person localization in first-person videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7593-7602, 2018. [cited by applicant]
Lingyao Zhang, Po-Hsun Su, Jerrick Hoang, Galen Clark Haynes, and Micol Marchetti-Bowick. Map-adaptive goal-based trajectory prediction. arXiv preprint arXiv:2009.04450, 2020. [cited by applicant]
Hang Zhao, Jiyang Gao, Tian Lan, Chen Sun, Benjamin Sapp, Balakrishnan Varadarajan, Yue Shen, Yi Shen, Yuning Chai, Cordelia Schmid, et al. Tnt: Target-driven trajectory prediction. arXiv preprint arXiv:2008.08294, 2020. [cited by applicant]