IP Library Granted Patent US 12,377,863
Granted Patent B2
US 12,377,863 · App. 17/988,361 · Granted Aug 5, 2025

System and method for future forecasting using action priors

Inventors: Srikanth Malla (Sunnyvale, CA); Chiho Choi (San Jose, CA); Behzad Dariush (San Ramon, CA)
Assignee: Honda Motor Co., Ltd.
B60W50/0097B60W60/00272G06T7/70G06V10/764G06V10/82G06V20/58B60W2554/4029B60W2554/4049G06T2207/20081G06T2207/20084G06T2207/30261G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,377,863
App. No.
17/988,361
Granted
Aug 5, 2025
Kind
B2
Abstract

A system and method for future forecasting using action priors that include receiving image data associated with a surrounding environment of an ego vehicle and dynamic data associated with dynamic operation of the ego vehicle. The system and method also include analyzing the image data to classify dynamic objects as agents and to detect and annotate actions that are completed by the agents that are located within the surrounding environment of the ego vehicle and analyzing the dynamic data to process an ego motion history that is associated with the ego vehicle that includes vehicle dynamic parameters during a predetermined period of time. The system and method further include predicting future trajectories of the agents located within the surrounding environment of the ego vehicle and a future ego motion of the ego vehicle within the surrounding environment of the ego vehicle based on the annotated actions.

Claims (39)

1. A computer-implemented method for future forecasting using action priors comprising:

receiving image data associated with a surrounding environment of an ego vehicle and dynamic data associated with dynamic operation of the ego vehicle;

analyzing the image data to classify dynamic objects as agents and to detect and annotate actions that are completed by the agents that are located within the surrounding environment of the ego vehicle;

analyzing the dynamic data to process an ego motion history that is associated with the ego vehicle that includes vehicle dynamic parameters during a predetermined period of time;

predicting future trajectories of the agents located within the surrounding environment of the ego vehicle and a future ego motion of the ego vehicle within the surrounding environment of the ego vehicle based on the annotated actions that are completed by the agents and the ego motion history of the ego vehicle; and

controlling a vehicle controller of the ego vehicle to process and execute autonomous driving commands to operate the ego vehicle based on the future trajectories and the future ego motion, wherein

predicting the future trajectories of the agents and the future ego motion of the ego vehicle includes concurrently evaluating each agent within the surrounding environment as a distinct target agent and modeling pairwise interactions between that target agent and each agent in the surrounding environment, thereby generating a unique interaction context for the target agent that reflects individualized interactions between the target agent and each agent in the surrounding environment.

2. The computer-implemented method of claim 1 , wherein analyzing the image data includes inputting the image data to an image encoder of a neural network to determine the dynamic objects that are captured within at least one egocentric image of the surrounding environment of the ego vehicle.

3. The computer-implemented method of claim 2 , wherein analyzing the image data to classify the dynamic objects as the agents and to detect and annotate the actions includes computing bounding boxes around each of the agents within an image scene, wherein the computed bounding boxes are analyzed by obtaining a sequence of image patches from each bounding box.

4. The computer-implemented method of claim 3 , wherein analyzing the image data to classify the dynamic objects as the agents and to detect and annotate the actions includes accessing a localization dataset that includes a plurality of pre-trained action sets that are associated with various types of pedestrian related actions and vehicle related actions, wherein the plurality of pre-trained action sets are analyzed to detect and annotate the actions associated with the agents.

5. The computer-implemented method of claim 4 , wherein analyzing the image data to classify the dynamic objects as the agents and to detect and annotate the actions includes analyzing the computed bounding boxes with respect to the localization dataset to analyze the plurality of pre-trained action sets to detect atomic actions and contextual actions.

6. The computer-implemented method of claim 5 , wherein analyzing the dynamic data to process the ego motion history includes receiving a real-time ego motion of the ego vehicle and retrieving prior ego motions stored at one or more prior time stamps, wherein the real-time ego motion of the ego vehicle and the prior ego motions are aggregated to process the ego motion history, wherein the ego motion history includes vehicle dynamic parameters of the ego vehicle during a predetermined period of time.

7. The computer-implemented method of claim 6 , wherein predicting future trajectories of the agents located within the surrounding environment of the ego vehicle and the future ego motion of the ego vehicle includes the neural network concatenating data derived from the image data and the dynamic data to generate concatenated data, wherein the derived data includes actions of the agents, an ego motion of the ego vehicle, and interactions between the agents.

8. The computer-implemented method of claim 7 , wherein the concatenated data is inputted to the neural network to encode past object locations, wherein the encoded past object locations are inputted to a future ego motion decoder of the neural network to decode future bounding boxes of the agents at future time steps to predict the future trajectories of the agents located within the surrounding environment of the ego vehicle.

9. The computer-implemented method of claim 8 , wherein the future bounding boxes of the agents and the ego motion history of the ego vehicle are analyzed to predict the future ego motion of the ego vehicle during at least one future time step.

10. A system for future forecasting using action priors comprising:

a memory storing instructions when executed by a processor cause the processor to:

receive image data associated with a surrounding environment of an ego vehicle and dynamic data associated with dynamic operation of the ego vehicle;

analyze the image data to classify dynamic objects as agents and to detect and annotate actions that are completed by the agents that are located within the surrounding environment of the ego vehicle;

analyze the dynamic data to process an ego motion history that is associated with the ego vehicle that includes vehicle dynamic parameters during a predetermined period of time;

predict future trajectories of the agents located within the surrounding environment of the ego vehicle and a future ego motion of the ego vehicle within the surrounding environment of the ego vehicle based on the annotated actions that are completed by the agents and the ego motion history of the ego vehicle; and

control a vehicle controller of the ego vehicle to process and execute autonomous driving commands to operate the ego vehicle based on the future trajectories and the future ego motion, wherein

predicting the future trajectories of the agents and the future ego motion of the ego vehicle includes concurrently evaluating each agent within the surrounding environment as a distinct target agent and modeling pairwise interactions between that target agent and each agent in the surrounding environment, thereby generating a unique interaction context for the target agent that reflects individualized interactions between the target agent and each agent in the surrounding environment.

11. The system of claim 10 , wherein analyzing the image data includes inputting the image data to an image encoder of a neural network to determine the dynamic objects that are captured within at least one egocentric image of the surrounding environment of the ego vehicle.

12. The system of claim 11 , wherein analyzing the image data to classify the dynamic objects as the agents and to detect and annotate the actions includes computing bounding boxes around each of the agents within an image scene, wherein the computed bounding boxes are analyzed by obtaining a sequence of image patches from each bounding box.

13. The system of claim 12 , wherein analyzing the image data to classify the dynamic objects as the agents and to detect and annotate the actions includes accessing a localization dataset that includes a plurality of pre-trained action sets that are associated with various types of pedestrian related actions and vehicle related actions, wherein the plurality of pre-trained action sets are analyzed to detect and annotate the actions associated with the agents.

14. The system of claim 13 , wherein analyzing the image data to classify the dynamic objects as the agents and to detect and annotate the actions includes analyzing the computed bounding boxes with respect to the localization dataset to analyze the plurality of pre-trained action sets to detect atomic actions and contextual actions.

15. The system of claim 14 , wherein analyzing the dynamic data to process the ego motion history includes receiving a real-time ego motion of the ego vehicle and retrieving prior ego motions stored at one or more prior time stamps, wherein the real-time ego motion of the ego vehicle and the prior ego motions are aggregated to process the ego motion history, wherein the ego motion history includes vehicle dynamic parameters of the ego vehicle during a predetermined period of time.

16. The system of claim 15 , wherein predicting future trajectories of the agents located within the surrounding environment of the ego vehicle and the future ego motion of the ego vehicle includes the neural network concatenating data derived from the image data and the dynamic data to generate concatenated data, wherein the derived data includes actions of the agents, an ego motion of the ego vehicle, and interactions between the agents.

17. The system of claim 16 , wherein the concatenated data is inputted to the neural network to encode past object locations, wherein the encoded past object locations are inputted to a future ego motion decoder of the neural network to decode future bounding boxes of the agents at future time steps to predict the future trajectories of the agents located within the surrounding environment of the ego vehicle.

18. The system of claim 17 , wherein the future bounding boxes of the agents and the ego motion history of the ego vehicle are analyzed to predict the future ego motion of the ego vehicle during at least one future time step.

19. A non-transitory computer readable storage medium storing instructions that when executed by a computer, which includes a processor perform a method, the method comprising:

receiving image data associated with a surrounding environment of an ego vehicle and dynamic data associated with dynamic operation of the ego vehicle;

analyzing the image data to classify dynamic objects as agents and to detect and annotate actions that are completed by the agents that are located within the surrounding environment of the ego vehicle;

analyzing the dynamic data to process an ego motion history that is associated with the ego vehicle that includes vehicle dynamic parameters during a predetermined period of time;

predicting future trajectories of the agents located within the surrounding environment of the ego vehicle and a future ego motion of the ego vehicle within the surrounding environment of the ego vehicle based on the annotated actions that are completed by the agents and the ego motion history of the ego vehicle; and

controlling a vehicle controller of the ego vehicle to process and execute autonomous driving commands to operate the ego vehicle based on the future trajectories and the future ego motion, wherein

predicting the future trajectories of the agents and the future ego motion of the ego vehicle includes concurrently evaluating each agent within the surrounding environment as a distinct target agent and modeling pairwise interactions between that target agent and each agent in the surrounding environment, thereby generating a unique interaction context for the target agent that reflects individualized interactions between the target agent and each agent in the surrounding environment.

20. The non-transitory computer readable storage medium of claim 19 , wherein a neural network concatenates data derived from the image data and the dynamic data to generate concatenated data, wherein the derived data includes actions of the agents, an ego motion of the ego vehicle, and interactions between the agents, and the data is inputted to the neural network to encode past object locations, wherein the encoded past object locations are inputted to a future ego motion decoder of the neural network to decode future bounding boxes of the agents at future time steps to predict the future trajectories of the agents located within the surrounding environment of the ego vehicle.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2022
From: MALLA, SRIKANTH; CHOI, CHIHO; DARIUSH, BEHZAD
To: HONDA MOTOR CO., LTD.
Reel/Frame 061796/0445 →
Continuity (3)
Continuation 16913260 · Jun 26, 2020
Provisional Application 62929296 · Nov 1, 2019
Related Publication 20230081247A1 · Mar 16, 2023
References Cited (78)
US 10625748B1 · Dong · 2020 [cited by examiner]
US 10678240B2 · Pollach et al. · 2020 [cited by applicant]
US 10884409B2 · Mercep et al. · 2021 [cited by applicant]
US 11087477B2 · Choi · 2021 [cited by applicant]
US 11308338B2 · Yang et al. · 2022 [cited by applicant]
US 11710352B1 · Ulutan · 2023 [cited by examiner]
US 20170010618A1 · Shashua et al. · 2017 [cited by applicant]
US 20180024562A1 · Bellaiche · 2018 [cited by applicant]
US 20180025235A1 · Fridman · 2018 [cited by applicant]
US 20180307935A1 · Rao · 2018 [cited by examiner]
US 20190228316A1 · Felsen · 2019 [cited by examiner]
US 20190258251A1 · Ditty et al. · 2019 [cited by applicant]
US 20190291728A1 · Shalev-Shwartz et al. · 2019 [cited by applicant]
US 20190333381A1 · Shalev-Shwartz et al. · 2019 [cited by applicant]
US 20190361439A1 · Zeng et al. · 2019 [cited by applicant]
US 20190371025A1 · Sukthankar · 2019 [cited by examiner]
US 20190384294A1 · Shashua et al. · 2019 [cited by applicant]
US 20200026282A1 · Choe et al. · 2020 [cited by applicant]
US 20200086863A1 · Rosman · 2020 [cited by examiner]
US 20200184233A1 · Berberian et al. · 2020 [cited by applicant]
US 20200276988A1 · Graves · 2020 [cited by applicant]
US 20200309541A1 · Lavy et al. · 2020 [cited by applicant]
US 20200353943A1 · Siddiqui · 2020 [cited by examiner]
US 20210101616A1 · Hayat et al. · 2021 [cited by applicant]
US 20220082403A1 · Shapira et al. · 2022 [cited by applicant]
US 20220126864A1 · Moustafa et al. · 2022 [cited by applicant]
JP 2017142735 · 2017 [cited by applicant]
WO WO2019106789 · 2019 [cited by applicant]
Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. “Social Istm: Human trajectory prediction in crowded spaces.” In Proceedings of the IEEE Conference on Computer V… [cited by applicant]
Apratim Bhattacharyya, Mario Fritz, and Bernt Schiele. “Long-term on-board prediction of people in traffic scenes under uncertainty.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.… [cited by applicant]
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. “Activitynet: A large-scale video benchmark for human activity understanding.” In Proceedings of the IEEE Conference on Computer Vision and… [cited by applicant]
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. “nuscenes: A multimodal dataset for autonomous driving.” arXiv preprint … [cited by applicant]
Joao Carreira and Andrew Zisserman. “Quo vadis, action recognition? a new model and the kinetics dataset.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6299-6308, 2017. [cited by applicant]
Rohan Chandra, Uttaran Bhattacharya, Aniket Bera, and Dinesh Manocha. “Traphic: Trajectory prediction in dense and heterogeneous traffic using weighted interactions.” In Proceedings of the IEEE Conference on Computer Vi… [cited by applicant]
Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, et al. “Argoverse: 3d tracking and forecasting with rich maps.” In Proceedin… [cited by applicant]
Chiho Choi and Behzad Dariush. “Looking to relations for future trajectory forecast.” In The IEEE International Conference on Computer Vision (ICCV), Oct. 2019. [cited by applicant]
Chiho Choi, Abhishek Patil, and Srikanth Malla. “Drogon: A causal reasoning framework for future trajectory forecast.” arXiv preprint arXiv:1908.00024, 2019. [cited by applicant]
Nachiket Deo and Mohan M. Trivedi. “Multi-modal trajectory prediction of surrounding vehicles with maneuver based Istms.” In 2018 IEEE Intelligent Vehicles Symposium (IV), pp. 1179-1184. IEEE, 2018. [cited by applicant]
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. “Vision meets robotics: The kitti dataset.” The International Journal of Robotics Research, 32(11):1231-1237, 2013. [cited by applicant]
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. “Ava: A video dataset of spatio-temporally localize… [cited by applicant]
Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. “Social gan: Socially acceptable trajectories with generative adversarial networks.” In Proceedings of the IEEE Conference on Computer Visio… [cited by applicant]
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?” In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 6546-6555… [cited by applicant]
Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders. “Timeception for complex action recognition.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 254-263, 2019. [cited by applicant]
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. “Large-scale video classification with convolutional neural networks.” In Proceedings of the IEEE conference on Computer … [cited by applicant]
R. Kesten, M. Usman, J. Houston, T. Pandya, K. Nadhamuni, A. Ferreira, M. Yuan, B. Low, A. Jain, P. Ondruska, S. Omari, S. Shah, A. Kulkarni, A. Kazakova, C. Tao, L. Platinsky, W. Jiang, and V. Shet. “Lyft level 5 av da… [cited by applicant]
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. “HMDB: a large video database for human motion recognition.” In Proceedings of the International Conference on Computer Vision (ICCV), 2011. [cited by applicant]
Tian Lan, Tsung-Chuan Chen, and Silvio Savarese. “A hierarchical representation for future action prediction.” In European Conference on Computer Vision, pp. 689-704. Springer, 2014. [cited by applicant]
Namhoon Lee, Wongun Choi, Paul Vernaza, Christopher B. Choy, Philip HS Torr, and Manmohan Chandraker. “Desire: Distant future prediction in dynamic scenes with interacting agents.” In Proceedings of the IEEE Conference … [cited by applicant]
Alon Lerner, Yiorgos Chrysanthou, and Dani Lischinski. “Crowds by example.” In Computer graphics forum, vol. 26, pp. 655-664. Wiley Online Library, 2007. [cited by applicant]
Jiachen Li, Hengbo Ma, and Masayoshi Tomizuka. “Conditional generative neural system for probabilistic trajectory prediction.” In 2019 IEEE Conference on Robotics and Systems (IROS), 2019. [cited by applicant]
Jiachen Li, Hengbo Ma, and Masayoshi Tomizuka. “Interaction-aware multi-agent tracking and probabilistic behavior prediction via adversarial learning.” In 2019 IEEE In-ternational Conference on Robotics and Automation (… [cited by applicant]
Yuexin Ma, Xinge Zhu, Sibo Zhang, Ruigang Yang, Wenping Wang, and Dinesh Manocha. “Trafficpredict: Trajectory prediction for heterogeneous traffic-agents.” In Proceedings of the AAAI Conference on Artificial Intelligenc… [cited by applicant]
Srikanth Malla and Chiho Choi. “Nemo: Future object localization using noisy ego priors.” arXiv preprint arXiv:1909.08150, 2019. [cited by applicant]
Sangmin Oh, Anthony Hoogs, Amitha Perera, Naresh Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, JK Aggarwal, Hyungtae Lee, Larry Davis, et al. “A large-scale benchmark dataset for event recognition in surve… [cited by applicant]
Seong Hyeon Park, ByeongDo Kim, Chang Mook Kang, Chung Choo Chung, and Jun Won Choi. “Sequence-to-sequence prediction of vehicle trajectory via lstm encoder-decoder architecture.” In 2018 IEEE Intelligent Vehicles Sympo… [cited by applicant]
Abhishek Patil, Srikanth Malla, Haiming Gang, and Yi-Ting Chen. The h3d dataset for full-surround 3d multi-object detection and tracking in crowded urban scenes. arXiv preprint arXiv:1903.01568, 2019. [cited by applicant]
Stefano Pellegrini, Andreas Ess, and Luc Van Gool. “Improving data association by joint modeling of pedestrian trajectories and groupings.” In European conference on computer vision, pp. 452-465. Springer, 2010. [cited by applicant]
Amir Rasouli, Iuliia Kotseruba, Toni Kunic, and John K. Tsotsos. “Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction.” In The IEEE International Conference on Computer Vi… [cited by applicant]
Amir Rasouli, Iuliia Kotseruba, and John K Tsotsos. “Are they going to cross? a benchmark dataset and baseline for pedestrian crosswalk behavior.” In Proceedings of the IEEE International Conference on Computer Vision, … [cited by applicant]
Nicholas Rhinehart, Kris M Kitani, and Paul Vernaza. “R2p2: A reparameterized pushforward policy for diverse, precise generative path forecasting.” In Proceedings of the European Conference on Computer Vision (ECCV), pp… [cited by applicant]
Alexandre Robicquet, Amir Sadeghian, Alexandre Alahi, and Silvio Savarese. “Learning social etiquette: Human trajectory understanding in crowded scenes.” In European conference on computer vision, pp. 549-565. Springer,… [cited by applicant]
Christoph Schöller, Vincent Aravantinos, Florian Lay, and Alois Knoll. The simpler the better: Constant velocity for pedestrian motion prediction. arXiv preprint arXiv:1903.07933, 2019. [cited by applicant]
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. “Actor and observer: Joint modeling of first and third-person videos.” In Proceedings of the IEEE Conference on Computer Vision and … [cited by applicant]
Gunnar A. Sigurdsson, Gul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. “Hollywood in homes: Crowdsourcing data collection for activity understanding.” In European Conference on Computer Vision, 201… [cited by applicant]
Karen Simonyan and Andrew Zisserman. “Two-stream convolutional networks for action recognition in videos.” In Advances in neural information processing systems, pp. 568-576, 2014. [cited by applicant]
Chaoming Song, Zehui Qu, Nicholas Blumm, and Albert-László Barabási. “Limits of predictability in human mobility.” Science, 327(5968):1018-1021, 2010. [cited by applicant]
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. “Ucf101: A dataset of 101 human actions classes from videos in the wild.” arXiv preprint arXiv:1212.0402, 2012. [cited by applicant]
Chen Sun, Abhinav Shrivastava, Carl Vondrick, Kevin Murphy, Rahul Sukthankar, and Cordelia Schmid. “Actor-centric relation network.” In Proceedings of the European Conference on Computer Vision (ECCV), pp. 318-334, 2018. [cited by applicant]
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. “Learning spatiotemporal features with 3d convolutional networks.” In Proceedings of the IEEE international conference on computer vision, pp.… [cited by applicant]
Gül Varol, Ivan Laptev, and Cordelia Schmid. “Long-term temporal convolutions for action recognition.” IEEE transactions on pattern analysis and machine intelligence, 40(6):1510-1517, 2017. [cited by applicant]
Anirudh Vemula, Katharina Muelling, and Jean Oh. “Social attention: Modeling attention in human crowds.” In 2018 IEEE International Conference on Robotics and Automation (ICRA), pp. 1-7. IEEE, 2018. [cited by applicant]
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. “Non-local neural networks.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7794-7803, 2018. [cited by applicant]
Yanyu Xu, Zhixin Piao, and Shenghua Gao. “Encoding crowd interaction with deep neural network for pedestrian trajectory prediction.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. … [cited by applicant]
Hao Xue, Du Q. Huynh, and Mark Reynolds. “Ss-Istm: A hierarchical Istm model for pedestrian trajectory prediction.” In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1186-1194. IEEE, 2018. [cited by applicant]
Yu Yao, Mingze Xu, Chiho Choi, David J Crandall, Ella M Atkins, and Behzad Dariush. “Egocentric vision-based future vehicle localization for intelligent driving assistance systems.” In 2019 International Conference on R… [cited by applicant]
Wei Zhan, Liting Sun, Di Wang, Haojie Shi, Aubrey Clausse, Maximilian Naumann, Julius Kummerle, Hendrik Konigshof, Christoph Stiller, Arnaud de La Fortelle, et al. “Interaction dataset: An international, adversarial and… [cited by applicant]
Pu Zhang, Wanli Ouyang, Pengfei Zhang, Jianru Xue, and Nanning Zheng. “Sr-Istm: State refinement for Istm towards pedestrian trajectory prediction.” In Proceedings of the IEEE Conference on Computer Vision and Pattern R… [cited by applicant]
Waymo open dataset: An autonomous driving dataset, 2019. [cited by applicant]