IP Library Granted Patent US 12,548,331
Granted Patent B2
US 12,548,331 · App. 18/118,995 · Granted Feb 10, 2026

Techniques to perform trajectory predictions

Inventors: Xinshuo Weng (North York, CA); Boris Ivanovic (Mountain View, CA); Marco Pavone (Stanford, CA)
Assignee: NVIDIA Corporation
G06V20/41G01C21/3874
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,331
App. No.
18/118,995
Granted
Feb 10, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to perform trajectory predictions within one or more images. In at least one embodiment, a processor comprises one or more circuits to cause one or more neural networks to perform trajectory predictions of two or more objects detected within a plurality of frames without tracking the two or more objects based, at least in part, on processing a sequence of data of the one or more objects as a whole.

Claims (54)

1 . A system, comprising:

at least one processor;

at least one memory comprising instructions that, in response to execution by the at least one processor, cause the system to at least:

obtain first data that indicates that two or more objects detected within consecutive frames have similar properties;

obtain second data that indicates positions of the two or more objects detected within each frame of the consecutive frames;

assign, using a first portion of a neural network, weights to a first portion of the second data based, at least in part, on a second portion of the second data and the first data;

generate context information of the two or more objects based, at least in part, on the first portion of the second data with the weights; and

generate, using a second portion of the neural network, one or more sets of trajectories associated with the two or more objects based, at least in part, on the context information.

2 . The system of claim 1 , wherein:

the neural network is a transformer neural network;

the first portion includes an encoder of the transformer neural network; and

the second portion includes a decoder of the transformer neural network.

3 . The system of claim 1 , wherein the second data is generated based, at least in part, on map data that indicates an area where the two or more objects are located.

4 . The system of claim 1 , wherein the at least one memory comprises further instructions that, in response to execution by the at least one processor, cause the system to at least generate at least one latent value based, at least in part, on the context information.

5 . The system of claim 1 , wherein the at least one memory comprises further instructions that, in response to execution by the at least one processor, cause the system to at least generate an embedded data based, at least in part, on the second data and one or more prior sets of trajectories associated with the two or more objects.

6 . The system of claim 1 , wherein at least one frame of the consecutive frames are generated based, at least in part, on interpolation of two or more frames of the consecutive frames.

7 . The system of claim 1 , wherein the neural network is trained based, at least in part, on at least two distinct loss functions.

8 . A method comprising:

generating first data that indicates that two or more objects detected within consecutive frames have similar properties;

obtaining second data that indicates position of the two or more objects detected within each frame of consecutive frames;

assigning, using a first neural network, weights to a first portion of the second data based, at least in part, on a second portion of the second data and the first data;

generating context information of the two or more objects based, at least in part, on the first portion of the second data with the weights; and

generating, using a second neural network, one or more sets of trajectories associated with the two or more objects based, at least in part, on the context information.

9 . The method of claim 8 , wherein,

the consecutive frames comprise a first frame, a second frame, and a third frame; and

the first data includes a first data structure and a second data structure, wherein the first data structure indicates similar properties between the two or more objects detected within the first frame and the second frame and the second data structure indicates similar properties between the two or more objects detected within the second frame and the third frame.

10 . The method of claim 8 , wherein the second data comprises data associated with a timestamp.

11 . The method of claim 8 , wherein the first neural network and the second neural network are parts of a transformer neural network.

12 . The method of claim 8 , wherein at least one frame of the consecutive frames are generated based, at least in part, on interpolation of two or more frames of the consecutive frames.

13 . The method of claim 8 , further comprising generating at least one latent value based, at least in part, on the context information.

14 . The method of claim 8 , wherein generating the one or more sets of trajectories further comprises:

generating embedded data based, at least in part, on one or more prior sets of trajectories, second data, and context information; and

generating the one or more sets of trajectories associated with the two or more objects based, at least in part, on the embedded data.

15 . A non-transitory computer-readable medium comprising instructions that, when performed by at least one processor of a computing device, cause the computing device to at least:

obtain a first data structure comprising a plurality of values that indicate that two or more objects detected within consecutive frames are similar;

obtain a second data structure that indicates position of the two or more objects detected within each frame of the consecutive frames;

assign, using a first layer of a neural network, weights to a first portion of the second data structure based, at least in part, on a second portion of the second data structure and the first data structure;

generate context information of the two or more objects based, at least in part, on the first portion of the second data structure and the weights; and

generate, using a second layer of the neural network, one or more sets of trajectories associated with the two or more objects based, at least in part, on the context information.

16 . The non-transitory computer-readable medium of claim 15 , wherein the second data structure is generated based at least in part on map data that indicates an area where the two or more objects are located.

17 . The non-transitory computer-readable medium of claim 15 , wherein at least one frame of the consecutive frames are generated based, at least in part, on interpolation of two or more frames of the consecutive frames.

18 . The non-transitory computer-readable medium of claim 15 , wherein the instructions, when performed by the at least one processor of the computing device, further cause the computing device to at least generate embedded data based, at least in part, on ground truth of trajectories and map data that indicates an area where the two or more objects are located.

19 . The non-transitory computer-readable medium of claim 15 , wherein the instructions, when performed by the at least one processor of the computing device, further cause the computing device to at least generate, using a third layer of the neural network, one or more latent variables based, at least in part, on ground truth of trajectories and the context information.

20 . The non-transitory computer-readable medium of claim 15 , wherein:

the first layer is part of an encoder; and

the second layer is part of a decoder.

21 . A method of calculating trajectories of a plurality of objects which are represented in sequential image frames, the method comprising:

detecting the objects in the sequential image frames;

calculating position information for the detected objects;

generating first data identifying at least one similar property of the detected objects;

assigning, using a first neural network, weights to a first portion of the position information based, at least in part, on a second portion of the position information and the first data;

generating context information of the detected objects based, at least in part, on the first portion of the position information with the weights; and

generating, using a second neural network, one or more sets of trajectories associated with the detected objects based, at least in part, on the context information.

22 . The method of claim 21 , wherein the sequential image frames comprise consecutive image frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 9, 2023
From: WENG, XINSHUO; IVANOVIC, BORIS; PAVONE, MARCO
To: NVIDIA CORPORATION
Reel/Frame 062931/0820 →
Continuity (2)
Provisional Application 63349019 · Jun 3, 2022
Related Publication 20230394823A1 · Dec 7, 2023
References Cited (56)
US 20190103026A1 · Liu · 2019 [cited by examiner]
“NuScenes Prediction Challenge Guidelines,” GitHub, retrieved from https://www.nuscenes.org/prediction?externalData=all&mapData=all&modalities=Any, 2020, 5 pages. [cited by applicant]
“NuScenes Prediction Data Split,” GitHub, retrieved from https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/eval/prediction/splits.py, 2020, 2 pages. [cited by applicant]
Alahi et al., “Social LSTM: Human Trajectory Prediction in Crowded Spaces,” Proceedings of the IEEE Conference an Computer Vision and Pattern Recognition 2016, 11 pages. [cited by applicant]
Benbarka et al., “Score Refinement for Confidence-Based 3D Multi-Object Tracking,” IEEE/RSJ International Conference on Intelligent Robots & Systems, 2021, 8 pages. [cited by applicant]
Bhat et al., “Trajformer: Trajectory Prediction with Local Self-Attentive Contexts for Autonomous Driving,” Machine Learning for Autonomous Driving Workshop at Conference on Neural Information Processing Systems, 2020, … [cited by applicant]
Caesar et al., “nuScenes: A Multimodal Dataset for Autonomous Driving,” Nov. 22, 2019, 16 pages. [cited by applicant]
Cao et al., “Spectral Temporal Graph Neural Network for Multivariate Time-series Forecasting,” IEEE Conference on Robotics and Automation, 2021, 13 pages. [cited by applicant]
Cheng et al., “Exploring Dynamic Context for Multi-path Trajectory Prediction,” IEEE Conference on Robotics and Automation, 2021, 7 pages. [cited by applicant]
Cox et al., “An Efficient Implementation of Reid's Multiple Hypothesis Tracking Algorithm and Its Evaluation for the Purpose of Visual Tracking,” IEEE Transactions on Pattern Analysis & Machine Intelligence, 1996, 13 pa… [cited by applicant]
Cui et al., “LookOut: Diverse Multi-Future Prediction and Planning for Self-Driving,” International Conference on Computer Vision, 2021, 10 pages. [cited by applicant]
Eiffert et al., “Probabilistic Crowd GAN: Multimodal Pedestrian Trajectory Prediction using a Graph Vehicle-Pedestrian Attention Network,” IEEE Robotics and Automation Letters, 2020, 8 pages. [cited by applicant]
Geiger et al., “Are we Ready for Autonomous Driving? The KITTI Vision Benchmark Suite,” CVPR, 2012, 8 pages. [cited by applicant]
Guo et al., “3D Object Detection and Tracking on Streaming Data,” IEEE Conference on Robotics and Automation, 2020, 7 pages. [cited by applicant]
Gupta et al., “Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks,” IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Ivanovic et al., “The Trajectron: Probabilistic Multi-Agent Trajectory Modeling With Dynamic Spatiotemporal Graphs,” IEEE/CVF International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Kim et al., “EagerMOT: 3D Multi-Object Tracking via Sensor Fusion,” IEEE Conference on Robotics and Automation, 2021, 7 pages. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes,” Dec. 27, 2013, 14 pages. [cited by applicant]
Kitani et al., “Activity Forecasting,” European Conference on Computer Vision, 2012, 14 pages. [cited by applicant]
KITTI, “KITTI Tracking Dataset,” retrieved from https://www.cvlibs.net/datasets/kitti/eval_tracking.php, 2021, 14 pages. [cited by applicant]
Kosaraju et al., “Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks,” Conference on Neural Information Processing Systems, 2019, 10 pages. [cited by applicant]
Kuhn, “The Hungarian Method for the Assignment Problem,” Naval Research Logistics Quarterly, 2(1-2): Mar. 1955, 19 pages. [cited by applicant]
Lee et al., “DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents,” IEEE Conference on Computer Vision and Pattern Recognition, 2017, 10 pages. [cited by applicant]
Lerner et al., “Crowds by example,” Computer Graphics Forum, vol. 26, 2007, 10 pages. [cited by applicant]
Liang et al., “PnPNet: End-to-End Perception and Prediction with Tracking in the Loop,” IEEE Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Liu et al., “Multimodal Motion Prediction with Stacked Transformers,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, 10 pages. [cited by applicant]
Long et al., “An Algorithm for Tracking Multiple Targets,” IEEE Transactions on Automatic Control, Dec. 1979, 12 pages. [cited by applicant]
Mohamed et al., “Social-STGCNN: A Social Spatio-Temporal Graph Convolutional Neural Network for Human Trajectory Prediction,” IEEE Conference on Computer Vision and Pattern Recognition, 2020, 9 pages. [cited by applicant]
Pellegrini et al., “You'll Never Walk Alone: Modeling Social Behavior for Multi-target Tracking,” IEEE 12th International Conference on Computer Vision, 2009, 8 pages. [cited by applicant]
Pöschmann et al., “Factor Graph based 3D Multi-Object Tracking in Point Clouds,” IEEE/RSJ International Conference on Intelligent Robots & Systems, 2020, 8 pages. [cited by applicant]
Rhinehart et al., “PRECOG: PREdiction Conditioned On Goals in Visual Multi-Agent Settings,” IEEE/CVF International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Rhinehart et al., “R2P2: A ReparameteRized Pushforward Policy for Diverse, Precise Generative Path Forecasting,” European Conf. on Computer Vision, 2018, 17 pages. [cited by applicant]
Robicquet et al., “Learning Social Etiquette: Human Trajectory Understanding In Crowded Scenes,” European Confernce on Computer Vision, 2016, 17 pages. [cited by applicant]
Salzmann et al., “Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data,” European Conference on Computer Vision, 2020, 17 pages. [cited by applicant]
Scheidegger et al., “Mono-Camera 3D Multi-Object Tracking Using Deep Learning Detections and PMBM Filtering,” IEEE Intelligent Vehicles Symposium, 2018, 8 pages. [cited by applicant]
Shi et al., “PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud,” IEEE Conference on Computer Vision and Pattern Recognition, 2019, 10 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Tang et al., “Multiple futures prediction,” Advances in Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Vaswani et al., “Attention is All You Need,” Dec. 6, 2017, 15 pages. [cited by applicant]
Wang et al., “Joint Object Detection and Multi-Object Tracking with Graph Neural Networks,” IEEE Conference on Robotics and Automation, 2021, 8 pages. [cited by applicant]
Wang et al., “PointTrackNet: An End-to-End Network for 3D Object Detection and Tracking from Point Cloud,” IEEE Conf. on Robotics and Automation, 2020, 7 pages. [cited by applicant]
Weng et al., “3D Multi-Object Tracking: A Baseline and New Evaluation Metrics,” IEEE/RSJ International Conference on Intelligent Robots & Systems, 2020, 9 pages. [cited by applicant]
Weng et al., “GNN3DMOT: Graph Neural Network for 3D Multi-Object Tracking with 2D-3D Multi-Feature Learning,” EEE Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Weng et al., “MTP: Multi-hypothesis Tracking and Prediction for Reduced Error Propagation,” 2021, 7 pages. [cited by applicant]
Weng et al., “PTP: Parallelized Tracking and Prediction with Graph Neural Networks and Diversity Sampling,” IEEE Robotics and Automation Letters, 2021, 8 pages. [cited by applicant]
Yu et al., “Spatio-Temporal Graph Transformer Networks for Pedestrian Trajectory Prediction,” European Conference on Computer Vision, 2020, 17 pages. [cited by applicant]
Yu et al., “Towards Robust Human Trajectory Prediction in Raw Videos,” 2021, 8 pages. [cited by applicant]
Yuan et al., “AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent Forecasting,” IEEE/CVF International Conference on Computer Vision, 2021, 14 pages. [cited by applicant]
Zeng et al., “LaneRCNN: Distributed Representations for Graph-Centric Motion Forecasting,” IEEE/RSJ Int. Conference on Intelligent Robots & Systems, 2021, 14 pages. [cited by applicant]
Vaswani et al., “Attention Is All You Need,” In Advances in Neural Information Processing Systems, Dec. 6, 2017, 15 pages. [cited by applicant]
Zhang et al., “Map-Adaptive Goal-Based Trajectory Prediction,” Conference on Robot Learning, 2020, 13 pages. [cited by applicant]
Zhang et al., “Robust Multi-Modality Multi-Object Tracking,” IEEE International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Zhang et al., “Trajectory Forecasting from Detection with Uncertainty-Aware Motion Encoding,” 2022, 11 pages. [cited by applicant]
Zhu et al., “Class-Balanced Grouping and Sampling for Point Cloud 3D Object Detection,” IEEE Conference on Computer Vision and Pattern Recognition, 2019, 8 pages. [cited by applicant]