IP Library Granted Patent US 12,726,799
Granted Patent B2
US 12,726,799 · App. 18/886,553 · Granted Sep 1, 2026

Systems and methods for mitigating vehicle pose error across an aggregated feature map

Inventors: Nicholas Baskar Vadivelu (Markham, CA); Mengye Ren (Toronto, CA); Xuanyuan Tu (Milton, CA); Raquel Urtasun (Toronto, CA); Jingkang Wang (Toronto, CA)
Assignee: AURORA OPERATIONS, INC.
H04W4/46B60W60/00272B60W60/00276G05D1/0088G05D1/0274G05D1/246G05D1/81G06F18/2415G06N20/00G06V20/56G06N3/045G06N3/09
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,726,799
App. No.
18/886,553
Granted
Sep 1, 2026
Kind
B2
Abstract

Systems and methods for improved vehicle-to-vehicle communications are provided. A system can obtain sensor data depicting its surrounding environment and input the sensor data (or processed sensor data) to a machine-learned model to perceive its surrounding environment based on its location within the environment. The machine-learned model can generate an intermediate environmental representation that encodes features within the surrounding environment. The system can receive a number of different intermediate environmental representations and corresponding locations from various other systems, aggregate the representations based on the corresponding locations, and perceive its surrounding environment based on the aggregated representations. The system can determine relative poses between the each of the systems and an absolute pose for each system based on the representations. Each representation can be aggregated based on the relative or absolute poses of each system and weighted according to an estimated accuracy of the location corresponding to the representation.

Claims (75)

1 . A computer-implemented method, the method comprising:

generating, based at least in part on first data that describes outputs from a first sensor of a first vehicle, one or more first representations of a first portion of an environment of the first vehicle;

receiving one or more second representations of a second portion of the environment that overlaps the first portion of the environment, the one or more second representations generated using second data that describes outputs from a second sensor of a second vehicle;

generating, by a first model trained to regress relative poses based on input representations, and based at least in part on the one or more first representations and the one or more second representations, a first model output;

generating, based at least in part on the first model output, a first pose, wherein the first pose describes a first relative pose for the first vehicle and the second vehicle;

generating, based at least in part on the first pose, a second pose that describes an absolute pose for the first vehicle;

generating, based at least in part on the first pose, a third pose that describes an absolute pose for the second vehicle;

generating a fourth pose based at least in part on the first pose, the second pose, and the third pose, wherein the fourth pose describes a second relative pose for the first vehicle and the second vehicle;

generating one or more third representations of a third portion of the environment based at least in part on the one or more first representations, the one or more second representations, and the fourth pose;

generating, by a second model trained to generate outputs that predict attributes of objects based on input representations that describe environments, and based at least in part on the one or more third representations, a second model output;

generating, based at least in part on the second model output, third data that describes an attribute for an object in the environment; and

controlling the first vehicle based at least in part on the third data.

2 . The computer-implemented method of claim 1 , comprising:

receiving the one or more second representations from the second vehicle.

3 . The computer-implemented method of claim 1 , comprising:

warping at least a portion of the one or more second representations based at least in part on the fourth pose into a warped representation; and

generating the one or more third representations based at least in part on the warped representation.

4 . The computer-implemented method of claim 3 , comprising:

generating, by a third model trained to aggregate representations of environments, a third model output;

aggregating, based at least in part on the third model output, the one or more first representations and the warped representation.

5 . The computer-implemented method of claim 1 , wherein the third portion of the environment comprises at least a portion of the second portion of the environment that is not included in the first portion of the environment.

6 . The computer-implemented method of claim 5 , wherein the third portion of the environment corresponds to a union of the first portion of the environment and the second portion of the environment.

7 . The computer-implemented method of claim 1 , wherein:

at least one of the one or more first representations is a first feature map encoded with a first plurality of encoded features representative of the first portion of the environment; and

at least one of the one or more second representations is a second feature map encoded with a second plurality of encoded features representative of the second portion of the environment.

8 . The computer-implemented method of claim 1 , wherein the attribute corresponds to a bounding box for an object in the environment.

9 . A computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that store instructions that are executable by the one or more processors to cause the computing system to perform operations, the operations comprising:

generating, based at least in part on first data that describes outputs from a first sensor of a first vehicle, one or more first representations of a first portion of an environment of the first vehicle;

receiving one or more second representations of a second portion of the environment that overlaps the first portion of the environment, the one or more second representations generated using second data that describes outputs from a second sensor of a second vehicle;

generating, by a first model trained to regress relative poses based on input representations, and based at least in part on the one or more first representations and the one or more second representations, a first model output;

generating, based at least in part on the first model output, a first pose, wherein the first pose describes a first relative pose for the first vehicle and the second vehicle;

generating, based at least in part on the first pose, a second pose that describes an absolute pose for the first vehicle;

generating, based at least in part on the first pose, a third pose that describes an absolute pose for the second vehicle;

generating a fourth pose based at least in part on the first pose, the second pose, and the third pose, wherein the fourth pose describes a second relative pose for the first vehicle and the second vehicle;

generating one or more third representations of a third portion of the environment based at least in part on the one or more first representations, the one or more second representations, and the fourth pose;

generating, by a second model trained to generate outputs that predict attributes of objects based on input representations that describe environments, and based at least in part on the one or more third representations, a second model output;

generating, based at least in part on the second model output, third data that describes an attribute for an object in the environment; and

controlling the first vehicle based at least in part on the third data.

10 . The computing system of claim 9 , wherein the computing system is onboard the first vehicle.

11 . The computing system of claim 9 , the operations comprising:

receiving the one or more second representations from the second vehicle.

12 . The computing system of claim 9 , the operations comprising:

warping at least a portion of the one or more second representations based at least in part on the fourth pose into a warped representation; and

generating the one or more third representations based at least in part on the warped representation.

13 . The computing system of claim 12 , the operations comprising:

generating, by a third model trained to aggregate representations of environments, a third model output;

aggregating, based at least in part on the third model output, the one or more first representations and the warped representation.

14 . The computing system of claim 9 , wherein the third portion of the environment comprises at least a portion of the second portion of the environment that is not included in the first portion of the environment.

15 . The computing system of claim 9 , wherein:

at least one of the one or more first representations is a first feature map encoded with a first plurality of encoded features representative of the first portion of the environment; and

at least one of the one or more second representations is a second feature map encoded with a second plurality of encoded features representative of the second portion of the environment.

16 . The computing system of claim 9 , wherein the attribute corresponds to a bounding box for an object in the environment.

17 . An autonomous vehicle comprising:

one or more processors; and

one or more non-transitory computer-readable media that store instructions that are executable by the one or more processors to cause the autonomous vehicle to perform operations, the operations comprising:

generating, based at least in part on first data that describes outputs from a first sensor of the autonomous vehicle, one or more first representations of a first portion of an environment of the autonomous vehicle;

receiving one or more second representations of a second portion of the environment that overlaps the first portion of the environment, the one or more second representations generated using second data that describes outputs from a second sensor of a second vehicle;

generating, by a first model trained to regress relative poses based on input representations, and based at least in part on the one or more first representations and the one or more second representations, a first model output;

generating, based at least in part on the first model output, a first pose, wherein the first pose describes a first relative pose for the autonomous vehicle and the second vehicle;

generating, based at least in part on the first pose, a second pose that describes an absolute pose for the autonomous vehicle;

generating, based at least in part on the first pose, a third pose that describes an absolute pose for the second vehicle;

generating a fourth pose based at least in part on the first pose, the second pose, and the third pose, wherein the fourth pose describes a second relative pose for the autonomous vehicle and the second vehicle;

generating one or more third representations of a third portion of the environment based at least in part on the one or more first representations, the one or more second representations, and the fourth pose;

generating, by a second model trained to generate outputs that predict attributes of objects based on input representations that describe environments, and based at least in part on the one or more third representations, a second model output;

generating, based at least in part on the second model output, third data that describes an attribute for an object in the environment; and

controlling the autonomous vehicle based at least in part on the third data.

18 . The autonomous vehicle of claim 17 , the operations comprising:

receiving the one or more second representations from the second vehicle.

19 . The autonomous vehicle of claim 17 , wherein the third portion of the environment comprises at least a portion of the second portion of the environment that is not included in the first portion of the environment.

20 . The autonomous vehicle of claim 17 , the operations comprising:

warping at least a portion of the one or more second representations based at least in part on the fourth pose into a warped representation;

generating, by a third model trained to aggregate representations of environments, a third model output; and

aggregating, based at least in part on the third model output, the one or more first representations and the warped representation into the one or more third representations.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2024
From: VADIVELU, NICHOLAS BASKAR; REN, MENGYE; WANG, JINGKANG
To: UBER TECHNOLOGIES, INC.
Reel/Frame 069081/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2024
From: TU, XUANYUAN
To: UBER TECHNOLOGIES, INC.
Reel/Frame 069081/0642 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2024
From: URTASUN, RAQUEL
To: UATC, LLC
Reel/Frame 069081/0661 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2024
From: UBER TECHNOLOGIES, INC.
To: UATC, LLC
Reel/Frame 069293/0709 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 069293/0765 →
Continuity (4)
Continuation 17150998 · Jan 15, 2021
Provisional Application 63132792 · Dec 31, 2020
Provisional Application 63058040 · Jul 29, 2020
Related Publication 20250016534A1 · Jan 9, 2025
References Cited (49)
US 10282861B2 · Kwant · 2019 [cited by applicant]
US 20150228077A1 · Menashe · 2015 [cited by applicant]
US 20190095716A1 · Shrestha · 2019 [cited by examiner]
US 20190171871A1 · Zhang · 2019 [cited by examiner]
US 20190362157A1 · Cambias · 2019 [cited by applicant]
US 20200217972A1 · Kim · 2020 [cited by examiner]
DE 102018105293 · 2018 [cited by applicant]
DE 102018105293A1 · 2018 [cited by examiner]
Agrawal et al., “Learning to See by Moving”, International Conference on Computer Vision, Dec. 11, 2015-Dec. 18, 2015, Santiago, Chile, pp. 37-45. [cited by applicant]
Arrigoni et al., “Robust synchronization in SO(3) and SE(3) via low-rank and sparse matrix decomposition”, Computer Vision and Image Understanding, vol. 174, 2018, pp. 95-113. [cited by applicant]
Arrigoni et al., “Spectral Synchronization of Multiple Views in SE(3)”, SIAM Journal of Imaging Sciences, vol. 9, No. 4, Nov. 2016, pp. 1963-1990. [cited by applicant]
Balachandar et al., “Collaboration of AI Agents via Cooperative Multi-Agent Deep Reinforcement Learning”, arXiv:1907.00327v1, Jun. 30, 2019, 9 pages. [cited by applicant]
Bernard et al., “A Solution for Multi-Alignment by Transformation Synchronisation”, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 7-12, 2015, Boston, MA, pp. 2161-2169. [cited by applicant]
Besag, “On the Statistical Analysis of Dirty Pictures”, Journal of the Royal Statistical Society. Series B (Methodological), vol. 48, No. 3, 1986, pp. 259-279. [cited by applicant]
Birdal et al., “Bayesian Pose Graph Optimization via Bingham Distributions and Tempered Geodesic MCMC”, Conference on Neural Information Processing Systems, Dec. 3-8, 2018, Montreal, Canada, 12 pages. [cited by applicant]
Chen et al., “Cooper: Cooperative Perception for Connected Autonomous Vehicles based on 3D Point Clouds”, International Conference on Distributed Computing Systems, Jul. 7-10, 2019, Dallas, TX, pp. 514-524. [cited by applicant]
Choi et al., “A Large Dataset of Object Scans”, arXiv:1602.02481v3, May 5, 2016, 7 pages. [cited by applicant]
Dornhege et al., “Visual Odometry for Tracked Vehicles”, 2008 IEEE International Workshop on Safety, Security & Rescue Robotics (SSRR), Oct. 21-24, 2008, Sendai, Japan, 6 pages. [cited by applicant]
Fitzgibbon, “Robust registration of 2D and 3D point sets”, Image and Vision Computing, vol. 21, 2003, pp. 1145-1153. [cited by applicant]
Gojcic et al., “Learning Multiview 3D point cloud registration”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 1759-1769. [cited by applicant]
Huang et al., “Learning Transformation Synchronization”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, CA, pp. 8082-8091. [cited by applicant]
Kaess et al., “Flow Separation for Fast and Robust Stereo Odometry”, IEEE International Conference on Robotics and Automation, May 12-17, 2009, Kobe, Japan, pp. 3539-3544. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization”, arXiv:1412.6980v9, Jan. 30, 2017, 15 pages. [cited by applicant]
Li et al., “Gated Graph Sequence Neural Networks”, arXiv:1511.05493v4, Sep. 22, 2017, 20 pages. [cited by applicant]
Liang et al., “Deep Continuous Fusion for Multi-Sensor 3D Object Detection”, European Conference on Computer Vision, Sep. 8-14, 2018, Munich, Germany, 16 pages. [cited by applicant]
Liu et al., “ML Estimation of the t Distribution Using EM and its Extensions, ECM and ECME”, Statistica Sinica, vol. 5, 1995, pp. 19-39. [cited by applicant]
Lu et al., “DeepICP: An End-to-End Deep Neural Network for 3D Point Cloud Registration”, arXiv:1905.04153v2, Sep. 16, 2019, 10 pages. [cited by applicant]
Luo et al., “Efficient Deep Learning for Stereo Matching”, Conference on Computer Vision and Pattern Recognition, Jun. 26-Jul. 1, 2016, Las Vegas, NV, pp. 5695-5703. [cited by applicant]
Manivasagam et al., “LiDARsim: Realistic LiDAR Simulation by Leveraging the Real World”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 11167-11176. [cited by applicant]
Matthies et al., “Error Modeling in Stereo Navigation”, Autonomous Robot Vehicles, Springer, 1990, 12 pages. [cited by applicant]
Mohanty et al., “DeepVO: A Deep Learning approach for Monocular Visual Odometry”, arXiv:1611.06069v1, Nov. 18, 2016, 9 pages. [cited by applicant]
Obst et al., “Multi-Sensor Data Fusion for Checking Plausibility of V2V Communications by Vision-based Multiple-Object Tracking” IEEE Vehicular Networking Conference, Dec. 3-5, 2014, Paderborn, Germany, pp. 143-150. [cited by applicant]
Omidshafiei et al., “Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability”, International Conference on Machine Learning, Aug. 6-11, 2017, Sydney, Australia, 10 pages. [cited by applicant]
Pomerleau et al., “A Review of Point Cloud Registration Algorithms for Mobile Robotics”, Foundations and Trends in Robotics, vol. 4, No. 1, 2013, 107 pages. [cited by applicant]
Purkait et al., “NeuRoRA: Neural Robust Rotation Averaging”, arXiv:1912.04485v1, Dec. 10, 2019, 10 pages. [cited by applicant]
Rauch et al., “Car2X-Based Perception in a High-Level Fusion Architecture for Cooperative Perception Systems”, Intelligent Vehicles Symposium, Jun. 3-7, 2012, Alcala de Henares, Spain, pp. 270-275. [cited by applicant]
Rawashdeh et al., “Collaborative Automated Driving: A Machine Learning-Based Method to Enhance the Accuracy of Shared Information”, International Conference on Intelligent Transportation Systems (ITSC), Nov. 4-7, 2018, … [cited by applicant]
Rockl et al., “V2V Communications in Automotive Multi-sensor Multi-target Tracking”, 68th Vehicular Technology Conference, Sep. 21-24, 2008, Calgary, Canada, 5 pages. [cited by applicant]
Rosen et al., “A Certifiably Correct Algorithm for Synchronization over the Special Euclidean Group”, arXiv:1611.00128v3, Feb. 10, 2017, 16 pages. [cited by applicant]
Singer, “Angular synchronization by eigenvectors and semidefinite programming”, Applied and Computational Harmonic Analysis, vol. 30, 2011, pp. 20-36. [cited by applicant]
Smith et al., “Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates”, arXiv:1708.07120v3, May 17, 2018, 18 pages. [cited by applicant]
Sukhbaatar et al., “Learning Multiagent Communication with Backpropagation”, Conference on Neural Information Processing Systems, Dec. 5-10, 2016, Barcelona, Spain, 9 pages. [cited by applicant]
Talukder et al., “Real-time detection of moving objects in a dynamic scene from moving robotic vehicles”, International Conference on Intelligent Robots and Systems, Oct. 27-Nov. 1, 2003, Las Vegas, NV, pp. 1308-1313. [cited by applicant]
Wang et al., “Deep VO: Towards End-to-End Visual Odometry with Deep Recurrent Convolutional Neural Networks”, IEEE International Conference on Robotics and Automation (ICRA), May 29-Jun. 3, 2017, Singapore, pp. 2043-205… [cited by applicant]
Wang et al., “Robust Probabilistic Modeling with Bayesian Data Reweighting”, International Conference on Machine Learning, Aug. 6-11, 2017, Sydney, Australia, 10 pages. [cited by applicant]
Wang et al., “V2VNet: Vehicle-to-Vehicle Communication for Joint Perception and Prediction”, European Conference on Computer Vision, Aug. 23-28, 2020, Virtual, 17 pages. [cited by applicant]
Yang et al., “A Polynomial-time Solution for Robust Registration with Extreme Outlier Rates”, Robotics: Science and Systems, Jun. 22-26, 2016, Freiburg im Breisgau, Germany, 10 pages. [cited by applicant]
Yew et al., “3DFeat-Net: Weakly Supervised Local 3D Features for Point Cloud Registration”, European Conference on Computer Vision, Sep. 8-14, 2018, Munich, Germany, 17 pages. [cited by applicant]
Yousif et al., “An Overview to Visual Odometry and Visual Slam: Applications to Mobile Robotics”, Intelligent Industrial Systems, vol. 1, No. 4, 2015, pp. 289-311. [cited by applicant]