IP Library Granted Patent US 12,367,681
Granted Patent B2
US 12,367,681 · App. 17/448,247 · Granted Jul 22, 2025

Simulating viewpoint transformations for sensor independent scene understanding in autonomous systems

Inventors: Zongyi Yang (Eatontown, NJ); Mariusz Bojarski (Lincroft, NJ); Bernhard Firner (Highland Park, NJ); Urs Muller (Keyport, NJ)
Assignee: NVIDIA Corporation
G06V20/56G05B13/0265G06F18/214G06F18/251G06T3/60G06T5/80G06T7/11G06V10/25G06T2207/20081G06T2207/20132G06T2207/30241G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,681
App. No.
17/448,247
Granted
Jul 22, 2025
Kind
B2
Abstract

In various examples, sensor data used to train an MLM and/or used by the MLM during deployment, may be captured by sensors having different perspectives (e.g., fields of view). The sensor data may be transformed—to generate transformed sensor data—such as by altering or removing lens distortions, shifting, and/or rotating images corresponding to the sensor data to a field of view of a different physical or virtual sensor. As such, the MLM may be trained and/or deployed using sensor data captured from a same or similar field of view. As a result, the MLM may be trained and/or deployed—across any number of different vehicles with cameras and/or other sensors having different perspectives—using sensor data that is of the same perspective as the reference or ideal sensor.

Claims (58)

1. A computer implemented method comprising:

receiving image data generated using a first camera having a first perspective view from a first mounting location of a vehicle associated with the first camera;

applying a transformation to the image data to generate transformed image data, the transformation converting at least the first mounting location of the first perspective view to a second mounting location to simulate a second camera having a second perspective view being mounted to the vehicle at the second mounting location based at least on the second mounting location being used with respect to training images used to train a machine learning model;

applying the transformed image data to the machine learning model trained using the training images corresponding to the second mounting location;

computing, using the machine learning model and based at least on the machine learning model processing the transformed image data, output data representative of one or more predictions corresponding to the transformed image data; and

transmitting the output data to cause the vehicle to perform one or more operations based at least on the one or more predictions.

2. The computer implemented method of claim 1 , wherein the training images were generated using a physical camera mounted to the vehicle or a different vehicle at the second mounting location.

3. The computer implemented method of claim 1 , wherein:

the converting includes converting a first lens distortion captured by the image data and associated with the first camera to a second lens distortion associated with the second camera; and

the applying the transformed image data includes applying the transformed image data having the second lens distortion as an input to the machine learning model.

4. The computer implemented method of claim 1 , wherein:

the converting includes:

determining boundaries that define a region of interest corresponding to the second perspective view in world space; and

extracting the region of interest from one or more images that correspond to the image data; and

the applying the transformed image data includes applying the region of interest extracted from the one or more images as an input to the machine learning model.

5. The computer implemented method of claim 1 , wherein:

the converting includes:

determining boundaries that define a region of interest corresponding to the second perspective view in one or more images that corresponds to the transformed image data;

identifying pixels that correspond to the region of interest using the boundaries; and

applying a viewpoint transformation to the region of interest using the pixels identified using the boundaries to extract the region of interest from the one or more images; and

the applying the transformed image data includes applying the region of interest extracted from the one or more images as an input to the machine learning model.

6. The computer implemented method of claim 1 , wherein the method further includes:

determining, using the machine learning model, one or more second predictions from second transformed image data generated using a third camera having a third perspective view and converted to the second perspective view of the second camera; and

fusing the one or more predictions with the one or more second predictions to determine one or more fused predictions, wherein the one or more operations are based at least on the one or more fused predictions.

7. The computer implemented method of claim 1 , wherein the one or more predictions include one or more trajectory points in world space of a trajectory for the vehicle through an environment.

8. The computer implemented method of claim 1 , wherein the second camera has a same facing direction relative to the vehicle as the first camera and the second camera has a fixed pose.

9. A system comprising:

one or more processing units to execute operations comprising:

identifying image coordinates that define a region of interest within an image generated using a first camera;

transforming the region of interest based at least on converting the first camera to a mounting location to simulate a second camera being mounted to a machine at the mounting location based at least on the mounting location being used in training images used to train a machine learning model;

determining one or more predictions based at least on applying the region of interest that corresponds to the mounting location to the machine learning model; and

transmitting data to cause the machine to perform one or more operations based at least on the one or more predictions.

10. The system of claim 9 , wherein the transforming further includes removing lens distortions in at least the region of interest caused by a lens of the first camera to concert the region of interest to a lens-independent format, and the region of interest that is applied to the machine learning model has the lens-independent format.

11. The system of claim 9 , wherein the transforming further includes converting one or more intrinsics of the first camera to one or more intrinsics of the second camera used to generate the training images, and the region of interest that is applied to the machine learning model has the one or more intrinsics of the second camera.

12. The system of claim 9 , wherein the converting includes:

selecting a set of pixels in the image that correspond to the region of interest based at least on the identifying; and

based at least on the selecting, adjusting a second field of view depicted in the set of pixels to generate an input image of the region of interest depicting the mounting location, wherein the one or more predictions are determined from the input image.

13. The system of claim 9 , wherein the transforming includes applying a viewpoint transform to a subset of the image, the subset representing the region of interest.

14. The system of claim 9 , wherein identifying the region of interest includes determining one or more boundaries of the region of interest using one or more reference lines in an environment depicted in the image.

15. A processor comprising:

one or more circuits to:

receive image data generated using a first camera associated with a machine, the first camera having a first field of view in an environment,

apply a transformation to the image data to generate transformed image data, the transformation converting a first mounting location corresponding to the first field of view to a second field of view that comprises a perspective view to simulate a second camera mounted to the machine at a second mounting location, and

update one or more parameters of a machine learning model such that the machine learning model is trained to generate one or more predictions from the transformed image data capturing the second field of view that comprises the perspective view from the second mounting location.

16. The processor of claim 15 , wherein the second camera corresponds to a physical camera and the processor is further to update the one or more parameters of the machine learning model based at least on applying at least one image generated using the physical camera, mounted at the second mounting location, to the machine learning model.

17. The processor of claim 15 , wherein the simulated second camera has a fixed pose throughout training the machine learning model.

18. The processor of claim 15 , wherein:

the converting includes:

determining boundaries that define a region of interest corresponding to the second field of view in world space; and

extracting the region of interest from one or more images that correspond to the image data;

the updating the one or more parameters includes applying the region of interest extracted from the one or more images as an input to the machine learning model.

19. The processor of claim 15 , wherein:

the converting includes:

determining boundaries that define a region of interest corresponding to the second field of view in one or more images that corresponds to the image data;

identifying pixels that correspond to the region of interest using the boundaries; and

applying a viewpoint transformation to the pixels identified using the boundaries to extract the region of interest from the one or more images; and

the updating the one or more parameters includes applying the region of interest extracted from the one or more images as an input to the machine learning model.

20. The processor of claim 15 , wherein the transformation uses one or more references points corresponding to one or more axles of the machine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2021
From: YANG, ZONGYI; BOJARSKI, MARIUSZ; FIRNER, BERNHARD; MULLER, URS
To: NVIDIA CORPORATION
Reel/Frame 057605/0148 →
Continuity (2)
Provisional Application 63081008 · Sep 21, 2020
Related Publication 20220092317A1 · Mar 24, 2022
References Cited (64)
US 10007269B1 · Gray · 2018 [cited by applicant]
US 10134278B1 · Konrardy et al. · 2018 [cited by applicant]
US 10832418B1 · Karasev · 2020 [cited by examiner]
US 10885698B2 · Muthler et al. · 2021 [cited by applicant]
US 10997433B2 · Xu · 2021 [cited by examiner]
US 11042163B2 · Chen et al. · 2021 [cited by applicant]
US 11107228B1 · Shrivastava · 2021 [cited by examiner]
US 11170524B1 · Mishra · 2021 [cited by examiner]
US 11610115B2 · Kar · 2023 [cited by examiner]
US 11698272B2 · Kroepfl · 2023 [cited by examiner]
US 11804042B1 · Alokhina · 2023 [cited by examiner]
US 11823433B1 · Paz-Perez · 2023 [cited by examiner]
US 20160321074A1 · Hung et al. · 2016 [cited by applicant]
US 20170010108A1 · Shashua · 2017 [cited by applicant]
US 20170259801A1 · Abou-Nasr et al. · 2017 [cited by applicant]
US 20170364083A1 · Yang et al. · 2017 [cited by applicant]
US 20180121273A1 · Fortino et al. · 2018 [cited by applicant]
US 20190071101A1 · Emura et al. · 2019 [cited by applicant]
US 20190310650A1 · Halder · 2019 [cited by applicant]
US 20190384304A1 · Towal et al. · 2019 [cited by applicant]
US 20200082567A1 · Liu · 2020 [cited by examiner]
US 20200257301A1 · Weiser · 2020 [cited by examiner]
US 20220092349A1 · Yang · 2022 [cited by examiner]
US 20230110713A1 · Degirmenci · 2023 [cited by examiner]
US 20230259540A1 · Das · 2023 [cited by examiner]
US 20230274151A1 · Xu · 2023 [cited by examiner]
US 20240077331A1 · Yin · 2024 [cited by examiner]
CN 107121952A · 2017 [cited by applicant]
CN 111373458A · 2020 [cited by applicant]
Lin et al. (“A Vision Based Top-View Transformation Model for a Vehicle Parking Assistant”, Sensors 2012,). (Year: 2012). [cited by examiner]
How can we get the driver view?⋅ Issue #1636 ⋅ carla-simulator/carla, https://github.com/carla-simulator/carla/issues/1636, May 2019 (Year: 2019). [cited by examiner]
Dekel et al. “Moving Camera, Moving People: A Deep Learning Approach to Depth Prediction”, Google Research, https://research.google/blog/moving-camera-moving-people-a-deep-learning-approach-to-depth-prediction/, May 201… [cited by examiner]
International Preliminary Report on Patentability for PCT Application No. PCT/US2021/051286, filed Sep. 21, 2021, mailed Mar. 30, 2023, 8 pgs. [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, National Highway Traffic Safety Administration (NHTSA), A Division of the US Department of Transportation, and the S… [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, National Highway Traffic Safety Administration (NHTSA), A Division of the US Department of Transportation, and the S… [cited by applicant]
Abdi, L., and Meddeb, A., “Driver information system: a combination of augmented reality, deep learning and vehicular Ad-hoc networks”, Multimedia Tools Applications, vol. 77, No. 12, pp. 14763-14703 (2018) , published … [cited by applicant]
Bojarski, M., et al., “End to End Learning for Self-Driving Cars”, arXiv:1604.073v16v1 [cs.CV], pp. 1-9 (Apr. 25, 2016), XP055570062. [cited by applicant]
Lecun, Y., et al., “Off-Road Obstacle Avoidance through End-to-End Learning”, In Advances in Neural Information Processing Systems, pp. 1-8 (2006). [cited by applicant]
Pomerleau, D. A., “ALVINN: An Autonomous Land Vehicle in a Neural Network”, In Advances in Neural Information Processing Systems, pp. 305-313 (Jan. 1989). [cited by applicant]
Reiher, L., et al., “A Sim2Real Deep Learning Approach for the Transformation of Images from Multiple Vehicle-Mounted Cameras to a Semantically Segmented Image in Bird's Eye View”, arXiv:200504078v1 [cs.CV], Cornell Uni… [cited by applicant]
Tian, Y., et al., “Training and Testing Object Detectors With Virtual Images”, IEEE/CAA Journal of Automatica Sinica, Chinese Association of Automation (CAA), vol. 5, No. 2, pp. 539-546 (Mar. 2018), XP011676953. [cited by applicant]
“Methodology of Using a Single Controller (ECU) for a Fault-Tolerant/Fail-Operational Self-Driving System”, U.S. Appl. No. 62/524,283, filed Jun. 23, 2017. [cited by applicant]
“Systems and Methods for Safe and Reliable Autonomous Vehicles”, U.S. Appl. No. 62/584,549, filed Nov. 10, 2017. [cited by applicant]
“System and Method for Controlling Autonomous Vehicles”, U.S. Appl. No. 62/614,466, filed Jan. 7, 2018. [cited by applicant]
“System and Method for Safe Operation of Autonomous Vehicles”, U.S. Appl. No. 62/625,351, filed Feb. 2, 2018. [cited by applicant]
“Conservative Control for Zone Driving of Autonomous Vehicles Using Safe Time of Arrival”, U.S. Appl. No. 62/628,831, filed Feb. 9, 2018. [cited by applicant]
“System and Method for Sharing Camera Data Between Primary and Backup Controllers in Autonomous Vehicle Systems”, U.S. Appl. No. 62/629,822, filed Feb. 13, 2018. [cited by applicant]
“Pruning Convolutional Neural Networks for Autonomous Vehicles and Robotics”, U.S. Appl. No. 62/630,445, filed Feb. 14, 2018. [cited by applicant]
“Methods for accurate real-lime object detection and for determining confidence of object detection suitable for Autonomous vehicles”, U.S. Appl. No. 62/631,781, filed Feb. 18, 2018. [cited by applicant]
“System and Method for Autonomous Shuttles, Robo-Taxis, Ride-Sharing and On-Demand Vehicles”, U.S. Appl. No. 62/635,503, filed Feb. 26, 2018. [cited by applicant]
“Convolutional Neural Networks to Detect Drivable Freespace for Autonomous Vehicles”, U.S. Appl. No. 62/643,665, filed Mar. 15, 2018. [cited by applicant]
“Geometric Shadow Filter for Denoising Ray-Traced Shadows”, U.S. Appl. No. 62/644,385, filed Mar. 17, 2018. [cited by applicant]
“Energy Based Reflection Filter for Denoising Ray-Traced Glossy Reflections”, U.S. Appl. No. 62/644,386, filed Mar. 17, 2018. [cited by applicant]
“Distance Based Ambient Occlusion Filter for Denoising Ambient Occlusions”, U.S. Appl. No. 62/644,601, filed Mar. 19, 2018. [cited by applicant]
“Adaptive Occlusion Sampling of Rectangular Area Lights with Voxel Cone Tracing”, U.S. Appl. No. 62/644,806, filed Mar. 19, 2018. [cited by applicant]
“Deep Neural Network for Estimating Depth from Stereo Using Semi-Supervised Learning”, U.S. Appl. No. 62/646,148, filed Mar. 21, 2018. [cited by applicant]
“Video Prediction Using Spatially Displaced Convolution”, U.S. Appl. No. 62/646,309, filed Mar. 21, 2018. [cited by applicant]
“Video Prediction Using Spatially Displaced Convolution”, U.S. Appl. No. 62/647,545, filed Mar. 23, 2018. [cited by applicant]
“System and Methods for Advanced AI-Assisted Vehicles”, U.S. Appl. No. 62/648,358, filed Mar. 26, 2018. [cited by applicant]
“System and Method for Training, Testing, Verifying, and Validating Autonomous and Semi-Autonomous Vehicles”, U.S. Appl. No. 62/648,399, filed Mar. 27, 2018. [cited by applicant]
“Method and System of Remote Operation of a Vehicle Using an Immersive Virtual Reality Environment”, U.S. Appl. No. 62/648,493, filed Mar. 27, 2018. [cited by applicant]
“System and Methods for Virtualized Intrusion Detection and Prevent System in Autonomous Vehicles”, U.S. Appl. No. 62/682,803, filed Jun. 8, 2018. [cited by applicant]
“Deep Learning for Path Detection in Autonomous Vehicles”, U.S. Appl. No. 62/684,328, filed Jun. 13, 2018. [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US2021/051286, mailed on Dec. 6, 2021, 11 pages. [cited by applicant]
Cited By (1)
US 12,703,373