IP Library › Granted Patent US 12,437,203
Granted Patent B2
US 12,437,203 · App. 17/179,214 · Granted Oct 7, 2025

Apparatus, system and method for translating sensor label data between sensor domains

Inventors: Nasim Souly (San Mateo, CA); Pratik Prabhanjan Brahma (Belmont, CA)
Assignee: Volkswagen Aktiengesellschaft
G06N3/088G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,203
App. No.
17/179,214
Granted
Oct 7, 2025
Kind
B2
Abstract

Technologies and techniques for converting sensor data, used in a vehicle or other device. A machine-learning model is applied to first sensor data, including a first operational characteristic capability and first sensor label data, wherein the machine-learning model is trained to second sensor data including a second operational characteristic capability. New sensor data is generated that corresponds to the applied machine-learning model, wherein the new sensor data includes translated first sensor label data. A loss function may be applied to the new sensor data to determine the accuracy of the new sensor data and translated first sensor label data. In some examples, a multi-dimensional matrix of camera sensor parameters may be applied to the first sensor data labels to transform the first sensor data labels to second sensor data labels.

Claims (37)

1. A method of converting sensor data for autonomous vehicle perception, comprising:

receiving, by a processing system of an autonomous vehicle, first sensor data from a first sensor, wherein the first sensor data comprises a first operational characteristic capability and first sensor label data comprising at least one of semantic segmentation labels or bounding box masks;

applying, by the processing system, a machine-learning model comprising an encoder-decoder network to the first sensor data, wherein the machine-learning model is trained to translate the first sensor data to emulate second sensor data comprising a second operational characteristic capability different from the first operational characteristic capability;

generating, by the processing system, new sensor data corresponding to the applied machine-learning model, wherein the new sensor data comprises translated first sensor label data emulating the second operational characteristic capability; and

applying, by the processing system, a loss function to the new sensor data to validate the accuracy of the new sensor data and translated first sensor label data; and

configuring the processing system of the autonomous vehicle to perform object detection using the validated new sensor data and translated first sensor label data.

2. The method of claim 1 , wherein applying the machine-learning model comprises applying a deep neural network (DNN).

3. The method of claim 2 , wherein applying the DNN comprises processing the first sensor data via an encoder-decoder network and a generative adversarial network.

4. The method of claim 1 , wherein applying the loss function to the new sensor data comprises applying a consistency loss function to the translated first sensor label data relative to the first sensor label data.

5. The method of claim 1 , wherein applying the loss function to the new sensor data comprises

applying a second machine-learning model to the new sensor data, wherein the second machine-learning model is trained to the first sensor label data to produce modified new sensor label data, and wherein the second machine-learning model comprises a deep convolutional neural network (DNN).

6. The method of claim 5 , wherein the DNN comprises an encoder-decoder network and a generative adversarial network, configured to process the modified new sensor data in the opposite direction of the first machine learning model.

7. The method of claim 1 , wherein the machine-learning model is trained to translate the first sensor data to emulate second sensor data in an unpaired data environment.

8. A method of converting sensor label data for autonomous vehicle perception, comprising:

receiving, by a processing system of an autonomous vehicle, first sensor data from a first sensor, wherein the first sensor data comprises the first operational characteristic capability and first sensor label data comprising at least one of semantic segmentation labels or bounding box masks;

receiving, by the processing system, second sensor data from a second sensor, wherein the second sensor data comprises a second operational characteristic capability different from the first operational characteristic capability;

applying, by the processing system, a machine-learning model comprising an encoder-decoder network to the first sensor data and second sensor data, wherein the machine-learning model is trained to infer first sensor label data to the second sensor data;

generating, by the processing system, new sensor data corresponding to the applied machine-learning model, wherein the new sensor data comprises the inferred first sensor label data emulating the second operational characteristic capability; and

configuring the processing system of the autonomous vehicle to perform object detection using the new sensor data and inferred first sensor label data.

9. The method of claim 8 , further comprising applying a loss function to the new sensor data to determine the accuracy of the inferred first sensor label data relative to the first sensor label data.

10. The method of claim 9 , wherein the loss function comprises a consistency loss function.

11. The method of claim 8 , wherein applying the machine-learning model comprises applying a deep neural network to the first sensor data, second sensor data, and first sensor label data.

12. The method of claim 8 , wherein applying the machine-learning model to infer first sensor label data comprises applying extrinsic parameters of camera matrix data.

13. The method of claim 12 , wherein applying the extrinsic parameters comprises transferring sensor labels to the new sensor data.

14. The method of claim 8 , wherein the second sensor data is translated from the first sensor data via a generative adversarial network.

15. A method for translating sensor data for autonomous vehicle perception, comprising:

receiving, by a processing system of an autonomous vehicle, first sensor data from a first sensor, wherein the first sensor data comprises a first operational characteristic capability and first sensor label data comprising at least one of semantic segmentation labels or bounding box masks;

receiving, by the processing system, second sensor data from a second sensor, wherein the second sensor data comprises a second operational characteristic capability different from the first operational characteristic capability; and

applying, by the processing system, a machine-learning model comprising an encoder-decoder network to the first sensor data and second sensor data in an unpaired data environment, wherein the machine-learning model is trained to translate the first sensor label data to emulate the second operational characteristic capability;

generating, by the processing system, new sensor data corresponding to the applied machine-learning model, wherein the new sensor data comprises translated first sensor label data emulating the second operational characteristic capability;

applying, by the processing system, a loss function to the new sensor data to validate the accuracy of the new sensor data and translated first sensor label data; and

configuring the processing system of the autonomous vehicle to perform object detection using the validated new sensor data and translated first sensor label data.

16. The method of claim 15 , wherein the camera sensor parameters comprise intrinsic and extrinsic coefficients representing three-dimensional world points and their corresponding two-dimensional image points.

17. The method of claim 16 , wherein the extrinsic parameters represent a location of a first and second sensor in a three-dimensional scene, and the intrinsic parameters represent an optical center and focal length of the first and second sensor.

18. The method of claim 15 , wherein translating the first sensor data to the second sensor data comprises applying a color mapping transformation to adjust coloration differences between the first sensor data and the second sensor data.

19. The method of claim 15 , wherein translating the first sensor data to the second sensor data comprises adjusting for differences in temporal sampling intervals between the first sensor and the second sensor.

20. The method of claim 15 , wherein the machine-learning algorithm comprises a reinforcement learning model configured to refine the accuracy of the transformed sensor data and associated label data using feedback derived from a validation process.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2021
From: VOLKSWAGEN GROUP OF AMERICA, INC.
To: VOLKSWAGEN AKTIENGESELLSCHAFT
Reel/Frame 055430/0618 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2021
From: SOULY, NASIM; BRAHMA, PRATIK PRABHANJAN
To: VOLKSWAGEN GROUP OF AMERICA, INC.
Reel/Frame 055325/0627 →
Continuity (1)
Related Publication 20220261658A1 · Aug 18, 2022
References Cited (31)
US 11537139B2 · Rankawat · 2022 [cited by examiner]
US 11551094B2 · Senn · 2023 [cited by examiner]
US 11693417B2 · George · 2023 [cited by examiner]
US 12073329B2 · Kapoor · 2024 [cited by examiner]
US 20190004534A1 · Huang et al. · 2019 [cited by applicant]
US 20190004535A1 · Huang · 2019 [cited by examiner]
US 20190147331A1 · Arditi · 2019 [cited by examiner]
US 20190197667A1 · Paluri · 2019 [cited by examiner]
US 20190228236A1 · Schlicht · 2019 [cited by examiner]
US 20200364572A1 · Senn · 2020 [cited by examiner]
US 20210286923A1 · Kristensen · 2021 [cited by examiner]
US 20220044118A1 · Hüger · 2022 [cited by examiner]
Yu et al. “Pu-net: Point cloud upsampling network.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) pp. 2790-2799 (2018). [cited by applicant]
Li et al. “PU-GAN: a Point Cloud Upsampling Adversarial Network.” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) pp. 7203-7212 (2019). [cited by applicant]
Qiu et al. “Deeplidar: Deep surface normal guided depth prediction for outdoor scene from sparse lidar data and single color image.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP… [cited by applicant]
Chen et al. “Learning joint 2D-3D representations for depth completion.” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) pp. 10023-10032 (2019). [cited by applicant]
Wu et al. “Point Cloud Super Resolution with Adversarial Residual Graph Networks.” arXiv preprint arXiv:1908.02111 (2019). [cited by applicant]
Isola et al. “Image-to-image translation with conditional adversarial networks.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) pp. 1125-1134 (2017). [cited by applicant]
Liu et al. “Unsupervised image-to-image translation networks.” 31st Conf. on Neural Information Processing Systems (NIPS), Long Beach, CA (2017). [cited by applicant]
Cherian et al. “Sem-GAN: semantically-consistent image-to-image translation.” IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2019, pp. 1797-1806, doi: 10.1109/WACV.2019.00196 (2019). [cited by applicant]
Pandey. “An information theoretic framework for camera and lidar sensor data fusion and its applications in autonomous navigation of vehicles.” Thesis, A dissertation submitted in partial fulfillment of the requirements… [cited by applicant]
Dou et al. “Asymmetric cyclegan for unpaired NIR-to-RGB face image translation.” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK pp. 1757-1761 (2019). [cited by applicant]
Kim et al. “Color Image Generation from LiDAR Reflection Data by Using Selected Connection UNET.” Sensors 20, 3387 (2020). [cited by applicant]
Kaji et al. “Overview of image-to-image translation using deep neural networks: denoising, super-resolution, modality-conversion, and reconstruction in medical imaging.” researchgate.net (May 21, 2019). [cited by applicant]
Vaswani et al. “Attention is all you need.” 31st Conference on Neural Information Processing Systems (NIPS), Long Beach, CA (2017). [cited by applicant]
Kim et al. “Asymmetric Encoder-Decoder Structured FCN Based LiDAR to Color Image Generation.” Sensors 2019, 19, 4818. [cited by applicant]
Zhou et al. “Graph Neural Networks:A Review of Methods and Applications.” arXiv.org > cs > arXiv:1812.08434 (Jul. 10, 2019). [cited by applicant]
Coors et al. “NoVA: Learning to see in Novel View Points and Domains.” 2019 International Conf. on 3D Vision (3DV), pp. 116-125, doi: 10.1109/3DV.2019.00022. [cited by applicant]
Bujwid et al. “GANtruth—an unpaired image-to-image translation method for driving scenarios.” 32nd Conference on Neural Information Processing Systems (NeurIPS), Machine Learning for Intelligent Transportation Systems W… [cited by applicant]
Tulsiani et al. “Layer-structured 3D Scene Inference via View Synthesis.” Proceedings of the European Conference on Computer Vision (ECCV), pp. 302-317 (2018). [cited by applicant]
PCT/EP2022/051978. International Search Report & Written Opinion (Jun. 1, 2022). [cited by applicant]