IP Library › Granted Patent US 12,592,038
Granted Patent B2
US 12,592,038 · App. 18/674,023 · Granted Mar 31, 2026

Editable semantic map with virtual camera for mobile robot learning

Inventors: Cheng Zhao (Cupertino, CA); Yuliang Guo (Redwood City, CA); Ruoyu Wang (Sunnyvale, CA); Xinyu Huang (San Jose, CA); Liu Ren (Saratoga, CA)
Assignee: Robert Bosch GmbH
G06T17/05G06T7/194G06T11/60G06V10/56G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,592,038
App. No.
18/674,023
Granted
Mar 31, 2026
Kind
B2
Abstract

A computer-implemented method and system relate to computer vision. A first semantic map of an environment is three-dimensional (3D). A foreground scene and a background scene are generated individually using the semantic data of the first semantic map. The foreground scene contains foreground components of the first semantic map. The background scene contains background components of the first semantic map. A machine learning model generates an enhanced background view by completing incomplete regions of the background components. Input data is received to modify the background components, the foreground components, or both. A second semantic map is generated in 3D using the enhanced background view, the foreground components, and the input data. The second semantic map is 3D. Virtual camera data is generated using the second semantic map. The virtual camera data includes at least new image data and corresponding new depth data.

Claims (83)

1 . A computer-implemented method comprising:

receiving a first semantic map of an environment, the first semantic map being three-dimensional (3D) and including semantic data;

generating a background scene by filtering out foreground components from the first semantic map using the semantic data, the background scene including background components;

generating a foreground scene by filtering out the background components from the first semantic map using the semantic data, the foreground scene including the foreground components;

generating, via a first machine learning model, an enhanced background view using the background scene, the first machine learning model generating map data for incomplete regions of the background scene, the incomplete regions including at least corresponding parts of the background components occluded by the foreground components in the first semantic map, the map data including image data and depth data;

receiving input data to edit the background components, the foreground components, or both the background components and the foreground components;

generating a second semantic map using at least the enhanced background view and the input data, the second semantic map being 3D and a modified version of the first semantic map with respect to the background components and the foreground components; and

generating virtual camera data using the second semantic map, the virtual camera data including at least new image data and new depth data of the second semantic map.

2 . The computer-implemented method of claim 1 , further comprising:

generating, via a second machine learning model, the map data for the incomplete regions of the foreground scene, the incomplete regions corresponding to each occluded part of each foreground component.

3 . The computer-implemented method of claim 2 , further comprising:

updating first parameters of the first machine learning model using a first truncation loss associated with completing the background scene; and

updating second parameters of the second machine learning model using a second truncation loss associated with completing the foreground scene,

wherein,

the first machine learning model includes a first Neural Radiance Fields (NeRF) model, and

the second machine learning model includes a second NeRF model.

4 . The computer-implemented method of claim 1 , wherein the input data includes at least (i) image data with a particular style or (ii) text data that specifies the particular style.

5 . The computer-implemented method of claim 1 , wherein:

the virtual camera data is captured from a viewpoint of a mobile robot; and

the virtual camera data further includes instance mask data, semantic mask data, bounding box data, and pose data.

6 . The computer-implemented method of claim 5 , further comprising:

generating annotated training data using the virtual camera data; and

training a deep neural network (DNN) using the annotated training data,

wherein the DNN is employed by the mobile robot.

7 . The computer-implemented method of claim 1 , further comprising:

generating, via another machine learning system, a new foreground component using a particular foreground component and the input data,

wherein,

the new foreground component is modified with respect to a selected feature based on the input data, and

the selected feature is color, shape, style, position, or pose.

8 . The computer-implemented method of claim 7 , wherein the another machine learning system includes (i) a first deep neural network (DNN) to predict a signed distance field (SDF) data from a subset of latent embeddings of features of the particular foreground component, (i) a second DNN to predict density data using the SDF data and (iii) a third DNN to predict 2D image data using the SDF data.

9 . A system comprising:

one or more processors;

one or more computer memory in data communication with the one or more processors, wherein the one or more computer memory being non-transitory and having computer readable data stored thereon, the computer readable data including instructions that, when executed by one or more processors, causes the one or more processors to perform a method, the method including

receiving a first semantic map of an environment, the first semantic map being three-dimensional (3D) and including semantic data;

generating a background scene by filtering out foreground components from the first semantic map using the semantic data, the background scene including background components;

generating a foreground scene by filtering out the background components from the first semantic map using the semantic data, the foreground scene including the foreground components;

generating, via a first machine learning model, an enhanced background view using the background scene, the first machine learning model generating map data for incomplete regions of the background scene, the incomplete regions including at least corresponding parts of the background components occluded by the foreground components in the first semantic map, the map data including image data and depth data;

receiving input data to edit the background components, the foreground components, or both the background components and the foreground components;

generating a second semantic map using at least the enhanced background view and the input data, the second semantic map being 3D and a modified version of the first semantic map with respect to the background components and the foreground components; and

generating virtual camera data using the second semantic map, the virtual camera data including at least new image data and new depth data of the second semantic map.

10 . The system of claim 9 , further comprising:

generating, via a second machine learning model, the map data for the incomplete regions of the foreground scene, the incomplete regions corresponding to each occluded part of each foreground component.

11 . The system of claim 10 , further comprising:

updating first parameters of the first machine learning model using a first truncation loss associated with completing the background scene; and

updating second parameters of the second machine learning model using a second truncation loss associated with completing the foreground scene,

wherein,

the first machine learning model includes a first Neural Radiance Fields (NeRF) model, and

the second machine learning model includes a second NeRF model.

12 . The system of claim 9 , wherein:

the virtual camera data is captured from a viewpoint of an mobile robot; and

the virtual camera data further includes instance mask data, semantic mask data, bounding box data, and pose data.

13 . The system of claim 12 , further comprising:

generating annotated training data using the virtual camera data; and

training a deep neural network (DNN) using the annotated training data,

wherein the DNN is employed by the mobile robot.

14 . The system of claim 9 , further comprising:

generating, via another machine learning system, a new foreground component using a particular foreground component and the input data,

wherein,

the new foreground component is modified with respect to a selected feature based on the input data; and

the selected feature is color, shape, style, position, or pose.

15 . The system of claim 14 , wherein the another machine learning system includes (i) a first deep neural network (DNN) to predict a signed distance field (SDF) data from a subset of latent embeddings of features of the particular foreground component, (i) a second DNN to predict density data using the SDF data and (iii) a third DNN to predict 2D image data using the SDF data.

16 . One or more non-transitory computer readable mediums having computer readable data stored thereon, the computer readable data including instructions that, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising:

receiving a first semantic map of an environment, the first semantic map being three-dimensional (3D) and including semantic data;

generating a background scene by filtering out foreground components from the first semantic map using the semantic data, the background scene including background components;

generating a foreground scene by filtering out the background components from the first semantic map using the semantic data, the foreground scene including the foreground components;

generating, via a first machine learning model, an enhanced background view using the background scene, the first machine learning model generating map data for incomplete regions of the background scene, the incomplete regions including at least corresponding parts of the background components occluded by the foreground components in the first semantic map, the map data including image data and depth data;

receiving input data to edit the background components, the foreground components, or both the background components and the foreground components;

generating a second semantic map using at least the enhanced background view and the input data, the second semantic map being 3D and a modified version of the first semantic map with respect to the background components and the foreground components; and

generating virtual camera data using the second semantic map, the virtual camera data including at least new image data and new depth data of the second semantic map.

17 . The one or more non-transitory computer readable mediums of claim 16 , further comprising:

generating, via a second machine learning model, the map data for the incomplete regions of the foreground scene, the incomplete regions corresponding to each occluded part of each foreground component.

18 . The one or more non-transitory computer readable mediums of claim 17 , further comprising:

updating first parameters of the first machine learning model using a first truncation loss associated with completing the background scene; and

updating second parameters of the second machine learning model using a second truncation loss associated with completing the foreground scene,

wherein,

the first machine learning model includes a first Neural Radiance Fields (NeRF) model, and

the second machine learning model includes a second NeRF model.

19 . The one or more non-transitory computer readable mediums of claim 16 , further comprising:

generating, via another machine learning system, a new foreground component using a particular foreground component and the input data,

wherein,

the new foreground component is modified with respect to a selected feature based on the input data; and

the selected feature is color, shape, style, position, or pose.

20 . The one or more non-transitory computer readable mediums of claim 19 , wherein the another machine learning system includes (i) a first deep neural network (DNN) to predict a signed distance field (SDF) data from a subset of latent embeddings of features of the particular foreground component, (i) a second DNN to predict density data using the SDF data and (iii) a third DNN to predict 2D image data using the SDF data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2025
From: ZHAO, CHENG; GUO, YULIANG; HUANG, XINYU; REN, LIU
To: ROBERT BOSCH GMBH
Reel/Frame 070809/0008 →
Continuity (1)
Related Publication 20250363737A1 · Nov 27, 2025
References Cited (17)
US 10891795B2 · Ha · 2021 [cited by examiner]
US 20200082219A1 · Li · 2020 [cited by examiner]
US 20240062645A1 · Dubinsky · 2024 [cited by examiner]
US 20240135612A1 · Hold-Geoffroy · 2024 [cited by examiner]
US 20240144520A1 · Gori · 2024 [cited by examiner]
US 20240144586A1 · Hold-Geoffroy · 2024 [cited by examiner]
US 20240144623A1 · Gori · 2024 [cited by examiner]
US 20240193727A1 · Yu · 2024 [cited by examiner]
US 20240362815A1 · Joachim · 2024 [cited by examiner]
US 20240378832A1 · Joachim · 2024 [cited by examiner]
US 20250225733A1 · Mech · 2025 [cited by examiner]
Kahneman, Daniel. Thinking, Fast and Slow. Macmillan, 2011, https://us.macmillan.com/books/9780374533557/thinkingfastandslow. [cited by applicant]
Dai et al., “BundleFusion: Real-time Globally Consistent 3D Reconstruction using On-the-fly Surface Re-integration,” arXiv:1604.01093v3 [cs.GR], Feb. 7, 2017, pp. 1-19. [cited by applicant]
Rukhovich et al., “FCAF3D: Fully Convolutional Anchor-Free 3D Object Detection,” European Conference on Computer Vision (ECCV), 2022, pp. 1-17. [cited by applicant]
Mildenhall et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” arXiv:2003.08943v2 [cs.CV], Aug. 3, 2020, pp. 1-25. [cited by applicant]
Azinovic et al., “Neural RGB-D Surface Reconstruction,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2002, pp. 1-21. [cited by applicant]
Muller et al., “AutoRF: Learning 3D Object Radiance Fields from Single View Observations,” Proceedings of hte IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, 10 pages. [cited by applicant]