IP Library Granted Patent US 12,012,189
Granted Patent B2
US 12,012,189 · App. 17/266,059 · Granted Jun 18, 2024

System and method of operation for remotely operated vehicles leveraging synthetic data to train machine learning models

Inventors: Manuel Alberto Parente Da Silva (Maia, PT); Pedro Miguel Vendas Da Costa (Oporto, PT)
Assignee: Ocean Infinity (Portugal), S.A.
B63G8/001G05B13/0265G05D1/0044B63G2008/005G05D1/0088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,012,189
App. No.
17/266,059
Granted
Jun 18, 2024
Kind
B2
Abstract

The present invention provides systems and methods for leveraging synthetic data to train machine learning models. A synthetic training engine may be used to train machine learning models. The synthetic training engine can automatically annotate real images for valuable tasks, such as object segmentation, depth map estimation, and classifying whether a structure is in an image. The synthetic training engine can also train the machine learning model with synthetic images in such a way that the machine learning model will work on real images. The output of the machine learning model may perform valuable tasks, such as the detection of integrity threats in underwater structures.

Claims (52)

1. A system for operating a remotely operated vehicle (ROV) leveraging synthetic data to train a machine learning model and to display classification labels on a display of a navigation interface, the system comprising:

a synthetic training engine comprising:

a video dataset including at least one of: video data or real images coming from the ROV;

a telemetry dataset including telemetry data coming from the ROV;

a 3D model dataset including 3D model of a scene where the ROV is configured to operate;

a synthetic dataset comprising synthetic images generated from different views of objects in the video data or different views of the 3D model of the scene, and associated training labels, the synthetic dataset providing additional data for the video dataset and the real images; and

a machine learning model configured to determine classification labels for the objects shown in the video data or the real images, the machine learning model trained using data comprising the synthetic dataset; and

a navigation interface configured to:

display an object within an environment of the ROV; and

annotate the displayed object using a corresponding classification label.

2. The system of claim 1 , wherein the synthetic training engine is operable to automatically annotate a real image from the real images for object segmentation, depth map estimation, and classifying whether a specific structure is in the real image.

3. A method of leveraging synthetic data to train a machine learning model for operating a remotely operated vehicles (ROV) the method comprising:

obtaining a video dataset including at least one of: video data or real images coming from the (ROV);

obtaining a 3D model dataset including 3D model of a scene where the ROV is configured to operate;

generating a synthetic dataset comprising synthetic images generated from different views of objects in the video data or different views of the 3D model of the scene, and associated training labels, the synthetic dataset providing additional data for the video dataset and the real images;

training a machine learning model using data comprising the synthetic dataset, the machine learning model configured to determine classification labels for the objects shown in the video data or the real images;

displaying, using a navigation interface, an object within an environment of the ROV; and

displaying, using the navigation interface, an annotation for the displayed object, the annotation corresponding to a classification label for the displayed object.

4. The system of claim 1 , wherein the synthetic training engine is configured to map the real images and the synthetic images to a common feature space, wherein the common feature space comprises image features.

5. The system of claim 4 , wherein the synthetic training engine further comprises:

a real image feature extraction model configured to extract real image features from a real image from the real images; and

a synthetic image feature extraction model configured to extract synthetic image features from a synthetic image from the synthetic images; and

wherein, the real image feature extraction model and the synthetic image feature extraction model are trained such that an Euclidean norm (L2) norm of a difference between the real image features and the synthetic image features is minimized.

6. The system of claim 5 , wherein the real image feature extraction model is a first convolutional neural network (CNN), and the synthetic image feature extraction model is a second CNN.

7. The system of claim 6 , wherein the machine learning model is configured to receive one of the extracted real image features or the extracted synthetic image features.

8. The system of claim 7 , wherein the machine learning model is a third CNN.

9. The system of claim 5 , wherein the machine learning model is configured to:

receive an input being a pair of sets of image features corresponding to an object, the pair comprising:

the extracted real image features for a real image from the real images; and

the extracted synthetic image features for a synthetic image from the synthetic images, the synthetic image corresponding to the real image; and

output at least one predicted classification label corresponding to the object.

10. The system of claim 9 , wherein the machine learning model is configured to output a first predicted classification label corresponding to the extracted real image features, and a second predicted classification label corresponding to the extracted synthetic image features, and wherein the machine learning model is trained to minimize a sum of two L2 norms, a first L2 norm being an L2 norm of a difference between the first predicted classification label and a corresponding training label from the training labels, and a second L2 norm being a norm of a difference between the second predicted classification label and the corresponding training label.

11. The system of claim 10 , wherein the real image feature extraction model and the synthetic image feature extraction model are trained jointly with the machine learning model.

12. The system of claim 1 , wherein the synthetic training engine further comprises a convolutional neural network configured to extract image features from one of a real image of the real images or a synthetic image of the synthetic images, and wherein the machine learning model is configured to output a predicted classification label corresponding to the object based on an input comprising the extracted image features.

13. The system of claim 12 , wherein the machine learning model is trained to minimize an L2 norm being the L2 norm of a difference between the predicted classification label and a corresponding training label from the training labels.

14. The system of claim 1 , wherein the machine learning model can be configured to receive one of:

a pair comprising a real image from the real images and a synthetic image corresponding to the real image from the synthetic images;

a real image from the real images; or

a synthetic image from the synthetic images.

15. The system of claim 1 , wherein the synthetic training engine is configured to replay a mission, and wherein the mission is replayed by retrieving ROV telemetry from the telemetry dataset and 3D model data from the 3D model dataset, denoising the telemetry data, and generating a synthetic video of the mission, the synthetic video including the classification labels for objects shown in the video.

16. The method of claim 3 , further comprising:

extracting real image features from a real image from the real images, using a real image feature extraction model; and

extracting synthetic image features from a synthetic image from the synthetic images, using a synthetic image feature extraction model; and

wherein, the real image feature extraction model and the synthetic image feature extraction model are trained such that an Euclidean norm (L2) norm of a difference between the real image features and the synthetic image features is minimized.

17. The method of claim 16 , wherein the real image feature extraction model is a first convolutional neural network (CNN), and the synthetic image feature extraction model is a second CNN.

18. The method of claim 17 , wherein the machine learning model is configured to receive one of the extracted real image features or the extracted synthetic image features.

19. The method of claim 17 , wherein the machine learning model is a third CNN.

20. The method of claim 17 , wherein the machine learning model is configured to:

receive an input being a pair of sets of image features corresponding to an object, the pair comprising:

the extracted real image features for a real image from the real images; and

the extracted synthetic image features for a synthetic image from the synthetic images, the synthetic image corresponding to the real image; and

output at least one predicted classification label corresponding to the object.

Assignments (3)
CHANGE OF NAME Recorded May 18, 2023
From: ABYSSAL, S.A.
To: OCEAN INFINITY (PORTUGAL), S.A.
Reel/Frame 063693/0762 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE'S INTERNAL ADDRESS FROM SALA A10, LEҫA DA PALMEIRA TO SALA A10, LECA DA PALMEIRA PREVIOUSLY RECORDED ON REEL 055153 FRAME 0892. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 8, 2021
From: PARENTE DA SILVA, MANUEL ALBERTO; VENDAS DA COSTA, PEDRO MIGUEL
To: ABYSSAL S.A.
Reel/Frame 055255/0686 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2021
From: PARENTE DA SILVA, MANUEL ALBERTO; VENDAS DA COSTA, PEDRO MIGUEL
To: ABYSSAL S.A.
Reel/Frame 055153/0892 →
Continuity (1)
Related Publication 20210309331A1 · Oct 7, 2021
Cited By (1)
US 12,717,334