IP Library Granted Patent US 12682629
Granted Patent B2
US 12682629 · App. 18/485,174 · Granted Jul 14, 2026

Device and method for determining an encoder configured image analysis

Inventors: Yumeng Li (Tuebingen, DE); Anna Khoreva (Stuttgart, DE); Dan Zhang (Leonberg, DE)
Assignee: ROBERT BOSCH GMBH
G06V10/82G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682629
App. No.
18/485,174
Granted
Jul 14, 2026
Kind
B2
Abstract

A computer-implemented method for training an encoder. The encoder is configured for determining a latent representation of an image. Training the encoder includes: determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image; masking out parts of the noise image, thereby determining a masked noise image; determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network; training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image.

Claims (66)

1 . A computer-implemented method for training an encoder, wherein the encoder is configured for determining a latent representation of an image, and the training of the encoder comprises the following steps of:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image;

masking out parts of the noise image to determine a masked noise image;

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network; and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image.

2 . The method according to claim 1 , wherein the masking out of parts of the noise image includes replacing values within the parts with randomly drawn values.

3 . The method according to claim 1 , wherein the loss value is determined based on a loss function, wherein a first term of the loss function characterizes the difference between the predicted image and the training image.

4 . The method according to claim 3 , wherein the first term further characterizes a masking of the difference, wherein the masking removes pixels from the difference that fall into the masked-out parts.

5 . The method according to claim 3 , wherein the loss function includes a second term that characterizes a norm of the noise image predicted by the encoder.

6 . The method according to claim 5 , wherein the loss function includes a third term characterizing a negative log likelihood of an output signal of a discriminator, wherein the output signal is determined by the discriminator by providing the predicted image to the discriminator.

7 . The method according to claim 6 , wherein the training image is determined by providing a randomly sampled latent representation or a user defined latent representation to the generator and wherein the loss function includes a fourth term characterizing a difference between the randomly sampled or user defined latent representation and the latent representation determined from the encoder.

8 . The method according to claim 3 , wherein the loss function includes a fifth term characterizing a difference of a first feature representation determined by providing the training image to a feature extractor and a second feature representation determined by providing the predicted image to the feature extractor, wherein the difference does not characterize features characterizing pixels in the masked-out parts.

9 . A computer-implemented method for determining an augmentation of an image, comprising the following steps:

obtaining a trained encoder, the encoder being configured for determining a latent representation of an image, the encoder being trained by:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image,

masking out parts of the noise image to determine a masked noise image,

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network, and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image;

determining a first latent representation and a noise image by providing the image to the trained encoder;

altering the first latent representation to determine a second latent representation; and

determining the augmentation by providing the second latent representation and the noise image used in the training the encoder, as input to the generator.

10 . A computer-implemented method for training a machine learning system, wherein the machine learning system is configured for determining an output signal characterizing a classification and/or regression analysis of an image, wherein the method comprises the following steps:

determining an augmentation of a second training image by:

obtaining a trained encoder, the encoder being configured for determining a latent representation of an image, the encoder being trained by:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image,

masking out parts of the noise image to determine a masked noise image,

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network, and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image;

determining a first latent representation and a noise image by providing the second training image to the trained encoder;

altering the first latent representation to determine a second latent representation;

determining the augmentation by providing the second latent representation and the noise image used in the training the encoder, as input to the generator; and

training the machine learning system based on the augmentation.

11 . A computer-implemented method for determining a control signal of an actuator using a machine learning system, wherein the machine learning system is configured for determining an output signal characterizing a classification and/or regression analysis of an image, wherein the method comprises the following steps:

determining an augmentation of a second training image by:

obtaining a trained encoder, the encoder being configured for determining a latent representation of an image, the encoder being trained by:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image,

masking out parts of the noise image to determine a masked noise image,

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network, and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image;

determining a first latent representation and a noise image by providing the second training image to the trained encoder;

altering the first latent representation to determine a second latent representation;

determining the augmentation by providing the second latent representation and the noise image used in the training the encoder, as input to the generator; and

training the machine learning system based on the augmentation; and

determining a control signal based on an output signal of the trained machine learning system, wherein the output signal is determined based on the image.

12 . A training system for training an encoder, wherein the encoder is configured for determining a latent representation of an image, and wherein the training system is configured to:

determine a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image;

mask out parts of the noise image to determine a masked noise image;

determine a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network; and

train the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image.

13 . A control system configured to determine a control signal of an actuator using a machine learning system, wherein the machine learning system is configured for determining an output signal characterizing a classification and/or regression analysis of an image, wherein the control system is configured to:

determine a control signal based on an output signal of the machine learning system, wherein the output signal is determined based on the image, wherein the machine learning system is trained by:

determining an augmentation of a second training image by:

obtaining a trained encoder, the encoder being configured for determining a latent representation of an image, the encoder being trained by:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image,

masking out parts of the noise image to determine a masked noise image,

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network, and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image;

determining a first latent representation and a noise image by providing the second training image to the trained encoder;

altering the first latent representation to determine a second latent representation;

determining the augmentation by providing the second latent representation and the noise image used in the training the encoder, as input to the generator; and

training the machine learning system based on the augmentation.

14 . A non-transitory machine-readable storage medium on which is stored a computer program for training an encoder, wherein the encoder is configured for determining a latent representation of an image, and the computer program, when executed by a computer, causing the computer to train the encoder by performing the following steps:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image;

masking out parts of the noise image to determine a masked noise image;

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network; and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image.