IP Library › Granted Patent US 12,682,629
Granted Patent B2
US 12,682,629 · App. 18/485,174 · Granted Jul 14, 2026

Device and method for determining an encoder configured image analysis

Inventors: Yumeng Li (Tuebingen, DE); Anna Khoreva (Stuttgart, DE); Dan Zhang (Leonberg, DE)
Assignee: ROBERT BOSCH GMBH
G06V10/82G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,629
App. No.
18/485,174
Filed
Oct 11, 2023
Granted
Jul 14, 2026
Kind
B2
Art Unit
2635
USPC
382/156
Abstract

A computer-implemented method for training an encoder. The encoder is configured for determining a latent representation of an image. Training the encoder includes: determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image; masking out parts of the noise image, thereby determining a masked noise image; determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network; training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image.

Claims (66)

1 . A computer-implemented method for training an encoder, wherein the encoder is configured for determining a latent representation of an image, and the training of the encoder comprises the following steps of:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image;

masking out parts of the noise image to determine a masked noise image;

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network; and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image.

2 . The method according to claim 1 , wherein the masking out of parts of the noise image includes replacing values within the parts with randomly drawn values.

3 . The method according to claim 1 , wherein the loss value is determined based on a loss function, wherein a first term of the loss function characterizes the difference between the predicted image and the training image.

4 . The method according to claim 3 , wherein the first term further characterizes a masking of the difference, wherein the masking removes pixels from the difference that fall into the masked-out parts.

5 . The method according to claim 3 , wherein the loss function includes a second term that characterizes a norm of the noise image predicted by the encoder.

6 . The method according to claim 5 , wherein the loss function includes a third term characterizing a negative log likelihood of an output signal of a discriminator, wherein the output signal is determined by the discriminator by providing the predicted image to the discriminator.

7 . The method according to claim 6 , wherein the training image is determined by providing a randomly sampled latent representation or a user defined latent representation to the generator and wherein the loss function includes a fourth term characterizing a difference between the randomly sampled or user defined latent representation and the latent representation determined from the encoder.

8 . The method according to claim 3 , wherein the loss function includes a fifth term characterizing a difference of a first feature representation determined by providing the training image to a feature extractor and a second feature representation determined by providing the predicted image to the feature extractor, wherein the difference does not characterize features characterizing pixels in the masked-out parts.

9 . A computer-implemented method for determining an augmentation of an image, comprising the following steps:

obtaining a trained encoder, the encoder being configured for determining a latent representation of an image, the encoder being trained by:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image,

masking out parts of the noise image to determine a masked noise image,

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network, and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image;

determining a first latent representation and a noise image by providing the image to the trained encoder;

altering the first latent representation to determine a second latent representation; and

determining the augmentation by providing the second latent representation and the noise image used in the training the encoder, as input to the generator.

10 . A computer-implemented method for training a machine learning system, wherein the machine learning system is configured for determining an output signal characterizing a classification and/or regression analysis of an image, wherein the method comprises the following steps:

determining an augmentation of a second training image by:

obtaining a trained encoder, the encoder being configured for determining a latent representation of an image, the encoder being trained by:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image,

masking out parts of the noise image to determine a masked noise image,

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network, and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image;

determining a first latent representation and a noise image by providing the second training image to the trained encoder;

altering the first latent representation to determine a second latent representation;

determining the augmentation by providing the second latent representation and the noise image used in the training the encoder, as input to the generator; and

training the machine learning system based on the augmentation.

11 . A computer-implemented method for determining a control signal of an actuator using a machine learning system, wherein the machine learning system is configured for determining an output signal characterizing a classification and/or regression analysis of an image, wherein the method comprises the following steps:

determining an augmentation of a second training image by:

obtaining a trained encoder, the encoder being configured for determining a latent representation of an image, the encoder being trained by:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image,

masking out parts of the noise image to determine a masked noise image,

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network, and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image;

determining a first latent representation and a noise image by providing the second training image to the trained encoder;

altering the first latent representation to determine a second latent representation;

determining the augmentation by providing the second latent representation and the noise image used in the training the encoder, as input to the generator; and

training the machine learning system based on the augmentation; and

determining a control signal based on an output signal of the trained machine learning system, wherein the output signal is determined based on the image.

12 . A training system for training an encoder, wherein the encoder is configured for determining a latent representation of an image, and wherein the training system is configured to:

determine a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image;

mask out parts of the noise image to determine a masked noise image;

determine a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network; and

train the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image.

13 . A control system configured to determine a control signal of an actuator using a machine learning system, wherein the machine learning system is configured for determining an output signal characterizing a classification and/or regression analysis of an image, wherein the control system is configured to:

determine a control signal based on an output signal of the machine learning system, wherein the output signal is determined based on the image, wherein the machine learning system is trained by:

determining an augmentation of a second training image by:

obtaining a trained encoder, the encoder being configured for determining a latent representation of an image, the encoder being trained by:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image,

masking out parts of the noise image to determine a masked noise image,

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network, and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image;

determining a first latent representation and a noise image by providing the second training image to the trained encoder;

altering the first latent representation to determine a second latent representation;

determining the augmentation by providing the second latent representation and the noise image used in the training the encoder, as input to the generator; and

training the machine learning system based on the augmentation.

14 . A non-transitory machine-readable storage medium on which is stored a computer program for training an encoder, wherein the encoder is configured for determining a latent representation of an image, and the computer program, when executed by a computer, causing the computer to train the encoder by performing the following steps:

determining a latent representation and a noise image by providing a training image to the encoder, wherein the encoder is configured for determining a latent representation and a noise image for a provided image;

masking out parts of the noise image to determine a masked noise image;

determining a predicted image by providing the latent representation and the masked noise image to a generator of a generative adversarial network; and

training the encoder by adapting parameters of the encoder based on a loss value, wherein the loss value characterizes a difference between the predicted image and the training image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2024
From: LI, YUMENG; KHOREVA, ANNA; ZHANG, DAN
To: ROBERT BOSCH GMBH
Reel/Frame 066031/0111 →
Priority Claims (1)
EP 22201998 · Oct 17, 2022 · regional
Continuity (1)
Related Publication 20240135699A1 · Apr 25, 2024
References Cited (29)
US 10726525B2 · El-Khamy · 2020 [cited by examiner]
US 11537881B2 · Choi · 2022 [cited by examiner]
US 11551034B2 · Karimi · 2023 [cited by examiner]
US 11763135B2 · Wang · 2023 [cited by examiner]
US 20200349393A1 · Zhong · 2020 [cited by examiner]
US 20200372308A1 · Anirudh · 2020 [cited by examiner]
US 20210118099A1 · Kearney · 2021 [cited by examiner]
US 20220245451A1 · Arik · 2022 [cited by examiner]
US 20230377213A1 · Naruniec · 2023 [cited by examiner]
US 20240013441A1 · Le · 2024 [cited by examiner]
US 20240078726A1 · Weber · 2024 [cited by examiner]
EP 3739515A1 · 2020 [cited by examiner]
EP 3477553B1 · 2023 [cited by examiner]
Weihao Xia et al.,“GAN Inversion: A Survey,”Jun. 9, 2022, IEEE Transactions On Pattern Analysis and Machine Intelligence, vol. 45, No. 3, Mar. 2023, pp. 3121-3133. [cited by examiner]
Yiping Gao et al.,“A Generative Adversarial Network Based Deep Learning Method for Low-Quality Defect Image Reconstruction and Recognition,” Jul. 13, 2020, IEEE Transactions on Industrial Informatics, vol. 17, No. 5, Ma… [cited by examiner]
Tero Karras et al.,“Analyzing and Improving the Image Quality of StyleGAN,” Jun. 2020, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 8110-8116. [cited by examiner]
Yu Li et al.,“Asymmetric GAN for Unpaired Image-to-Image Translation,” Jun. 19, 2019, IEEE Transactions on Image Processing, vol. 28, No. 12, Dec. 2019, pp. 5881-5894. [cited by examiner]
Changhee Han et al.,“Combining Noise-to-Image and Image-to-Image GANs: Brain MR Image Augmentation for Tumor Detection,” Oct. 15, 2019, IEEEAccess, vol. 7, 2019,pp. 156966-156974. [cited by examiner]
Han Zhang et al.,“Cross-Modal Contrastive Learning for Text-to-Image Generation,” Jun. 2021, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 833-840. [cited by examiner]
Elad Richardson et al.,“Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation,” Jun. 2021, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 2287-2293. [cited by examiner]
Nan Liang et al.,“End-To-End Retina Image Synthesis Based on CGAN Using Class Feature Loss and Improved Retinal Detail Loss,” Aug. 4, 2022, IEEEAccess, vol. 10, 2022,pp. 83125-83133. [cited by examiner]
Zhen Qin et al.,“Segmentation mask and feature similarity loss guided GAN for object-oriented image-to-image translation,” Apr. 4, 2022, Information Processing and Management 59 (2022),pp. 1-14. [cited by examiner]
Vu Nguyen et al.,“Shadow Detection with Conditional Generative Adversarial Networks,” Oct. 2017, Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017,pp. 4510-4516. [cited by examiner]
Atiye Sadat Hashemi et al.,“Secure deep neural networks using adversarial image generation and training with Noise-GAN,” Jul. 2, 2019, Computers and Security 86( 2019) , pp. 372-384. [cited by examiner]
Xia et al., “GAN Inversion: a Survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, pp. 1-21. [cited by applicant]
Richardson et al., “Encoding in Style: a Stylegan Encoder for Image-to-Image Translation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 2287-2296. <https://openacce… [cited by applicant]
Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4396-4405. <https://sci-hub.ru/10.1109/cvp… [cited by applicant]
Karras et al., “Analyzing and Improving the Image Quality of StyleGAN,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 8107-8116. <https://sci-hub.ru/10.1109/cvpr42600.2020.00813> … [cited by applicant]
Zhang et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 586-595. <https://sci-hub.ru/10.1109/cvpr.2018.00068… [cited by applicant]