IP Library › Granted Patent US 12,482,078
Granted Patent B2
US 12,482,078 · App. 18/013,802 · Granted Nov 25, 2025

Machine learning for high quality image processing

Inventor: Noritsugu Kanazawa (Campbell, CA)
Assignee: GOOGLE LLC
G06T5/77G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,078
App. No.
18/013,802
Granted
Nov 25, 2025
Kind
B2
Abstract

A system or method for inpainting can be aided through the use of machine learning and ground truth data training. The training of machine-learning inpainting models through the use of ground truth image data may add efficiency and precision to the field of image inpainting. Furthermore, machine-learning inpainting models can aid in the non-deterministic prediction of a variety of data types and can be applicable to the removing and/or replacing of a variety of data types. The trained models can be enabled to make predictions without ground truth reassurance due to calibrated parameters tuned through the training.

Claims (62)

1 . A computer-implemented method for training a conditional variational autoencoder to perform image processing, the method comprising:

obtaining, by one or more computing devices, a training example comprising ground truth image data, augmented image data derived from an addition of unwanted image data to the ground truth image data, and a mask that indicates one or more locations of the unwanted image data within the augmented image data;

processing, by the one or more computing devices, the augmented image data and the mask with a first encoder model of the conditional variational autoencoder to generate an embedding for the image data;

processing, by the one or more computing devices, the ground truth image data and the mask with a second encoder model to generate one or more distribution values;

processing, by the one or more computing devices, the embedding and the one or more distribution values with a decoder model of the conditional variational autoencoder to generate predicted image data that comprises replacement image data at the one or more locations indicated by the mask, wherein the replacement image data replaces the unwanted image data;

evaluating, by the one or more computing devices, one or more loss functions based on a comparison of the predicted image data with the ground truth image data wherein evaluating the one or more loss functions comprises evaluating an adversarial loss generated based on a discriminator output generated by a discriminator model based on the predicted image data and the ground truth image data; wherein the discriminator model comprises:

a first texture-level network that processes a portion of the ground truth image data at the one or more locations identified by the mask to generate a first texture discriminator output;

a second texture-level network that processes a portion of the predicted image data at the one or more locations identified by the mask to generate a second texture discriminator output; and

a semantic level network that comprises a shared network that processes the ground truth image data with the unwanted data removed therefrom to generate a semantic discriminator output;

wherein the semantic level network generates the discriminator output based on the first texture discriminator output, the second texture discriminator output, and the semantic discriminator output; and

modifying, by the one or more computing devices, one or more parameter values of the conditional variational autoencoder based at least in part on the one or more loss functions.

2 . The computer-implemented method of claim 1 , wherein the ground truth image data comprises a two-dimensional photograph.

3 . The computer-implemented method of claim 1 , wherein the ground truth image data comprises a ground truth three-dimensional point cloud, the augmented image data comprises one or more unwanted points added to the ground truth three-dimensional point cloud, and the mask identifies the one or more unwanted points.

4 . The computer-implemented method of claim 1 , wherein evaluating the one or more loss functions comprises evaluating a L1 loss between the predicted image data and the ground truth image data.

5 . The computer-implemented method of claim 1 , wherein the distribution values comprise a mean value and a standard deviation value.

6 . The computer-implemented method of claim 5 , wherein the distribution values are penalized by a KL divergence loss function.

7 . The computer-implemented method of claim 1 , further comprising:

multiplying the distribution values by a random value to generate modified distribution values, wherein the discriminator model processes the modified distribution values to generate the predicted image data.

8 . The computer-implemented method of claim 1 , wherein obtaining the augmented image data comprises:

identifying a set of unwanted data;

determining a location for the unwanted data to occlude the ground truth data;

replacing a portion of the ground truth data at the location with the set of unwanted data.

9 . The computer-implemented method of claim 1 , wherein:

the ground truth image data depicts a scene;

the unwanted image data comprises an occluding object; and

the replacement image data depicts one or more portions of the scene occluded by the occluded object.

10 . The computer-implemented method of claim 9 wherein the occluding object comprises artefacts from a denoising process applied to the augmented image data.

11 . A computing system, comprising:

one or more processors;

one or more non-transitory computer-readable storage media that collectively store:

a machine-learned conditional variational autoencoder model trained on a loss function based on a discriminator output generated by a discriminator model based on predicted image data and ground truth image data, comprising:

an encoder, wherein the encoder is configured to encode image data;

a decoder, wherein the decoder is configured to decode encoded image data;

wherein the discriminator model comprises:

a first texture-level network that processes a portion of the ground truth image data at one or more locations identified by a mask to generate a first texture discriminator output;

a second texture-level network that processes a portion of the predicted image data at one or more locations identified by a mask to generate a second texture discriminator output; and

a semantic level network that comprises a shared network that processes the ground truth image data with unwanted data removed therefrom to generate a semantic discriminator output;

wherein the semantic level network generates the discriminator output based on the first texture discriminator output, the second texture discriminator output, and the semantic discriminator output; and

instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

inputting image data and a mask into the encoder, wherein the image data comprises unwanted image data, and wherein the mask indicates a location and size of the unwanted image data;

receiving an embedding from the encoder, wherein the embedding comprises the encoded image data;

inputting the embedding and a conditioning vector into the decoder; and

receiving predicted image data as an output of the decoder, wherein the predicted image data replaces the unwanted image data with predicted replacement data based at least in part on the image data and the conditioning vector.

12 . The computing system of claim 11 , wherein the conditioning vector comprises a zero vector that replaces a set of randomized feature vectors used during training of the machine-learned conditional variational autoencoder model.

13 . The computing system of claim 11 , wherein the machine-learned conditional variational autoencoder model has been trained based on a loss function that compares ground truth training image data against predicted training image data generated by the machine-learned conditional variational autoencoder model based on augmented training image data, the augmented training image data created through insertion of unwanted image data into the ground truth training image data.

14 . The computing system of claim 11 , wherein the input image data comprises a two-dimensional photograph.

15 . The computing system of claim 11 , wherein the input image data comprises a three-dimensional point cloud.

16 . The computing system of claim 11 , wherein:

the unwanted image data comprises an occluding object; and

the predicted replacement image data depicts one or more portions of the scene occluded by the occluded object.

17 . One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations comprising:

obtaining, by one or more computing devices, a training example comprising ground truth data, augmented data derived from an addition of unwanted data to the ground truth data, and a mask that indicates one or more locations of the unwanted data within the augmented data;

processing, by the one or more computing devices, the augmented data and the mask with a first encoder model of the conditional variational autoencoder to generate an embedding for the data;

processing, by the one or more computing devices, the ground truth data and the mask with a second encoder model to generate one or more distribution values;

processing, by the one or more computing devices, the embedding and the one or more distribution values with a decoder model of the conditional variational autoencoder to generate predicted data that comprises replacement data at the one or more locations indicated by the mask, wherein the replacement data replaces the unwanted data;

evaluating, by the one or more computing devices, one or more loss functions based on a comparison of the predicted image data with the ground truth data wherein evaluating the one or more loss functions comprises evaluating an adversarial loss generated based on a discriminator output generated by a discriminator model based on the predicted image data and the ground truth image data; wherein the discriminator model comprises:

a first texture-level network that processes a portion of the ground truth image data at the one or more locations identified by the mask to generate a first texture discriminator output;

a second texture-level network that processes a portion of the predicted image data at the one or more locations identified by the mask to generate a second texture discriminator output; and

a semantic level network that comprises a shared network that processes the ground truth image data with the unwanted data removed therefrom to generate a semantic discriminator output;

wherein the semantic level network generates the discriminator output based on the first texture discriminator output, the second texture discriminator output, and the semantic discriminator output; and

modifying, by the one or more computing devices, one or more parameter values of the conditional variational autoencoder based at least in part on the one or more loss functions.

18 . The one or more non-transitory computer-readable media of claim 17 , wherein the ground truth data comprises ground truth audio waveform data, the augmented data comprises augmented audio waveform data, the replacement data comprises replacement audio waveform data, and the unwanted data comprises unwanted audio waveform data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2023
From: KANAZAWA, NORITSUGU
To: GOOGLE LLC
Reel/Frame 062870/0408 →
Continuity (1)
Related Publication 20230360181A1 · Nov 9, 2023
References Cited (28)
US 11508042B1 · Knuffman · 2022 [cited by examiner]
US 11537134B1 · Wiest · 2022 [cited by examiner]
US 20190251723A1 · Coppersmith, III · 2019 [cited by examiner]
US 20200104640A1 · Poole · 2020 [cited by examiner]
US 20210287352A1 · Calderon · 2021 [cited by examiner]
US 20210343080A1 · Kim · 2021 [cited by examiner]
US 20220277491A1 · Lee · 2022 [cited by examiner]
US 20230267330A1 · Sandler et al. · 2023 [cited by applicant]
WO WO2018102717 · 2018 [cited by examiner]
Oleg et al (“Variational Autoencoder with Arbitrary Conditioning”, arXiv.org preprint arXiv:1806.02382 (2019), pp. 1-25) (Year: 2019). [cited by examiner]
International Preliminary Report on Patentability for Application No. PCT/US2020/040104, mailed Jan. 12, 2023, 9 pages. [cited by applicant]
Machine Translated Chinese Search Report Corresponding to Application No. 2020801020275 on Jun. 21, 2024. [cited by applicant]
Du et al, “Conditional Variational Image Deraining”, Transactions on Image Processing, vol. 29, 2020, pp. 6288-6301. [cited by applicant]
International Search Report for Application No. PCT/ US2021/040104, mailed on Mar. 30, 2021, 2 pages. [cited by applicant]
Ivanov et al, “Variational Autoencoder with Arbitrary Conditioning”, arXiv:1806.02382v3, Jun. 27, 2019, 26 pages. [cited by applicant]
Qin et al, “Automatic Semantic Content Removal by Learning to Neglect”, arXiv:1807.07696v1, Jul. 20, 2018, 12 pages. [cited by applicant]
Zheng et al, “Pluralistic Image Completion”, arXiv:1903.04227v2, Apr. 5, 2019, 22 pages. [cited by applicant]
Du et al., “Conditional Variational Image Deraining”, arXiv:2004.11373v2, May 8, 2020, 14 pages. [cited by applicant]
Github, “JiahuiYu/generative_inpainting”, https://github.com/JiahuiYu/generative_inpainting, retrieved on Mar. 14, 2023, 4 pages. [cited by applicant]
Goodfellow et al., “Generative Adversarial Networks”, arXiv:1406.2661vl, Jun. 10, 2014, 9 pages. [cited by applicant]
Ham et al., “Variational Image Inpainting”, Third Workshop on Bayesian Deep Learning, Conference on Neural Information Processing Systems, Montréal, Canada, Dec. 7, 2018, 6 pages. [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Nets”, arXiv:1611.07004v3, Nov. 26, 2018, 17 pages. [cited by applicant]
Ivanov et al., “Variational Autoencoder with Arbitrary Conditioning”, arXiv:1806.02382v3, Jun. 27, 2019, 25 pages. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes”, arXiv:1312.6114v11, Dec. 10, 2022, 14 pages. [cited by applicant]
Nvidia, “Image Inpainting”, 2018, https://www.nvidia.com/research/inpainting/index.html, retrieved on Mar. 14, 2023, 1 page. [cited by applicant]
Sohn et al., “Learning Structured Output Representation using Deep Conditional Generative Models”, Twenty-eighth International Conference on Neural Information Processing Systems, vol. 2, Dec. 2015, 9 pages. [cited by applicant]
Walker et al., “An Uncertain Future: Forecasting from Static Images using Variational Autoencoders”, arXiv:1606.07873v1, Jun. 25, 2016, 17 pages. [cited by applicant]
Zheng et al., “Pluralistic Image Completion”, arXiv:1903.04227v2, Apr. 5, 2019, 21 pages. [cited by applicant]
Cited By (1)
US 12,682,616