IP Library Granted Patent US 12,731,216
Granted Patent B2
US 12,731,216 · App. 19/339,292 · Granted Sep 8, 2026

Class-based training of a machine learning model for image processing

Inventors: Sanchit Sanchit (Vienna, AT); Alexander Tack (Vienna, AT)
Assignee: Canva Pty Ltd
G06T5/60G06T11/60G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,216
App. No.
19/339,292
Granted
Sep 8, 2026
Kind
B2
Abstract

Methods of training a machine learning model for image processing are described. A method of training includes utilising as a learning objective a reduction or minimisation of a combination of both an image loss and a classification loss. A method of training includes utilising unsupervised images pairs generated by applying a selected degradation model to a target image, the selected degradation model being selected based on classification information associated with the target image. Methods for generating unsupervised image pairs and methods for image processing using a trained machine learning model are also described, together with computer systems and computer-readable storage for performing the various methods.

Claims (50)

1 . A method of training a machine learning model for image processing, the method including, by a computer processing system implementing a machine learning model:

for an image pair comprising a first image and a second image, wherein the first image is a degraded image, comprising degraded image characteristics relative to the second image, and the second image is a target image for machine learning:

a) applying a current machine learning model to the degraded image to produce a processed image output;

b) determining a loss for training, the loss for training comprising a loss between the processed image output and the target image;

c) updating parameters of the machine learning model based on the loss for training; and

d) performing processes a) to c) for a plurality of other image pairs until an end condition is met, each of the other image pairs being different to the first image pair and each other, and each image pair comprising a said degraded image and a said target image;

wherein:

a plurality of, up to all of, the image pairs are unsupervised image pairs;

a said unsupervised image pair is one in which the degraded image has been generated by a computational process based on the target image of the unsupervised image pair;

the computational process comprises applying a selected degradation model to the target image;

the selected degradation model is one of a plurality of degradation models available for selection;

the selected degradation model for each of the plurality of unsupervised image pairs is selected based on scene classification information associated with the target image of that unsupervised image pair, the scene classification information being or defining a scene class of the target image.

2 . The method of claim 1 , wherein the first image has a plurality of visual parameters with an associated parameter value, affecting how the first image appears relative to the second image.

3 . The method of claim 2 , wherein the plurality of visual parameters include at least one of: (i) brightness, (ii) contrast, (iii) saturation, (iv) vibrance, (v) whites, (vi) blacks, (vii) shadows and (viii) highlights.

4 . The method of claim 2 , wherein the image pairs comprise a first image pair with a first set of the plurality of visual parameters and a second image pair with a second set of the plurality of visual parameters, the first set being different from the second set.

5 . The method of claim 4 , wherein the first set and the second set are mutually exclusive.

6 . The method of claim 4 , wherein the first set and the set include at least one common visual parameter.

7 . The method of claim 2 , wherein the plurality of visual parameters were selected according to a random or quasi-random process.

8 . The method of claim 2 , wherein a first degradation model of the plurality of degradation models is associated with a first range of values for a first visual parameter of the plurality of visual parameters and a second degradation model of the plurality of degradation models is associated with a second range of values, different to the first range of values, for the first of the plurality of visual parameters and wherein the applying either the first or the second degradation model to the target image comprises determining a value for the first visual parameter from the first or the second range of values respectively.

9 . The method of claim 8 , wherein determining a value for the first visual parameter within the first or second range of values comprises a random or quasi-random selection process.

10 . The method of claim 2 , wherein the visual parameters are expressed as differentiable functions.

11 . The method of claim 1 , wherein the classification information associated with at least one of the target images identifies a class of one of: (i) people, (ii) nature, (iii) sunrise and sunset, (iv) animals, (v) city, (vi) food, and (vii) night.

12 . The method of claim 1 , wherein:

a) applying the current machine learning model to the degraded image also produces a first classification output; and

b) the loss for training also comprises a loss between the first classification output and the classification information.

13 . The method of claim 12 , wherein the loss for training is a mathematical combination of the loss between the processed image output and the target image and the loss between the first classification output and the classification information.

14 . A computer processing system including one or more computer processors and computer-readable storage, the computer processing system configured to perform a method comprising:

for an image pair comprising a first image and a second image, wherein the first image is a degraded image, comprising degraded image characteristics relative to the second image, and the second image is a target image for machine learning:

a) applying a current machine learning model to the degraded image to produce a processed image output;

b) determining a loss for training, the loss for training comprising a loss between the processed image output and the target image;

c) updating parameters of the machine learning model based on the loss for training; and

d) performing processes a) to c) for a plurality of other image pairs until an end condition is met, each of the other image pairs being different to the first image pair and each other, and each image pair comprising a said degraded image and a said target image;

wherein:

a plurality of, up to all of, the image pairs are unsupervised image pairs;

a said unsupervised image pair is one in which the degraded image has been generated by a computational process based on the target image of the unsupervised image pair;

the computational process comprises applying a selected degradation model to the target image;

the selected degradation model is one of a plurality of degradation models available for selection;

the selected degradation model for each of the plurality of unsupervised image pairs is selected based on scene classification information associated with the target image of that unsupervised image pair, the scene classification information being or defining a scene class of the target image.

15 . Non-transitory computer-readable storage storing instructions for a computer processing system, wherein the instructions, when executed by the computer processing system cause the computer processing system to perform a method comprising:

for an image pair comprising a first image and a second image, wherein the first image is a degraded image, comprising degraded image characteristics relative to the second image, and the second image is a target image for machine learning:

a) applying a current machine learning model to the degraded image to produce a processed image output;

b) determining a loss for training, the loss for training comprising a loss between the processed image output and the target image;

c) updating parameters of the machine learning model based on the loss for training; and

d) performing processes a) to c) for a plurality of other image pairs until an end condition is met, each of the other image pairs being different to the first image pair and each other, and each image pair comprising a said degraded image and a said target image;

wherein:

a plurality of, up to all of, the image pairs are unsupervised image pairs;

a said unsupervised image pair is one in which the degraded image has been generated by a computational process based on the target image of the unsupervised image pair;

the computational process comprises applying a selected degradation model to the target image;

the selected degradation model is one of a plurality of degradation models available for selection;

the selected degradation model for each of the plurality of unsupervised image pairs is selected based on scene classification information associated with the target image of that unsupervised image pair, the scene classification information being or defining a scene class of the target image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2025
From: SANCHIT, SANCHIT; TACK, ALEXANDER
To: KALEIDO AI GMBH
Reel/Frame 072368/0651 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2025
From: KALEIDO AI GMBH
To: CANVA PTY LTD
Reel/Frame 072368/0677 →
Priority Claims (1)
AU 2023201686 · Mar 17, 2023 · national
Continuity (2)
Continuation 18600679 · Mar 9, 2024
Related Publication 20260024176A1 · Jan 22, 2026
References Cited (35)
US 10049308B1 · Dhua et al. · 2018 [cited by applicant]
US 11069030B2 · Shen et al. · 2021 [cited by applicant]
US 11182877B2 · Zhu et al. · 2021 [cited by applicant]
US 11341367B1 · Barbosa et al. · 2022 [cited by applicant]
US 11455753B1 · Alemi et al. · 2022 [cited by applicant]
US 11983853B1 · Zhu · 2024 [cited by examiner]
US 20160284070A1 · Pettigrew et al. · 2016 [cited by applicant]
US 20190043210A1 · Chui et al. · 2019 [cited by applicant]
US 20190139262A1 · Wang et al. · 2019 [cited by applicant]
US 20190362475A1 · Lin · 2019 [cited by examiner]
US 20200110994A1 · Goto et al. · 2020 [cited by applicant]
US 20200226421A1 · Almazan · 2020 [cited by examiner]
US 20200334532A1 · Zuev · 2020 [cited by examiner]
US 20200342652A1 · Rowell et al. · 2020 [cited by applicant]
US 20200380293A1 · Finnie et al. · 2020 [cited by applicant]
US 20210004631A1 · Wang · 2021 [cited by applicant]
US 20210150764A1 · Kuo et al. · 2021 [cited by applicant]
US 20210201464A1 · Tariq · 2021 [cited by examiner]
US 20220004823A1 · Shoshan et al. · 2022 [cited by applicant]
US 20220261965A1 · Ji · 2022 [cited by examiner]
US 20230040122A1 · Park et al. · 2023 [cited by applicant]
US 20230298142A1 · Lin · 2023 [cited by examiner]
US 20240135496A1 · Dudhane et al. · 2024 [cited by applicant]
CN 111932462B · 2020 [cited by applicant]
CN 113763296A · 2021 [cited by applicant]
CN 115115910A · 2022 [cited by applicant]
CN 115424013A · 2022 [cited by applicant]
CN 115511733A · 2022 [cited by applicant]
CN 115760605A · 2023 [cited by applicant]
WO 2022179586A1 · 2022 [cited by applicant]
WO 2023000872A1 · 2023 [cited by applicant]
Sharma, Vivek, et al. “Classification-driven dynamic image enhancement.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. (Year: 2018). [cited by examiner]
Lore, Kin Gwn, Adedotun Akintayo, and Soumik Sarkar. “LLNet: A deep autoencoder approach to natural low-light image enhancement.” Pattern Recognition 61 (2017): 650-662. (Year: 2017). [cited by examiner]
Gao, Yuan et al., “NDDR-CNN: Layerwise Feature Fusing in Multi-Task CNNs by Neural Discriminative Dimensionality Reduction,” arXiv:1801.08297v4, pp. 3205-3214, Apr. 4, 2019. [cited by applicant]
Rad, Mohammad Saeed, “Benefiting from multitask learning to improve single image super-resolution,” Neurocomputing 398 (2020): 304-313. (Year: 2020). [cited by applicant]