IP Library Granted Patent US 12,190,233
Granted Patent B2
US 12,190,233 · App. 17/111,072 · Granted Jan 7, 2025

Data style transformation with adversarial models

Inventors: Christian Schäfer (Berlin, DE); Florian Kuhlmann (Berlin, DE)
Assignee: LEVERTON HOLDING LLC
G06N3/08G06N3/045G06V10/82G06V20/62G06V30/19013G06V30/19147G06V30/413G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,233
App. No.
17/111,072
Granted
Jan 7, 2025
Kind
B2
Abstract

Systems and methods for transforming data between multiple styles are provided. In one embodiment, a system is provided that includes a generator model, a discriminator model, and a preserver model. The generator model may be configured to receive data in a first style and generate converted data in a second style. The discriminator model may be configured to receive the converted data from the generator model, compare the converted data to original data in the second style, and compute a resemblance measure based on the comparison. The preserver model may be configured to receive the converted data from the generator model and compute an information measure of the converted data. The generator model may also be trained to optimize the resemblance measure and the information measure.

Claims (97)

1. A system comprising:

a processor; and

a memory storing instructions which, when executed by the processor, cause the processor to implement:

a generator model configured to receive data in a first style and generate converted data in a second style;

a discriminator model configured to receive the converted data from the generator model, compare the converted data to original data in the second style, and compute a resemblance measure based on the comparison; and

a preserver model configured to receive the converted data from the generator model, recognize information within the converted data, and compute an information measure of the converted data based on a proportion of the converted data for which information is recognized,

wherein the generator model is trained to optimize the resemblance measure and the information measure.

2. The system of claim 1 , wherein the discriminator model is further configured to:

receive the converted data and the original data; and

classify data items within the converted data and data items within the original data as being in the first style or the second style,

wherein the resemblance measure is computed based on a proportion of data items within the converted data classified as being in the second style.

3. The system of claim 1 , wherein the memory contains further instructions which, when executed by the processor, cause the processor to:

iteratively train the generator model based on either the resemblance measure or the information measure.

4. The system of claim 3 , wherein the memory contains further instructions which, when executed by the processor while training the generator model based on the resemblance measure, cause the processor to:

receive, at the generator model, first training data in the first style and generate first converted training data in the second style;

receive, at the discriminator model, the first converted training data, compare the first converted training data to the original data, and compute a training resemblance measure based on the comparison; and

receive the training resemblance measure and update the generator model based on the training resemblance measure.

5. The system of claim 4 , wherein the generator model is trained based on the training resemblance measure until the training resemblance measure exceeds a first predetermined threshold.

6. The system of claim 3 , wherein the memory contains further instructions which, when executed by the processor while training the generator model based on the information measure, cause the processor to:

receive, at the generator model, second training data in the first style and generate second converted training data in the second style;

receive, at the preserver model, the second converted training data and compute a training information measure of the second converted training data; and

receive the training information measure and update the generator model based on the training information measure.

7. The system of claim 6 , wherein the generator model is trained based on the training information measure until the training information measure exceeds a second predetermined threshold.

8. The system of claim 3 , wherein one or both of the discriminator model and the preserver model are separately trained prior to training the generator model.

9. The system of claim 1 , wherein data in the first style includes one or more types of data selected from the group consisting of: images of high quality, text images in a first font, images of low quality, spoken audio of low quality, spoken audio in a first language, video of high quality, and video of a low quality, and

wherein data in the second style includes one or more types of data selected from the group consisting of: images of lower quality, text images in a second font, images of higher quality, spoken audio of higher quality, spoken audio in a second language, video of lower quality, and video of higher quality.

10. The system of claim 9 , wherein data in the first style includes high-quality text images and wherein data in the second style includes text images of lower quality to resemble scanned text images.

11. The system of claim 10 , wherein the generator model is configured, while generating the converted data in the second style, to generate at least one image degradation resembling at least one type of error selected from the group consisting of: scanning artifacts, document damage, blurring errors, stray markings, and document blemishes.

12. The system of claim 10 , wherein the preserver model is configured to:

recognize values corresponding to characters within the converted data; and

compute the information measure based on a proportion of characters within the converted data for which corresponding values were successfully identified.

13. The system of claim 10 , wherein the memory stores further instructions which, when executed by the processor, cause the processor to:

store the converted data for use in training a model configured to recognize text within scanned text images.

14. A method comprising:

receiving, at a generator model, data in a first style;

generating, with the generator model, converted data in a second style;

comparing, with a discriminator model, the converted data to original data in the second style;

computing, with the discriminator model, a resemblance measure based on the comparison;

recognizing, with a preserver model, information within the converted data;

computing, with the preserver model, an information measure of the converted data based on a proportion of the converted data for which information is recognized; and

training the generator model to optimize the resemblance measure and the information measure.

15. The method of claim 14 , further comprising:

receiving, with the discriminator model, the converted data and the original data; and

classifying, with the discriminator model, data items within the converted data and data items within the original data as being in the first style or the second style, and

wherein the resemblance measure is computed based on a proportion of data items within the converted data classified as being in the second style.

16. The method of claim 14 , further comprising:

iteratively training the generator model based on either the resemblance measure or the information.

17. The method of claim 16 , wherein training the generator model based on the resemblance measure further comprises:

receiving, at the generator model, first training data in the first style;

generating, with the generator model, first converted training data in the second style;

comparing, with the discriminator model, the first converted training data to the original data;

computing, with the discriminator model, a training resemblance measure based on the comparison; and

updating the generator model based on the training resemblance measure.

18. The method of claim 17 , wherein the generator model is trained based on the training resemblance measure until the training resemblance measure exceeds a first predetermined threshold.

19. The method of claim 16 , wherein training the generator model based on the information measure further comprises:

receiving, at the generator model, second training data in the first style;

generating, with the generator model, second converted training data in the second style;

computing, with the preserver model, a training information measure of the second converted training data; and

updating the generator model based on the training information measure.

20. The method of claim 19 , wherein the generator model is trained based on the training information measure until the training information measure exceeds a second predetermined threshold.

21. The method of claim 16 , wherein one or both of the discriminator model and the preserver model are separately trained prior to training the generator model.

22. The method of claim 14 , wherein data in the first style includes one or more types of data selected from the group consisting of: images of high quality, text images in a first font, images of low quality, spoken audio of low quality, spoken audio in a first language, video of high quality, and video of a low quality, and

wherein data in the second style includes one or more types of data selected from the group consisting of: images of lower quality, text images in a second font, images of higher quality, spoken audio of higher quality, spoken audio in a second language, video of lower quality, and video of higher quality.

23. The method of claim 22 , wherein data in the first style includes high-quality text images and wherein data in the second style includes text images of lower quality to resemble scanned text images.

24. The method of claim 23 , wherein generating the converted data in the second style includes generating at least one image degradation resembling at least one type of error selected from the group consisting of: scanning artifacts, document damage, blurring errors, stray markings, and document blemishes.

25. The method of claim 23 , further comprising:

recognizing, with the preserver model, values corresponding to characters within the converted data; and

computing, with the preserver model, the information measure based on a proportion of characters within the converted data for which corresponding values were successfully identified.

26. The method of claim 23 , further comprising:

storing the converted data for use in training a model configured to recognize text within scanned text images.

27. A non-transitory, computer-readable medium storing instructions which, when executed by a processor, cause the processor to:

receive, at a generator model, data in a first style;

generate, with the generator model, converted data in a second style;

compare, with a discriminator model, the converted data to original data in the second style;

compute, with the discriminator model, a resemblance measure based on the comparison;

recognize, with a preserver model, information within the converted data;

compute, with the preserver model, an information measure of the converted data based on a proportion of the converted data for which information is recognized; and

train the generator model to optimize the resemblance measure and the information measure.

28. A system comprising:

a processor; and

a memory storing instructions which, when executed by the processor, cause the processor to implement:

a generator model configured to receive data in a first style and generate converted data in a second style;

a discriminator model configured to:

receive the converted data from the generator model and the original data, compare the converted data to original data in the second style,

compute a resemblance measure based on the comparison,

classify data items within the converted data and data items within the original data as being in the first style or the second style, wherein the resemblance measure is computed based on a proportion of data items within the converted data classified as being in the second style, and

a preserver model configured to receive the converted data from the generator model and compute an information measure of the converted data,

wherein the generator model is trained to optimize the resemblance measure and the information measure.

29. A method comprising:

receiving, at a generator model, data in a first style;

generating, with the generator model, converted data in a second style;

comparing, with a discriminator model, the converted data to original data in the second style;

receiving, with the discriminator model, the converted data and the original data;

classifying, with the discriminator model, data items within the converted data and data items within the original data as being in the first style or the second style;

computing, with the discriminator model, a resemblance measure based on the comparison, wherein the resemblance measure is computed based on a proportion of data items within the converted data classified as being in the second style;

computing, with a preserver model, an information measure of the converted data; and

training the generator model to optimize the resemblance measure and the information measure.

Assignments (2)
SECURITY INTEREST Recorded Oct 2, 2025
From: LEVERTON HOLDING, LLC
To: GOLUB CAPITAL MARKETS LLC, AS COLLATERAL AGENT
Reel/Frame 072447/0265 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2021
From: SCHÄFER, CHRISTIAN; KUHLMANN, FLORIAN
To: LEVERTON HOLDING LLC
Reel/Frame 057796/0787 →
Continuity (2)
Provisional Application 62942872 · Dec 3, 2019
Related Publication 20210166125A1 · Jun 3, 2021
References Cited (14)
US 9263036B1 · Graves · 2016 [cited by applicant]
US 10803646B1 · Bogan, III · 2020 [cited by examiner]
US 10810721B2 · Mech · 2020 [cited by examiner]
US 20040022123A1 · Foote et al. · 2004 [cited by applicant]
US 20050006064A1 · Glass et al. · 2005 [cited by applicant]
US 20140014605A1 · Kilgore et al. · 2014 [cited by applicant]
US 20170006125A1 · Yasuma et al. · 2017 [cited by applicant]
US 20190171908A1 · Salavon · 2019 [cited by examiner]
US 20190318474A1 · Han · 2019 [cited by applicant]
US 20200218937A1 · Visentini Scarzanella · 2020 [cited by examiner]
US 20210303925A1 · Hofmann · 2021 [cited by examiner]
The Extended European Search Report dated Nov. 8, 2023 issued for European Patent Application No. 20895706.8. [cited by applicant]
Bhunia et al., “Improving Document Binoarization via Adversarial Noise-Texture Augmentation”, 2019 IEEE International Conference on Image Processing (ICIP), IEEE, 22 Sep. 22, 2019, pp. 2721-2725. [cited by applicant]
International Search Report and Written Opinion dated Feb. 24, 2021 issued for International PCT Application No. PCT/US2020/062838 filed Dec. 2, 2020. [cited by applicant]