IP Library Granted Patent US 12,056,849
Granted Patent B2
US 12,056,849 · App. 17/466,711 · Granted Aug 6, 2024

Neural network for image style translation

Inventors: Michal Lukác (Boulder Creek, CA); Daniel Sýkora (Prague, CZ); David Futschik (Liberec, CZ); Zhaowen Wang (San Jose, CA); Elya Shechtman (Seattle, WA)
Assignees: Adobe Inc.; CZECH TECHNICAL UNIVERSITY IN PRAGUE
G06T5/50G06F18/214G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,056,849
App. No.
17/466,711
Granted
Aug 6, 2024
Kind
B2
Abstract

Embodiments are disclosed for translating an image from a source visual domain to a target visual domain. In particular, in one or more embodiments, the disclosed systems and methods comprise a training process that includes receiving a training input including a pair of keyframes and an unpaired image. The pair of keyframes represent a visual translation from a first version of an image in a source visual domain to a second version of the image in a target visual domain. The one or more embodiments further include sending the pair of keyframes and the unpaired image to an image translation network to generate a first training image and a second training image. The one or more embodiments further include training the image translation network to translate images from the source visual domain to the target visual domain based on a calculated loss using the first and second training images.

Claims (68)

1. A computer-implemented method comprising:

receiving a training input including a pair of keyframes of a training video sequence and an unpaired image of the training video sequence, the pair of keyframes representing a visual translation of an image, wherein a first keyframe in the pair of keyframes is a first version of the image in a source visual domain and a second keyframe in the pair of keyframes is a second version of the image in a target visual domain, and wherein the unpaired image is in the source visual domain;

generating, by an image translation network, a first training image from the first version of the image and a second training image from the unpaired image, wherein the first training image and the second training image are generated in the target visual domain based on the second version of the image; and

training the image translation network to translate images from the source visual domain to the target visual domain based on a calculated loss using the first training image and the second training image.

2. The computer-implemented method of claim 1 , wherein generating the first training image from the first version of the image and the second training image from the unpaired image comprises:

generating the first training image for the first version of the image by translating the first version of the image from the source visual domain to the target visual domain; and

generating the second training image for the unpaired image by translating the unpaired image from the source visual domain to the target visual domain.

3. The computer-implemented method of claim 1 , wherein the image translation network is configured to generate the second training image by:

analyzing the unpaired image to identify contents of the unpaired image;

analyzing the second version of the image from the pair of keyframes to identify a visual style of the second version of the image; and

applying the visual style of the second version of the image to the contents of the unpaired image.

4. The computer-implemented method of claim 1 , wherein training the image translation network based on the calculated loss comprises:

calculating, using a first loss function, a first loss of the first training image and the second version of the image from the pair of keyframes;

calculating, using a second loss function, a second loss of the second training image and the second version of the image from the pair of keyframes; and

determining the calculated loss from the first loss and the second loss.

5. The computer-implemented method of claim 4 , wherein calculating the first loss of the first training image and the second version of the image from the pair of keyframes comprises:

minimizing a sum of the absolute differences between the first training image and the second version of the image from the pair of keyframes.

6. The computer-implemented method of claim 4 , wherein calculating the second loss of the second training image and the second version of the image from the pair of keyframes comprises:

measuring a visual consistency between the second training image and the second version of the image from the pair of keyframes.

7. The computer-implemented method of claim 1 , further comprising:

receiving an input including a video sequence and a request to translate frames of the video sequence from the source visual domain to the target visual domain;

sending the frames of the video sequence to the trained image translation network to translate the frames of the video sequence from the source visual domain to the target visual domain; and

returning an output including a translated version of the video sequence in the target visual domain.

8. The computer-implemented method of claim 1 , further comprising:

receiving a second training input including a second pair of keyframes of a source image and a second unpaired image, the second pair of keyframes of the source image representing a visual translation of the source image, wherein a third keyframe in the second pair of keyframes is a first version of a portion of the source image in a second source visual domain and a fourth keyframe in the second pair of keyframes is a second version of the portion of the source image in a second target visual domain, and wherein the second unpaired image is in the source visual domain;

generating, by the image translation network, a third training image from the first version of the portion of the source image and a fourth training image from the second unpaired image, wherein the third training image and the fourth training image are generated in the second target visual domain based on the second version of the portion of the source image; and

training the image translation network to translate images from the second source visual domain to the second target visual domain based on a calculated loss using the third training image and the fourth training image.

9. A non-transitory computer-readable storage medium including instructions stored thereon which, when executed by at least one processor, cause the at least one processor to:

receive a training input including a pair of keyframes of a training video sequence and an unpaired image of the training video sequence, the pair of keyframes representing a visual translation of an image, wherein a first keyframe in the pair of keyframes is a first version of the image in a source visual domain and a second keyframe in the pair of keyframes is a second version of the image in a target visual domain, and wherein the unpaired image is in the source visual domain;

generating, by an image translation network, a first training image from the first version of the image and a second training image from the unpaired image, wherein the first training image and the second training image are generated in the target visual domain based on the second version of the image; and

train the image translation network to translate images from the source visual domain to the target visual domain based on a calculated loss using the first training image and the second training image.

10. The non-transitory computer-readable storage medium of claim 9 , wherein generating the first training image from the first version of the image and the second training image from the unpaired image comprises:

generating the first training image for the first version of the image by translating the first version of the image from the source visual domain to the target visual domain; and

generating the second training image for the unpaired image by translating the unpaired image from the source visual domain to the target visual domain.

11. The non-transitory computer-readable storage medium of claim 9 , wherein the image translation network is configured to generate the second training image by:

analyzing the unpaired image to identify contents of the unpaired image;

analyzing the second version of the image from the pair of keyframes to identify a visual style of the second version of the image; and

applying the visual style of the second version of the image to the contents of the unpaired image.

12. The non-transitory computer-readable storage medium of claim 9 , wherein training the image translation network based on the calculated loss comprises:

calculating, using a first loss function, a first loss of the first training image and the second version of the image from the pair of keyframes;

calculating, using a second loss function, a second loss of the second training image and the second version of the image from the pair of keyframes; and

determining the calculated loss from the first loss and the second loss.

13. The non-transitory computer-readable storage medium of claim 12 , wherein calculating the first loss of the first training image and the second version of the image from the pair of keyframes comprises:

minimizing a sum of the absolute differences between the first training image and the second version of the image from the pair of keyframes.

14. The non-transitory computer-readable storage medium of claim 12 , wherein calculating the second loss of the second training image and the second version of the image from the pair of keyframes comprises:

measuring a visual consistency between the second training image and the second version of the image from the pair of keyframes.

15. The non-transitory computer-readable storage medium of claim 9 , further comprising:

receiving an input including a video sequence and a request to translate frames of the video sequence from the source visual domain to the target visual domain;

sending the frames of the video sequence to the trained image translation network to translate the frames of the video sequence from the source visual domain to the target visual domain; and

returning an output including a translated version of the video sequence in the target visual domain.

16. The non-transitory computer-readable storage medium of claim 9 , further comprising:

receiving a second training input including a second pair of keyframes of a source image and a second unpaired image, the second pair of keyframes of the source image representing a visual translation of the source image, wherein a third keyframe in the second pair of keyframes is a first version of a portion of the source image in a second source visual domain and a fourth keyframe in the second pair of keyframes is a second version of the portion of the source image in a second target visual domain, and wherein the second unpaired image is in the source visual domain;

generating, by the image translation network, a third training image from the first version of the portion of the source image and a fourth training image from the second unpaired image, wherein the third training image and the fourth training image are generated in the second target visual domain based on the second version of the portion of the source image; and

training the image translation network to translate images from the second source visual domain to the second target visual domain based on a calculated loss using the third training image and the fourth training image.

17. A computer-implemented method comprising:

receiving, by a machine-learning backed service, an input including a video sequence and a request to translate frames of the video sequence from a source visual domain to a target visual domain;

sending the frames of the video sequence to an image translation network trained to translate the frames of the video sequence from the source visual domain to the target visual domain, the image translation network trained by generating training images in a target visual domain from a training input that includes a pair of keyframes of a training video sequence and an unpaired image of the training video sequence, the pair of keyframes representing a visual translation of an image, wherein a first keyframe in the pair of keyframes is a first version of the image in a source visual domain and a second keyframe in the pair of keyframes is a second version of the image in the target visual domain, wherein the unpaired image is in the source visual domain; and

returning an output including a translated version of the video sequence in the target visual domain.

18. The computer-implemented method of claim 17 , wherein the image translation network is trained using a training system configured to:

generate, by the image translation network, a first training image from the first version of the image and a second training image from the unpaired image, wherein the first training image and the second training image are generated in the target visual domain based on the second version of the image; and

train the image translation network to translate images from the source visual domain to the target visual domain based on a calculated loss using the first training image and the second training image.

19. The computer-implemented method of claim 18 , wherein training the image translation network based on the calculated loss comprises:

generating the first training image for the first version of the image by translating the first version of the image from the source visual domain to the target visual domain; and

generating the second training image for the unpaired image by translating the unpaired image from the source visual domain to the target visual domain.

20. The computer-implemented method of claim 18 , wherein the image translation network is configured to generate the second training image for the unpaired image by:

analyzing the unpaired image to identify contents of the unpaired image;

analyzing the second version of the image from the pair of keyframes to identify a visual style of the second version of the image; and

applying the visual style of the second version of the image to the contents of the unpaired image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2022
From: LUKÁC, MICHAL; FUTSCHIK, DAVID; WANG, ZHAOWEN; SHECHTMAN, ELYA
To: ADOBE INC.
Reel/Frame 060325/0497 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2022
From: SÝKORA, DANIEL
To: CZECH TECHNICAL UNIVERSITY IN PRAGUE
Reel/Frame 060325/0563 →
Continuity (1)
Related Publication 20230070666A1 · Mar 9, 2023