IP Library Granted Patent US 12,725,320
Granted Patent B2
US 12,725,320 · App. 18/425,371 · Granted Sep 1, 2026

Image style transfer

Inventors: Pablo Pernias Pascual de Pobil (San Juan de Alicante, ES); Santiago Iglesias Navarro (Palma de Mallorca, ES); Robert B. Moore (Clermont, FL); David N. Juboor (Johnson City, TN)
Assignee: DISNEY ENTERPRISES, INC.
G06T11/10G06T2211/441
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,320
App. No.
18/425,371
Granted
Sep 1, 2026
Kind
B2
Abstract

Techniques for generating modified images using content information and style information are disclosed. First image data comprising image content information is received, and a content encoder generates a first embedding by extracting the image content information from the first image data. A second embedding generated by a style encoder is received, the second embedding comprising style information of second image data. The style information comprises color information and texture information. A decoder generates a modified image using the first embedding and the second embedding, the modified image comprising the image content information of the first image data and the style information of the second image data.

Claims (80)

1 . A computer-implemented method of generating modified images using content information and style information, the method comprising:

receiving first image data comprising image content information, wherein the image content information comprises positional information of a first object and a second object in the first image data;

generating, using a content encoder, a first embedding comprising the image content information, wherein the content encoder extracts the image content information from the first image data to generate the first embedding;

receiving a second embedding and a third embedding generated by a style encoder, the second embedding comprising first style information of second image data, the third embedding comprising second style information of third image data, wherein the first style information and the second style information each comprises at least one of color information or texture information; and

generating, by a decoder, a modified image using the first embedding, the second embedding, and the third embedding, wherein the modified image comprises the image content information of the first image data, the first style information of the second image data, and the second style information of the third image data, wherein the generating the modified image comprises:

applying a first style to the first object based on the first style information, wherein the applying the first style comprises applying a first user-selected weight of the first style to the first object; and

applying a second style to the second object based on the second style information, wherein the applying the second style comprises applying a second user-selected weight of the second style to the second object.

2 . The computer-implemented method of claim 1 , wherein the first object and the second object correspond to different facial features in the first image data.

3 . The computer-implemented method of claim 1 , wherein the content encoder, the style encoder, and the decoder are included in a machine-learned (ML) model.

4 . The computer-implemented method of claim 3 , further comprising:

receiving a plurality of training image data comprising a plurality of image content information and a plurality of style information;

generating a training dataset using the received plurality of training image data, wherein the received plurality of training image data is pre-processed by performing segmentation to identify features in the plurality of image content information; and

training the ML model using the generated training dataset, wherein the training comprises determining a set of loss functions and corresponding weights for the loss functions.

5 . The computer-implemented method of claim 4 , further comprising:

evaluating accuracy of the trained ML model using a testing dataset comprising at least a portion of the training dataset; and

retraining the trained ML model when the accuracy does not exceed a threshold accuracy, wherein the retraining comprises adjusting a set of weights or training the ML model using a different training dataset.

6 . The computer-implemented method of claim 1 , wherein at least one of the first style information or the second style information is associated with an animation style or an artistic style or technique.

7 . The computer-implemented method of claim 1 , wherein the style encoder extracts the first style information and the second style information from different images.

8 . The computer-implemented method of claim 1 , wherein the image content information comprises positional information of a third object in the first image data, and wherein the method further comprises applying a third style to the third object based on third style information, wherein the first style, the second style, and the third style are different styles.

9 . The computer-implemented method of claim 8 , wherein the first object, the second object, and the third object correspond to different facial features in the first image data.

10 . A non-transitory computer-readable medium carrying instructions that, when executed by a computing system, cause the computing system to perform operations comprising:

receive first image data comprising image content information, wherein the image content information comprises positional information of a first object and a second object in the first image data;

generate, using a content encoder, a content embedding comprising the image content information, wherein the content encoder extracts the image content information from the first image data to generate the content embedding;

receive a first style embedding and a second style embedding generated by a style encoder, the first style embedding comprising first style information of second image data, the second style embedding comprising second style information of third image data, wherein the first style information comprises at least one of first color information or texture information, and wherein the second style information comprises at least one of second color information or texture information; and

generate, by a decoder, a modified image using the content embedding, the first style embedding, and the second style embedding, wherein the modified image comprises the image content information of the first image data, the first style information of the second image data, and the second style information of the third image data, wherein the modified image comprises a first style applied to the first object and a second style applied to the second object, wherein the first style is applied based on a first user-selected weight of the first style to the first object, and wherein the second style is applied based on a second user-selected weight of the second style to the second object.

11 . The non-transitory computer-readable medium of claim 10 , wherein the content encoder, the style encoder, and the decoder are included in a machine-learned (ML) model.

12 . The non-transitory computer-readable medium of claim 11 , wherein the operations further comprise:

receive a plurality of training image data comprising a plurality of image content information and a plurality of style information;

generate a training dataset using the received plurality of training image data, wherein the received plurality of training image data is pre-processed by performing segmentation to identify features in the plurality of image content information; and

train the ML model using the generated training dataset, wherein the training comprises determining a set of loss functions and corresponding weights for the loss functions.

13 . The non-transitory computer-readable medium of claim 12 , wherein the operations further comprise:

evaluate accuracy of the trained ML model using a testing dataset comprising at least a portion of the training dataset; and

retrain the trained ML model when the accuracy does not exceed a threshold accuracy, wherein the retraining comprises adjusting a set of weights or training the ML model using a different training dataset.

14 . The non-transitory computer-readable medium of claim 10 , wherein each of the first style information and the second style information is associated with an animation style or an artistic style or technique.

15 . A computing system comprising:

at least one processor; and

at least one non-transitory memory carrying instructions that, when executed by the at least one processor, cause the computing system to perform operations comprising:

receive first image data comprising image content information, wherein the image content information comprises positional information of a first object and a second object in the first image data;

generate, using a content encoder, a content embedding comprising the image content information, wherein the content encoder extracts the image content information from the first image data to generate the content embedding;

receive a first style embedding and a second style embedding generated by a style encoder, the first style embedding comprising first style information of second image data, the second style embedding comprising second style information of third image data, wherein the first style information comprises at least one of first color information or texture information, and wherein the second style information comprises at least one of second color information or texture information; and

generate, by a decoder, a modified image using the content embedding, the first style embedding, and the second style embedding, wherein the modified image comprises the image content information of the first image data, the first style information of the second image data, and the second style information of the third image data, wherein the modified image comprises a first style applied to the first object based on a first weight and a second style applied to the second object based on a second weight, wherein the first weight and the second weight are selected by a user.

16 . The computing system of claim 15 , wherein the content encoder, the style encoder, and the decoder are included in a machine-learned (ML) model, and wherein the operations further comprise:

receive a plurality of training image data comprising a plurality of image content information and a plurality of style information;

generate a training dataset using the received plurality of training image data, wherein the received plurality of training image data is pre-processed by performing segmentation to identify features in the plurality of image content information; and

train the ML model using the generated training dataset, wherein the training comprises determining a set of loss functions and corresponding weights for the loss functions.

17 . The computing system of claim 16 , wherein the operations further comprise:

evaluate accuracy of the trained ML model using a testing dataset comprising at least a portion of the training dataset; and

retrain the trained ML model when the accuracy does not exceed a threshold accuracy, wherein the retraining comprises adjusting a set of weights or training the ML model using a different training dataset.

18 . The computing system of claim 15 , wherein the second style information is associated with an animation style or an artistic style or technique different than the first style information.

19 . The computing system of claim 15 , further comprising extracting the first style information and the second style information from different images.

20 . A computer-implemented method of generating modified images, the method comprising:

receiving first image data comprising image content information, wherein the image content information comprises positional information of a first object and a second object in the first image data;

generating, using a content encoder, a first embedding comprising the image content information, wherein the content encoder extracts the image content information from the first image data to generate the first embedding;

receiving a second embedding and a third embedding generated by a style encoder, the second embedding comprising first style information of second image data, the third embedding comprising second style information of third image data, wherein the first style information and the second style information each comprises at least one of color information or texture information; and

generating, by a decoder, a modified image using the first embedding, the second embedding, and the third embedding, wherein the modified image comprises the image content information of the first image data, the first style information of the second image data, and the second style information of the third image data, wherein the generating the modified image comprises:

applying a first weight of a first style to the first object based on the first style information; and

applying a second weight of a second style to the second object based on the second style information,

wherein the first weight and the second weight are selected by a user.

21 . The computer-implemented method of claim 20 , wherein the generating the modified image comprises discarding superfluous information of the first image data based on the second embedding or the third embedding.

22 . The computer-implemented method of claim 20 , wherein the first object comprises a first facial feature of a face in the first image data, and wherein the second object comprises a second facial feature of the face.

23 . A computer-implemented method of generating modified images, the method comprising:

receiving first image data comprising image content information, wherein the image content information comprises positional information of a first object and a second object in the first image data;

generating, using a content encoder, a first embedding comprising the image content information, wherein the content encoder extracts the image content information from the first image data to generate the first embedding;

receiving a second embedding and a third embedding generated by a style encoder, the second embedding comprising first style information of second image data, the third embedding comprising second style information of third image data, wherein the first style information and the second style information each comprises at least one of color information or texture information; and

generating, by a decoder, a modified image using the first embedding, the second embedding, and the third embedding, wherein the modified image comprises the image content information of the first image data, the first style information of the second image data, and the second style information of the third image data, wherein the generating the modified image comprises:

applying a first weight of a first style to the first object based on the first style information;

applying a second weight of a second style to the second object based on the second style information; and

discarding superfluous information of the first image data based on the second embedding or the third embedding.

24 . The computer-implemented method of claim 23 , wherein the first weight and the second weight are selected by a user.

25 . The computer-implemented method of claim 23 , wherein: the image content information comprises positional information of a third object in the first image data; the generating the modified image comprises applying a third style to the third object based on third style information; and the first style, the second style, and the third style are different styles.

26 . A computer-implemented method of generating modified images using content information and style information, the method comprising:

receiving first image data comprising image content information, wherein the image content information comprises positional information of a first object, a second object, and a third in the first image data;

generating, using a content encoder, a first embedding comprising the image content information, wherein the content encoder extracts the image content information from the first image data to generate the first embedding;

receiving a second embedding and a third embedding generated by a style encoder, the second embedding comprising first style information of second image data, the third embedding comprising second style information of third image data, wherein the first style information and the second style information each comprises at least one of color information or texture information; and

generating, by a decoder, a modified image using the first embedding, the second embedding, and the third embedding, wherein the modified image comprises the image content information of the first image data, the first style information of the second image data, and the second style information of the third image data, wherein the generating the modified image comprises:

applying a first style to the first object based on the first style information;

applying a second style to the second object based on the second style information; and

applying a third style to the third object based on third style information, wherein the first style, the second style, and the third style are different styles.

27 . The computer-implemented method of claim 26 , wherein the applying the first style comprises applying a first user-selected weight of the first style to the first object, and wherein the applying the second style comprises applying a second user-selected weight of the second style to the second object.

28 . The computer-implemented method of claim 26 , wherein the generating the modified image comprises discarding superfluous information of the first image data based on the second embedding or the third embedding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2024
From: PERNIAS PASCUAL DE POBIL, PABLO; IGLESIAS NAVARRO, SANTIAGO; MOORE, ROBERT B.; JUBOOR, DAVID N.
To: DISNEY ENTERPRISES, INC.
Reel/Frame 066279/0802 →
Continuity (1)
Related Publication 20250245878A1 · Jul 31, 2025
References Cited (19)
US 11610351B2 · Bethge et al. · 2023 [cited by applicant]
US 12159334B2 · Kuta et al. · 2024 [cited by applicant]
US 12340440B2 · Chandran et al. · 2025 [cited by applicant]
US 12345806B2 · Lyu et al. · 2025 [cited by applicant]
US 20200134150A1 · Hunegnaw · 2020 [cited by applicant]
US 20210358164A1 · Liu · 2021 [cited by examiner]
US 20220004921A1 · Balaraman · 2022 [cited by examiner]
US 20220156987A1 · Chandran et al. · 2022 [cited by applicant]
US 20240354895A1 · Ravi · 2024 [cited by examiner]
US 20250078349A1 · Cho · 2025 [cited by examiner]
US 20250111573A1 · Phan · 2025 [cited by applicant]
US 20250245884A1 · Iglesias Navarro · 2025 [cited by examiner]
US 20250371426A1 · Fortkort · 2025 [cited by examiner]
“U.S. Appl. No. 18/655,143, filed May 3, 2024, entitled “Speed and Flexibility in Style Transfer””. [cited by applicant]
“U.S. Appl. No. 18/655,147, filed May 3, 2024, entitled “Differentiable Composition of Attributes in Style Transfer””. [cited by applicant]
“U.S. Appl. No. 18/655,150, filed May 3, 2024, entitled “Semi-Supervised Style Transfer””. [cited by applicant]
Gatys, et al., “A Neural Algorithm of Artistic Style.”, Sep. 2, 2015 Journal of Vision 2016;16(12):326. https://doi.org/10.1167/16.12.326. [cited by applicant]
Gatys, et al., “Controlling Perceptual Factors in Neural Style Transfer”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 3730-3738, doi: 10.1109/CVPR.2017.397. [cited by applicant]
Gatys, et al., “Image style transfer using convolutional neural networks”, Proceedings ofthe IEEE conference on computervision and pattern recognition. 2016. pp. 2414-2423. doi: 10.1109/CVPR.2016.265. [cited by applicant]