IP Library › Granted Patent US 12,620,215
Granted Patent B1
US 12,620,215 · App. 18/188,799 · Granted May 5, 2026

Systems and methods for deep learning-based image and video modifications

Inventors: Oliver Dayun Liu (Irvine, CA); Wenbin Ouyang (Redmond, WA)
Assignee: Amazon Technologies, Inc.
G06V10/82G06T3/4046G06V10/751G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,215
App. No.
18/188,799
Granted
May 5, 2026
Kind
B1
Abstract

Systems and methods for deep learning-based image and video modifications are provided. Particularly, a combination of two different neural networks may be used to perform a modification to existing image and/or video content (for example, increasing the resolution of an older video). The original image and/or video content may be provided to the first neural network, which may extract information about the image and/or frames of the video content. This information may then be provided to a second neural network, which may use the information to produce the modified image and/or video content.

Claims (79)

1 . A method comprising:

receiving, by a first neural network, a sequence of frames of a video, the sequence of frames including at least a first frame and a second frame, wherein the first frame and the second frame are initially rendered at a lower resolution;

determining, by the first neural network, a first video information embedding for the first frame, wherein the first video information embedding includes an indication of at least one of: that a resolution modification is to be performed to the first frame or a modified resolution for the first frame;

determining, by the first neural network, a second video information embedding for the second frame, wherein the second video information embedding includes an indication that the second frame includes an intentional effect including at least one of: intentional blurriness, varied level of transparency, an intended lighting artifact, or a particle effect;

outputting, by the first neural network and to a second neural network, one or more vectors including the first video information embedding and the second video information embedding;

receiving, by the second neural network, the one or more vectors;

determining, by the second neural network and based on the first video information embedding for the first frame and the second video information embedding for the second frame, that the first frame is to be upscaled to an increased resolution and the second frame is to remain at a same resolution to prevent the intentional effect from being diminished or removed from the second frame by the upscaling;

modifying, by the second neural network, the first frame to produce a modified first frame at the increased resolution; and

outputting, by the second neural network, the modified first frame and the second frame.

2 . The method of claim 1 , further comprising:

receiving, by a first loss function, the one or more vectors;

determining a comparison between the first video information embedding and the second video information embedding and ground truth data; and

training the first neural network based on the comparison.

3 . The method of claim 1 , further comprising:

receiving, by a second loss function, the modified first frame;

outputting, by the second loss function, and indication of a likelihood that the modified first frame was produced by the second neural network; and

training the second neural network based on the indication.

4 . The method of claim 1 , further comprising:

receiving, by a third loss function, the modified first frame;

determining a second comparison between the modified first frame and ground truth image data; and

training the second neural network based on the second comparison.

5 . A method comprising:

receiving, by a first neural network, first image data and second image data, wherein the first image data is associated with a first video frame and the second image data is associated with a second video frame;

determining, by the first neural network, a first classification for the first image data and a second classification for the second image data, wherein the second classification indicates that the second video frame includes an intended effect;

receiving, by a second neural network, the first image data, the second image data, the first classification for the first image data, and the second classification for the second image data;

determining that modifying the second video frame would remove or diminish the intended effect;

determining, by the second neural network and based on the first classification for the first image data and the second classification for the second image data, a first type of modification to perform to first image data instead of the second image data based on the determination that modifying the second video frame would remove or diminish the intended effect; and

outputting, by the second neural network, the second image data and third image data including the first type of modification.

6 . The method of claim 5 , wherein the first type of modification includes an increase in a resolution of the first image data.

7 . The method of claim 5 , wherein the first image data is a first frame of a video and the second image data is a second frame of a video.

8 . The method of claim 5 , further comprising:

receiving, by the first neural network, third image data;

determining, by the first neural network, a third classification for the third image data;

receiving, by the second neural network, the third image data and the third classification for the third image data;

determining, by the second neural network and based on the third classification for the third image data, a second type of modification to perform to third image data; and

outputting, by the second neural network, fourth image data including the second type of modification, wherein the first type of modification is different than the second type of modification.

9 . The method of claim 5 , wherein the first classification includes an indication of at least one of: that a resolution modification is to be performed to the first image data, a modified resolution for the first image data, that the first image data includes intentional blurriness, that the first image data includes varied level of transparency, that the first image data and/or the second image data includes an intended lighting artifact, and that the first image data includes a particle effect.

10 . The method of claim 5 , further comprising:

receiving, by a first loss function, the first classification and the second classification;

determining a comparison between the first classification and the second classification and ground truth data; and

training the first neural network based on the comparison.

11 . The method of claim 5 , further comprising:

receiving, by a second loss function, the third image data;

outputting, by the second loss function, and indication of a likelihood that the third image data was produced by the second neural network; and

training the second neural network based on the indication.

12 . The method of claim 5 , further comprising:

receiving, by a third loss function, the third image data;

determining a second comparison between the third image data and ground truth image data; and

training the second neural network based on the second comparison.

13 . A system comprising:

memory that stores computer-executable instructions; and

one or more processors configured to access the memory and execute the computer-executable instructions to:

receive, by a first neural network, first image data and second image data, wherein the first image data is associated with a first video frame and the second image data is associated with a second video frame;

determine, by the first neural network, a first classification for the first image data and a second classification for the second image data, wherein the second classification indicates that the second video frame includes an intended effect;

determine that modifying the second video frame would remove or diminish the intended effect;

receive, by a second neural network, the first image data, the second image data, the first classification for the first image data, and the second classification for the second image data;

determine, by the second neural network and based on the first classification for the first image data and the second classification for the second image data, a first type of modification to perform to first image data instead of the second image data based on the determination that modifying the second video frame would remove or diminish the intended effect; and

output, by the second neural network, the second image data and third image data including the first type of modification.

14 . The system of claim 13 , wherein the first type of modification includes an increase in a resolution of the first image data.

15 . The system of claim 13 , wherein the first image data is a first frame of a video and the second image data is a second frame of a video.

16 . The system of claim 13 , wherein the one or more processors are further configured to execute the computer-executable instructions to:

receive, by the first neural network, third image data;

determine, by the first neural network, a third classification for the third image data;

receiving, by the second neural network, the third image data and the third classification for the third image data;

determine, by the second neural network and based on the third classification for the third image data, a second type of modification to perform to third image data; and

output, by the second neural network, fourth image data including the second type of modification, wherein the first type of modification is different than the second type of modification.

17 . The system of claim 13 , wherein the first classification includes an indication of at least one of: that a resolution modification is to be performed to the first image data, a modified resolution for the first image data, that the first image data includes intentional blurriness, that the first image data includes varied level of transparency, and that the first image data includes a particle effect.

18 . The system of claim 13 , wherein the one or more processors are further configured to execute the computer-executable instructions to:

receive, by a first loss function, the first classification and the second classification;

determine a comparison between the first classification and the second classification and ground truth data; and

train the first neural network based on the comparison.

19 . The system of claim 13 , wherein the one or more processors are further configured to execute the computer-executable instructions to:

receive, by a second loss function, the third image data;

output, by the second loss function, and indication of a likelihood that the third image data was produced by the second neural network; and

train the second neural network based on the indication.

20 . The system of claim 13 , wherein the one or more processors are further configured to execute the computer-executable instructions to:

receive, by a third loss function, the third image data;

determine a second comparison between the third image data and ground truth image data; and

train the second neural network based on the second comparison.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2023
From: LIU, OLIVER DAYUN; OUYANG, WENBIN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 063108/0192 →
References Cited (7)
US 20170347110A1 · Wang · 2017 [cited by examiner]
US 20200349681A1 · Andrei · 2020 [cited by examiner]
US 20210097646A1 · Choi · 2021 [cited by examiner]
US 20220108421A1 · Shacklett · 2022 [cited by examiner]
US 20230306739A1 · Yu · 2023 [cited by examiner]
US 20240104912A1 · Kim · 2024 [cited by examiner]
US 20240153033A1 · Wang · 2024 [cited by examiner]