IP Library › Granted Patent US 11,875,491
Granted Patent B2
US 11,875,491 · App. 17/696,535 · Granted Jan 16, 2024

Method and system for image processing

Inventors: Thomas Davies (Toronto, CA); Ali Mahdavi-Amiri (North Vancouver, CA); Matthew Panousis (Toronto, CA); Jonathan Bronfman (Toronto, CA); Lon Molnar (Toronto, CA); Paul Birulin (Toronto, CA); Debjoy Chowdhury (Toronto, CA); Ishrat Badami (Toronto, CA); Anton Skourides (Toronto, CA)
Assignee: MONSTERS ALIENS ROBOTS ZOMBIES INC.
G06T5/50G06T7/11G06T11/60G06V10/774G06V40/10G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,875,491
App. No.
17/696,535
Granted
Jan 16, 2024
Kind
B2
Abstract

An image processing system comprising: a computer readable medium and at least one processor configured to provide a machine learning architecture for image processing. In particular, keyframes are selected for modification by a visual artist, and the modifications are used for training the machine learning architecture. The modifications are then automatically propagated to remaining frames requiring modification through interpolation or extrapolation through processing remaining frames through the trained machine learning architecture. The generated modified frames or frame portions can then be inserted into an original video to generate a modified video where the modifications have been propagated. Example usages include automatic computational approaches for aging/de-aging and addition/removal of tattoos or other visual effects.

Claims (49)

1. A computer system configured to automatically interpolate or extrapolate visual modifications from a set of keyframes extracted from a set of target video frames, the system comprising:

a computer processor, operating in conjunction with computer memory maintaining a machine learning model architecture, the computer processor configured to:

receive the set of target video frames;

identify, from the set of target video frames, the set of keyframes;

provide the set of keyframes for visual modification by a human;

receive a set of modified keyframes;

train the machine learning model architecture using the set of modified keyframes and the set of keyframes, the machine learning model architecture including a first autoencoder configured for unity reconstruction of the set of modified keyframes from the set of keyframes to obtain a trained machine learning model architecture; and

process one or more frames of the set of target video frames to generate a corresponding set of modified target video frames having the automatically interpolated or extrapolated visual modifications.

2. The computer system of claim 1 , wherein the visual modifications include visual effects applied to a target human being or a portion of the target human being visually represented in the set of target video frames.

3. The computer system of claim 2 , wherein the visual modifications include at least one of eye-bag addition/removal, wrinkle addition/removal, or tattoo addition/removal.

4. The computer system of claim 2 , wherein the computer processor is further configured to:

pre-process the set of target video frames to obtain a set of visual characteristic values present in each frame of the set of target video frames; and

identify distributions or ranges of the set of visual characteristic values.

5. The computer system of claim 4 , wherein the set of visual characteristic values present in each frame of the set of target video frames is utilized to identify which frames of the set of target video frames form the set of keyframes.

6. The computer system of claim 4 , wherein the set of visual characteristic values present in each frame of the set of target video frames is utilized to perturb the set of modified keyframes to generate an augmented set of modified keyframes, the augmented set of modified keyframes representing an expanded set of additional modified keyframes having modified visual characteristic values generated across the ranges or distributions of the set of visual characteristic values, the augmented set of modified keyframes utilized for training the machine learning model architecture.

7. The computer system of claim 1 , wherein the visual modification is conducted across an identified region of interest in the set of target video frames, and the computer processor is configured to pre-process the set of target video frames to identify a corresponding region of interest in each frame of the set of target video frames.

8. The computer system of claim 7 , wherein the corresponding set of modified target video frames having the automatically interpolated or extrapolated visual modifications include modified frame regions of interest for combining into the set of target video frames, and wherein the corresponding region of interest in each frame of the set of target video frames is defined using a plurality of segmentation masks.

9. The computer system of claim 8 , wherein the machine learning model architecture includes a second autoencoder that is trained for identifying segmented target regions through comparing modifications in the set of modified keyframes with the corresponding frames of the set of keyframes, the second autoencoder, after training, configured to generate a new segmented target region when provided a frame of the set of target video frames; and

wherein outputs of the first autoencoder and the second autoencoder are combined together to conduct modifications of the provided frame of the set of target video frames to generate a final output frame having a modification generated by the first autoencoder applied in the new segmented target region generated by the second autoencoder.

10. The computer system of claim 1 , wherein the computer system is provided as a computing appliance coupled to a system implementing a post-production processing pipeline, and wherein the post-production processing pipeline includes manually assessing each frame of the corresponding set of modified target video frames having the automatically interpolated or extrapolated visual modifications to identify a set of incorrectly modified frames;

wherein for each frame of the set of incorrectly modified frames, a reviewer provides a corresponding revision frame; and

wherein the trained machine learning model architecture is further retrained using a combination of revision frames and a corresponding modified target video frame corresponding to each revision frame of the revision frames.

11. A computer implemented method for automatic interpolation or extrapolation of visual modifications from a set of keyframes extracted from a set of target video frames, the method comprising:

instantiating a machine learning model architecture;

receiving the set of target video frames;

identifying, from the set of target video frames, the set of keyframes;

providing the set of keyframes for visual modification by a human;

receiving a set of modified keyframes;

training the machine learning model architecture using the set of modified keyframes and the set of keyframes, the machine learning model architecture including a first autoencoder configured for unity reconstruction of the set of modified keyframes from the set of keyframes to obtain a trained machine learning model architecture; and

processing one or more frames of the set of target video frames to generate a corresponding set of modified target video frames having the automatically interpolated or extrapolated visual modifications.

12. The computer implemented method of claim 11 , wherein the visual modifications include visual effects applied to a target human being or a portion of the target human being visually represented in the set of target video frames.

13. The computer implemented method of claim 12 , wherein the visual modifications include at least one of eye-bag addition/removal, wrinkle addition/removal, or tattoo addition/removal.

14. The computer implemented method of claim 12 , wherein the method comprises:

pre-processing the set of target video frames to obtain a set of visual characteristic values present in each frame of the set of target video frames; and

identifying distributions or ranges of the set of visual characteristic values.

15. The computer implemented method of claim 14 , wherein the set of visual characteristic values present in each frame of the set of target video frames is utilized to identify which frames of the set of target video frames form the set of keyframes.

16. The computer implemented method of claim 14 , wherein the set of visual characteristic values present in each frame of the set of target video frames is utilized to perturb the set of modified keyframes to generate an augmented set of modified keyframes, the augmented set of modified keyframes representing an expanded set of additional modified keyframes having modified visual characteristic values generated across the ranges or distributions of the set of visual characteristic values, the augmented set of modified keyframes utilized for training the machine learning model architecture.

17. The computer implemented method of claim 11 , wherein the visual modification is conducted across an identified region of interest in the set of target video frames, and the method comprises configured to pre-processing the set of target video frames to identify a corresponding region of interest in each frame of the set of target video frames.

18. The computer implemented method of claim 17 , wherein the corresponding set of modified target video frames having the automatically interpolated or extrapolated visual modifications include modified frame regions of interest for combining into the set of target video frames, and wherein the corresponding region of interest in each frame of the set of target video frames is defined using a plurality of segmentation masks.

19. The computer implemented method of claim 18 , wherein the machine learning model architecture includes a second autoencoder that is trained for identifying segmented target regions through comparing modifications in the set of modified keyframes with the corresponding frames of the set of keyframes, the second autoencoder, after training, configured to generate a new segmented target region when provided a frame of the set of target video frames; and

wherein outputs of the first autoencoder and the second autoencoder are combined together to conduct modifications of the provided frame of the set of target video frames to generate a final output frame having a modification generated by the first autoencoder applied in the new segmented target region generated by the second autoencoder.

20. A non-transitory computer readable medium, storing machine-interpretable instruction sets which when executed by a processor, cause the processor to perform a method for automatic interpolation or extrapolation of visual modifications from a set of keyframes extracted from a set of target video frames, the method comprising:

instantiating a machine learning model architecture;

receiving the set of target video frames;

identifying, from the set of target video frames, the set of keyframes;

providing the set of keyframes for visual modification by a human;

receiving a set of modified keyframes;

training the machine learning model architecture using the set of modified keyframes and the set of keyframes, the machine learning model architecture including a first autoencoder configured for unity reconstruction of the set of modified keyframes from the set of keyframes to obtain a trained machine learning model architecture; and

processing one or more frames of the set of target video frames to generate a corresponding set of modified target video frames having the automatically interpolated or extrapolated visual modifications.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2023
From: PANOUSIS, MATTHEW; BIRULIN, PAUL; CHOWDHURY, DEBJOY; BADAMI, ISHRAT; SKOURIDES, ANTON; BRONFMAN, JONATHAN; MOLNAR, LON
To: MONSTERS ALIENS ROBOTS ZOMBIES INC.
Reel/Frame 065635/0688 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2023
From: DAVIES, THOMAS; MAHDAVI-AMIRI, ALI
To: MONSTERS ALIENS ROBOTS ZOMBIES INC.
Reel/Frame 065635/0764 →
Continuity (2)
Provisional Application 63161967 · Mar 16, 2021
Related Publication 20220309633A1 · Sep 29, 2022
Cited By (2)
US 12,211,183 US 12,475,537