Single image manipulation
Embodiments relate to a system which includes a training generator configured to: slice an input source file into source file time series crops, slice an input source reference file into source reference file time series crops corresponding to the source file time series crops. The system includes a dynamic flow detector configured to: determine a first flow output and a second flow output. The system includes an identity detector configured to: determine a first identity output and a second identity output. The system includes a generative preprocessor to generate a preprocessor output and a target crop generator configured to receive the preprocessor output, apply the preprocessor output to generate a target crop, adjust a vector of the target crop generator to minimize a loss metric, and apply a motion extracted from a new source file to a new target file.
1 . A method of manipulating digital media, comprising:
receiving a target media file, wherein the target media file is a still image that includes a target subject;
receiving a source media file, wherein the source media file is a video that includes a source subject;
cropping the target media file to produce a series of cropped target media files depicting the target subject in an unaltered state;
cropping the source media file to produce a series of cropped source media files;
for each cropped source media file, identifying a predetermined identifiable attribute of the input source subject;
determining, by an attribute detector, changes in the identifiable attribute of the source subject between the series of cropped source media files and storing the determined changes as computer-readable data;
providing at least a generator from a transformation generative adversarial network (GAN), wherein the transformation GAN has been trained on at least one reference subject, and the target subject and the source subject are not intentionally the same as the at least one reference subject on which the transformation GAN has been trained;
providing the computer-readable data of the changes in the identifiable attribute of the source subject between the series of cropped source media files to a generative processor neural network and the generator;
providing at least one of the cropped target media files to the generative preprocessor neural network;
the generative preprocessor neural network producing a preprocessor output that includes at least one of: (i) a vector flow image corresponding to a flow of each pixel of at least one cropped target media file, and (ii) a grayscale missing pixel image identifying pixels of at least one cropped target media file to be replaced;
providing the preprocessor output to the generator,
the generator producing a transformed media file based at least in part on the preprocessor output, wherein the transformed media file includes the source video with the source subject replaced with the target subject, wherein a corresponding set of identifiable attributes of the target subject are altered based on the changes in the identifiable attribute of the source subject between the series of cropped source media files; and
outputting a final transformed media file.
2 . The method of claim 1 , wherein the attribute detector is a dynamic flow detector configured to detect movement of at least a portion of the source subject between two or more cropped source media files.
3 . The method of claim 1 , wherein the attribute detector is an identity detector configured to detect changes in identity characteristics of the source subject between two or more cropped source media files.
4 . The method of claim 1 , wherein:
identifying the predetermined identifiable attribute includes identity characteristics and movement of at least a portion of the source subject between two or more cropped source media files; and
the attribute detector includes both:
an identity detector configured to detect changes in the identity characteristics of the source subject between two or more cropped source media files; and
a dynamic flow detector configured to detect the movement of at least the portion of the source subject between two or more cropped source media files.
5 . The method of claim 1 , wherein the source subject and the target subject are different.
6 . The method of claim 1 , wherein the transformation GAN has been trained on at least 50,000 media files and in comparing the transformed media file to the source media file during training, a percent error between one or more features of the target subject in the transformed media file and the source subject in the source media file was less than or equal to 20% error.
7 . The method of claim 1 , wherein the transformation GAN has not previously trained on the target subject.
8 . The method of claim 1 , further including creating reference files for each cropped source media files, wherein each reference file corresponds to an identifiable attribute and the step of determining changes in the identifiable attribute of the source subject is performed on each reference file.
9 . The method of claim 1 , further including analyzing the cropped source media files and determining which of the cropped source media files includes the source subject most similarly aligned to the alignment of the target subject in the target media file.
10 . The method of claim 9 , further including analyzing the cropped source media files and determining which of the cropped source media files includes the source subject with the most similar facial expression to the target subject in the target media file.
11 . The method of claim 10 , further including identifying the cropped source media file with the most similar subject alignment and facial expression as a starter image to be provided to the generator.
12 . The method of claim 1 , further including aligning the cropped source media files so that a face of the source subject is forward facing in each of the cropped source media files.
13 . The method of claim 12 , further including further transforming the final transformed media file to reverse the aligning that was performed on the cropped source media files.
14 . A media transformation system, comprising:
an input source, the input source configured to provide a target media file and a source media file, wherein the target media file includes a target subject, and the source media file includes a source subject;
wherein the target media file is a still image and the s ace media file is a video;
a cropping module configured to crop the source media file to produce a series of cropped source media files and crop the target media file to produce a series of cropped target media files depicting the target subject in an unaltered state;
a first attribute detector configured to identify changes in identity characteristics of the source subject between two or more cropped source media files and store the determined changes as computer-readable data;
a generator from a transformation generative adversarial network (GAN), wherein the transformation GAN has been trained on at least one reference subject and the target subject and the source subject are not intentionally the same as the at least one reference subject on which the transformation GAN has been trained;
a generative preprocessor neural network configured to generate a preprocessor output based at least in part on the computer-readable data and the series of cropped target media files, the preprocessor output including at least one of: (i) a vector flow image corresponding to a flow of each pixel of at least one cropped target media file, and (ii) a grayscale missing pixel image identifying pixels of at least one cropped target media file to be replaced,
the generator configured to receive the computer-readable data of the changes in the identifiable attribute of the source subject between the series of cropped input source media files, and to receive the series of cropped target media files depicting the target subject in the unaltered state;
the generator further configured to receive the preprocessor output;
the generator configured to produce a transformed media file based at least in part on the preprocessor output, wherein the transformed media file includes the source video with the source subject replaced with the target subject, wherein a corresponding set of identifiable attributes of the target subject are altered based on the changes in the identifiable attribute of the source subject between the series of cropped source media files;
wherein the system outputs a final transformed media file to a user.
15 . The system of claim 14 , further including a second attribute detector in the form of a dynamic flow detector configured to detect movement of at least a portion of the source subject between two or more cropped source media files.
16 . The system of claim 14 , further includes:
a reference file generator configured to create reference files for each cropped source media files, wherein each reference file corresponds to an identifiable attribute; and
at least a second attribute detector configured to determine changes in at least a second identifiable attribute of the source subject.
17 . The system of claim 14 , further including an alignment processor configured to align the cropped source media files so that a face of the source subject is forward facing in each of the cropped source media files.
18 . A method of manipulating digital media, comprising:
receiving a target media file, wherein the target media file includes a target subject;
cropping the target media file to produce a series of cropped target media files depicting the target subject in an unaltered state;
receiving a source media file, wherein the source media file includes a source subject that is different from the target subject;
wherein the source media file is a video;
cropping the source media file to produce a series of cropped source media files;
for each cropped source media file, identifying a predetermined identifiable attribute of the input source subject;
identifying changes in identity characteristics of the source subject between two or more cropped source media files and storing the determined changes in identity characteristics as computer-readable data;
detecting movement of at least a portion of the source subject between two or more cropped source media files and storing the determined changes in movement as computer-readable data;
providing at least a generator from a transformation generative adversarial network (GAN), wherein the transformation GAN has been trained on at least one reference subject and the target subject and the source subject are not intentionally the same as the at least one reference subject on which the transformation GAN has been trained;
providing the computer-readable data of the changes in the identity characteristics to the generator;
providing the computer-readable data of the changes in movement to the generator;
providing the series of cropped target media files to the generator;
providing the computer-readable data of the changes in the identity characteristics, the computer-readable data of the changes in movement, and the series of cropped target media files to a generative preprocessor neural network configured to generate a preprocessor output that includes at least one of: (i) a vector flow image corresponding to a flow of each pixel of at least one cropped target media file, and (ii) a grayscale missing pixel image identifying pixels of at least one cropped target media file to be replaced,
providing the preprocessor output to the generator;
producing, by the generator, a transformed media file based at least in part on the preprocessor output and the provided data corresponding to changes in movement and identity characteristics, wherein the transformed media file includes the source video with the source subject replaced with the target subject, wherein a corresponding set of identifiable attributes of the target subject are altered based on the changes in the identifiable attribute of the source subject between the series of cropped source media files; and
outputting a final transformed media file to a user.