High resolution synthesis using shaders
Methods and systems are disclosed for generating high resolution images using lower resolution machine learning models. The system receives one or more images depicting a real-world object in a real-world scene and receives a source image depicting a fashion item comprising a target. The system processes, using one or more machine learning models, the one or more images together with the source image to generate a new image depicting the real-world object wearing the fashion item depicted in the source image, the new image having a lower image resolution than an image resolution of the source image. The system selectively blends pixels of the new image with pixels of the source image to generate a virtual extended reality (XR) experience.
1 . A method comprising:
capturing, by a camera of a user device, one or more images depicting a real-world object in a real-world scene;
receiving a source image depicting a fashion item comprising a target;
processing the source image and the one or more images via a first portion of a video processing system and a second portion of the video processing system, the first portion comprising a generative machine learning model encoder and a generative machine learning model decoder, the second portion comprising a blending shader;
receiving, by the first portion of the video processing system comprising the generative machine learning model encoder and the generative machine learning model decoder that have previously been trained based on training data, a first input and a second input together, the first input comprising the one or more images that have been captured by the camera of the user device, the second input comprising the source image;
in response to the first portion of the video processing system receiving the first input and the second input together, generating by the first portion of the video processing system an output of the generative machine learning model decoder comprising, a new image depicting the real-world object depicted in the one or more images of the first input captured by the camera of the user device wearing the fashion item depicted in the source image of the second input, the new image having a lower image resolution than an image resolution of the source image;
generating, by the first portion of the video processing system, a map that specifies at least a first correspondence between one or more first pixels in the new image and one or more second pixels in the source image and at least a second correspondence between one or more third pixels in the new image and one or more fourth pixels in the one or more images; and
after generating the output of the generative machine learning model decoder comprising the new image by the first portion of the video processing system, receiving by the second portion of the video processing system, the new image and the source image and the one or more images that were also received by the first portion of the video processing system, the second portion processing the new image, the source image, and the one or more images that were also received by the first portion of the video processing system to generate an additional image by increasing the image resolution of a portion of the new image by selectively blending at least the one or more first pixels of the new image with the one or more second pixels of the source image and at least the one or more third pixels of the new image with the one or more fourth pixels of the one or more images based on the map, the additional image being used to generate a virtual extended reality (XR) experience on the user device that was used to capture the one or more images.
2 . The method of claim 1 , further comprising:
determining a pose of the real-world object depicted in the one or more images; and
adjusting a pose of the fashion item depicted in the source image to match the pose of the real-world object depicted in the one or more images.
3 . The method of claim 2 , further comprising:
increasing in part an image resolution of the new image in response to selectively blending the one or more first pixels of the new image with the one or more second pixels of the source image.
4 . The method of claim 3 , further comprising:
identifying a region of the source image that depicts the target; and
replacing the one or more first pixels of the new image with the one or more second pixels of the source image corresponding to the identified region.
5 . The method of claim 4 , wherein the target comprises an emblem, a logo, a graphical element or text.
6 . The method of claim 1 , further comprising:
communicating with the first portion of the video processing system to determine an image resolution of the first portion of the video processing system; and
using the determined image resolution of the first portion of the video processing system to generate the source image as one of multiple images of the target posed in a way that matches a pose of the real-world object, the multiple images being simultaneously generated to depict the target posed in a way that matches the pose of the real-world object depicted in the one or more images, a first of the multiple images being generated to have a first image resolution and a second of the multiple images being generated to have a second image resolution.
7 . The method of claim 2 , further comprising:
generating a first image depicting the fashion item in the adjusted pose having a first image resolution; and
generating a second image depicting the fashion item in the adjusted pose having a second image resolution lower than the first image resolution.
8 . The method of claim 7 , further comprising:
providing the second image to the first portion of the video processing system for processing together with the one or more images.
9 . The method of claim 8 , further comprising:
selectively blending pixels of the new image with pixels of the first image to generate the XR experience.
10 . The method of claim 9 , further comprising:
selectively blending the pixels of the first image based on the map, wherein the map identifies portions of the new image corresponding to the target.
11 . The method of claim 1 , wherein the first portion of the video processing system comprise a convolutional neural network associated with a fashion item XR experience.
12 . The method of claim 11 , wherein the first portion of the video processing system are trained by performing training operations comprising:
accessing training data comprising a training pair of a training image depicting a training object and a ground truth image depicting the training object wearing a training fashion item;
analyzing, using the first portion of the video processing system, the training image to predict an image depicting the training object wearing the training fashion item;
computing a loss based on a deviation between the predicted image and the ground truth image; and
updating one or more parameters of the first portion of the video processing system based on the computed loss.
13 . The method of claim 12 , further comprising repeating the training operations for additional training data until a stopping criterion is met.
14 . The method of claim 12 , further comprising generating the training pair by performing operations comprising:
accessing a training video depicting the training object performing different poses and wearing a target fashion item;
selecting a first frame of the training video and a second frame of the training video that is adjacent to the first frame;
removing at least a portion of the target fashion item from the first frame to generate the training image; and
storing the second frame as the ground truth image corresponding to the training image, the first portion of the video processing system being trained to predict the second frame from the first frame from which the at least a portion of the target fashion item has been removed.
15 . The method of claim 1 , wherein the fashion item is a virtual object.
16 . The method of claim 1 , wherein increasing the image resolution of the portion of the new image comprises:
receiving the map and the new image from first portion of the video processing system;
determining whether a portion of pixels identified by the map corresponds to either the source image or the one or more images depicting the real-world object in the real-world scene;
in response to determining whether the portion of pixels identified by the map corresponds to the source image or the one or more images depicting the real-world object in the real-world scene:
retrieving pixel values from either the source image or the one or more images at a location matching the location of the portion of the new image, and
replacing pixels in the portion of the new image with the retrieved pixel values to increase image resolution of the portion relative to a remaining area of the new image; and
wherein the blending shader selectively blends portions of the new image with pixel values retrieved from either or both the source image and the one or more images.
17 . A system comprising:
at least one processor; and
at least one memory component having instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:
capturing, by a camera of a user device, one or more images depicting a real-world object in a real-world scene;
receiving a source image depicting a fashion item comprising a target;
processing the source image and the one or more images via a first portion of a video processing system and a second portion of the video processing system, the first portion comprising a generative machine learning model encoder and a generative machine learning model decoder, the second portion comprising a blending shader;
receiving, by the first portion of the video processing system comprising the generative machine learning model encoder and the generative machine learning model decoder that have previously been trained based on training data, a first input and a second input together, the first input comprising the one or more images that have been captured by the camera of the user device, the second input comprising the source image;
in response to the first portion of the video processing system receiving the first input and the second input together, generating by the first portion of the video processing system an output of the generative machine learning model decoder comprising, a new image depicting the real-world object depicted in the one or more images of the first input captured by the camera of the user device wearing the fashion item depicted in the source image of the second input, the new image having a lower image resolution than an image resolution of the source image;
generating, by the first portion of the video processing system, a map that specifies at least a first correspondence between one or more first pixels in the new image and one or more second pixels in the source image and at least a second correspondence between one or more third pixels in the new image and one or more fourth pixels in the one or more images; and
after generating the output of the generative machine learning model decoder comprising the new image by the first portion of the video processing system, receiving by the second portion of the video processing system, the new image and the source image and the one or more images that were also received by the first portion of the video processing system, the second portion processing the new image, the source image, and the one or more images that were also received by the first portion of the video processing system to generate an additional image by increasing the image resolution of a portion of the new image by selectively blending at least the one or more first pixels of the new image with the one or more second pixels of the source image and at least the one or more third pixels of the new image with the one or more fourth pixels of the one or more images based on the map, the additional image being used to generate a virtual extended reality (XR) experience on the user device that was used to capture the one or more images.
18 . The system of claim 17 , the operations further comprising:
receiving the map and the new image from one or more machine learning models;
determining whether a portion of pixels in the new image identified by the map corresponds to the source image;
in response to determining that the portion of pixels in the new image corresponds to the source image:
retrieving pixel values from the source image at a location matching a location of the portion in the new image, and
replacing pixels in the portion of the new image with the retrieved pixel values from the source image to increase image resolution of the portion of the new image relative to a remaining area of the new image.
19 . A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:
capturing, by a camera of a user device, one or more images depicting a real-world object in a real-world scene;
receiving a source image depicting a fashion item comprising a target;
processing the source image and the one or more images via a first portion of a video processing system and a second portion of the video processing system, the first portion comprising a generative machine learning model encoder and a generative machine learning model decoder, the second portion comprising a blending shader;
receiving, by the first portion of the video processing system comprising the generative machine learning model encoder and the generative machine learning model decoder that have previously been trained based on training data, a first input and a second input together, the first input comprising the one or more images that have been captured by the camera of the user device, the second input comprising the source image;
in response to the first portion of the video processing system receiving the first input and the second input together, generating by the first portion of the video processing system an output of the generative machine learning model decoder comprising, a new image depicting the real-world object depicted in the one or more images of the first input captured by the camera of the user device wearing the fashion item depicted in the source image of the second input, the new image having a lower image resolution than an image resolution of the source image;
generating, by the first portion of the video processing system, a map that specifies at least a first correspondence between one or more first pixels in the new image and one or more second pixels in the source image and at least a second correspondence between one or more third pixels in the new image and one or more fourth pixels in the one or more images; and
after generating the output of the generative machine learning model decoder comprising the new image by the first portion of the video processing system, receiving by the second portion of the video processing system, the new image and the source image and the one or more images that were also received by the first portion of the video processing system, the second portion processing the new image, the source image, and the one or more images that were also received by the first portion of the video processing system to generate an additional image by increasing the image resolution of a portion of the new image by selectively blending at least the one or more first pixels of the new image with the one or more second pixels of the source image and at least the one or more third pixels of the new image with the one or more fourth pixels of the one or more images based on the map, the additional image being used to generate a virtual extended reality (XR) experience on the user device that was used to capture the one or more images.
20 . The non-transitory computer-readable storage medium of claim 19 , the operations further comprising generating the training data by performing training operations comprising:
accessing a training video that depicts an individual object wearing a training fashion item;
generating a pair of training images extracted from the training video, a first training image of the pair of training images depicting the individual object wearing the training fashion item, a second training image of the pair of training images depicting the individual object having a portion of the training fashion item removed, the first portion of the video processing system comprising the generative machine learning model encoder and the generative machine learning model decoder trained to predict the first training image depicting the individual object wearing the training fashion item, extracted from the training video, from the second training image of the training video depicting the individual object having the portion of the training fashion item removed which has also been extracted from the training video.