Video system with object replacement and insertion features
In one aspect, an example method includes (i) obtaining video that depicts an object across multiple frames of the video; (ii) detecting the object within the obtained video and determining object characteristic data associated with the detected object; (iii) determining user profile data associated with a viewer of the video; (iv) using at least the determined object characteristic data and the determined user profile data as a basis to select a replacement object from among a set of multiple candidate replacement objects; (v) replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video; and (vi) outputting for presentation the generated video.
1 . A method comprising:
obtaining video that depicts an object across multiple frames of the video;
detecting the object within the obtained video and determining object characteristic data associated with the detected object;
determining user profile data associated with a viewer of the video;
determining scene attribute data associated with the obtained video;
using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects;
replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video, wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video comprises applying a lighting normalization technique to blend the selected replacement object into the video, wherein applying the lighting normalization technique to blend the selected replacement object into the video comprises determining a shape of a shadow of the selected replacement object, determining a shape of a shadow of the detected object, and using the determined shape of the shadow of the selected replacement object and the determined shape of the shadow of the detected object as a basis to modify a shadow of the detected object, wherein the detected object and the selected replacement object differ in at least one object characteristic other than scale, and wherein the shadow of the detected object and the shadow of the selected replacement object differ in at least one characteristic other than scale; and
outputting for presentation the generated video.
2 . The method of claim 1 , wherein the object characteristic data indicates a size, shape, or orientation of the detected object.
3 . The method of claim 1 , wherein detecting the object within the obtained video and determining the object characteristic data associated with the detected object comprises detecting edges and/or boundaries of the object.
4 . The method of claim 1 , wherein using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects comprises using mapping data to map the determined object characteristic data, the determined user profile data, and the determined scene attribute data to a corresponding replacement object.
5 . The method of claim 1 , wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video further comprises:
obtaining a three-dimensional model of the selected replacement object;
using the obtained three-dimensional model of the selected replacement object and the determined object characteristic data, together with a time-based affine transform model, to generate a time-based two-dimensional projection of the selected replacement object;
determining object position data associated with the detected object; and
at a position indicated by the determined object position data, replacing the detected object with the corresponding time-based two-dimensional projection of the selected replacement object.
6 . The method of claim 1 , wherein outputting for presentation, the generated video comprises transmitting to a presentation device, video data representing the generated video for display by the presentation device.
7 . A computing system configured for performing a set of acts comprising:
obtaining video that depicts an object across multiple frames of the video;
detecting the object within the obtained video and determining object characteristic data associated with the detected object;
determining user profile data associated with a viewer of the video;
determining scene attribute data associated with the obtained video;
using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects;
replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video, wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video comprises applying a lighting normalization technique to blend the selected replacement object into the video, wherein applying the lighting normalization technique to blend the selected replacement object into the video comprises determining a shape of a shadow of the selected replacement object, determining a shape of a shadow of the detected object, and using the determined shape of the shadow of the selected replacement object and the determined shape of the shadow of the detected object as a basis to modify a shadow of the detected object, wherein the detected object and the selected replacement object differ in at least one object characteristic other than scale, and wherein the shadow of the detected object and the shadow of the selected replacement object differ in at least one characteristic other than scale; and
outputting for presentation the generated video.
8 . The computing system of claim 7 , wherein the object characteristic data indicates a size, shape, or orientation of the detected object.
9 . The computing system of claim 7 , wherein detecting the object within the obtained video and determining the object characteristic data associated with the detected object comprises:
providing video data representing the obtained video to a trained model, wherein the trained model is configured to use at least video data as runtime input-data to generate object characteristic data as runtime output-data, wherein the model was trained using (i) at least one training input-data set including video data representing video depicting a training object, and (ii) at least one corresponding training output-data set including object characteristic data of the training object; and
responsive to providing the video data to the trained model, receiving from the trained model, corresponding generated object characteristic data.
10 . The computing system of claim 7 , wherein using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects comprises using mapping data to map the determined object characteristic data and the determined user profile data, and the determined scene attribute data to a corresponding replacement object.
11 . The computing system of claim 7 , wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video further comprises:
obtaining a three-dimensional model of the selected replacement object;
using the obtained three-dimensional model of the selected replacement object and the determined object characteristic data, together with a time-based affine transform model, to generate a time-based two-dimensional projection of the selected replacement object;
determining object position data associated with the detected object; and
at a position indicated by the determined object position data, replacing the detected object with the corresponding time-based two-dimensional projection of the selected replacement object.
12 . The computing system of claim 7 , wherein outputting for presentation, the generated video comprises transmitting to a presentation device, video data representing the generated video for display by the presentation device.
13 . The computing system of claim 12 , wherein the presentation device is a television.
14 . The computing system of claim 7 , wherein outputting for presentation, the generated video comprises displaying the generated video.
15 . A non-transitory computer-readable medium having stored thereon program instructions that upon execution by a computing system, cause performance of a set of acts comprising:
obtaining video that depicts an object across multiple frames of the video;
detecting the object within the obtained video and determining object characteristic data associated with the detected object;
determining user profile data associated with a viewer of the video;
determining scene attribute data associated with the obtained video;
using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects;
replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video, wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video comprises applying a lighting normalization technique to blend the selected replacement object into the video, wherein applying the lighting normalization technique to blend the selected replacement object into the video comprises determining a shape of a shadow of the selected replacement object, determining a shape of a shadow of the detected object, and using the determined shape of the shadow of the selected replacement object and the determined shape of the shadow of the detected object as a basis to modify a shadow of the detected object, wherein the detected object and the selected replacement object differ in at least one object characteristic other than scale, and wherein the shadow of the detected object and the shadow of the selected replacement object differ in at least one characteristic other than scale; and
outputting for presentation the generated video.
16 . The method of claim 1 , wherein the scene attribute data specifies information about one or more people in the obtained video.
17 . The method of claim 1 , wherein the scene attribute data comprises scene scale data.
18 . The method of claim 17 , further comprising determining the scene scale, wherein determining the scene scale data comprises:
using a trained model to obtain scene scale data for the scene, wherein the trained model was trained with data video data and corresponding metadata specifying information about areas and/or objects in the scene as an input data set.
19 . The method of claim 1 , wherein determining object characteristic data associated with the detected object comprises:
detecting a brand and/or model of the detected object; and
using the detected brand and/or model of the detected object to look up size and/or scale data for the detected object.
20 . The method of claim 1 , wherein determining object characteristic data associated with the detected object comprises:
detecting a brand and/or or model of the detected object;
using the detected brand and/or model of the detected object to look up multiple candidate size and/or scale data sets for the detected object; and
based on an analysis of multiple objects within the obtained video, selecting a single size and/or scale data set from among the multiple candidate size and/or scale data sets.