IP Library Granted Patent US 12,739,449
Granted Patent B2
US 12,739,449 · App. 18/177,897 · Granted Sep 15, 2026

Video system with object replacement and insertion features

Inventors: Shashank C. Merchant (Sunnyvale, CA); Prateek Tandon (San Jose, CA); Michael Cutter (Golden, CO); Sunil Ramesh (Cupertino, CA); Karina Levitian (Austin, TX)
Assignee: Roku, Inc.
H04N21/23412H04N21/23418H04N21/251H04N21/25883H04N21/2668H04N21/8146H04N21/854
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,739,449
App. No.
18/177,897
Granted
Sep 15, 2026
Kind
B2
Abstract

In one aspect, an example method includes (i) obtaining video that depicts an object across multiple frames of the video; (ii) detecting the object within the obtained video and determining object characteristic data associated with the detected object; (iii) determining user profile data associated with a viewer of the video; (iv) using at least the determined object characteristic data and the determined user profile data as a basis to select a replacement object from among a set of multiple candidate replacement objects; (v) replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video; and (vi) outputting for presentation the generated video.

Claims (57)

1 . A method comprising:

obtaining video that depicts an object across multiple frames of the video;

detecting the object within the obtained video and determining object characteristic data associated with the detected object;

determining user profile data associated with a viewer of the video;

determining scene attribute data associated with the obtained video;

using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects;

replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video, wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video comprises applying a lighting normalization technique to blend the selected replacement object into the video, wherein applying the lighting normalization technique to blend the selected replacement object into the video comprises determining a shape of a shadow of the selected replacement object, determining a shape of a shadow of the detected object, and using the determined shape of the shadow of the selected replacement object and the determined shape of the shadow of the detected object as a basis to modify a shadow of the detected object, wherein the detected object and the selected replacement object differ in at least one object characteristic other than scale, and wherein the shadow of the detected object and the shadow of the selected replacement object differ in at least one characteristic other than scale; and

outputting for presentation the generated video.

2 . The method of claim 1 , wherein the object characteristic data indicates a size, shape, or orientation of the detected object.

3 . The method of claim 1 , wherein detecting the object within the obtained video and determining the object characteristic data associated with the detected object comprises detecting edges and/or boundaries of the object.

4 . The method of claim 1 , wherein using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects comprises using mapping data to map the determined object characteristic data, the determined user profile data, and the determined scene attribute data to a corresponding replacement object.

5 . The method of claim 1 , wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video further comprises:

obtaining a three-dimensional model of the selected replacement object;

using the obtained three-dimensional model of the selected replacement object and the determined object characteristic data, together with a time-based affine transform model, to generate a time-based two-dimensional projection of the selected replacement object;

determining object position data associated with the detected object; and

at a position indicated by the determined object position data, replacing the detected object with the corresponding time-based two-dimensional projection of the selected replacement object.

6 . The method of claim 1 , wherein outputting for presentation, the generated video comprises transmitting to a presentation device, video data representing the generated video for display by the presentation device.

7 . A computing system configured for performing a set of acts comprising:

obtaining video that depicts an object across multiple frames of the video;

detecting the object within the obtained video and determining object characteristic data associated with the detected object;

determining user profile data associated with a viewer of the video;

determining scene attribute data associated with the obtained video;

using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects;

replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video, wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video comprises applying a lighting normalization technique to blend the selected replacement object into the video, wherein applying the lighting normalization technique to blend the selected replacement object into the video comprises determining a shape of a shadow of the selected replacement object, determining a shape of a shadow of the detected object, and using the determined shape of the shadow of the selected replacement object and the determined shape of the shadow of the detected object as a basis to modify a shadow of the detected object, wherein the detected object and the selected replacement object differ in at least one object characteristic other than scale, and wherein the shadow of the detected object and the shadow of the selected replacement object differ in at least one characteristic other than scale; and

outputting for presentation the generated video.

8 . The computing system of claim 7 , wherein the object characteristic data indicates a size, shape, or orientation of the detected object.

9 . The computing system of claim 7 , wherein detecting the object within the obtained video and determining the object characteristic data associated with the detected object comprises:

providing video data representing the obtained video to a trained model, wherein the trained model is configured to use at least video data as runtime input-data to generate object characteristic data as runtime output-data, wherein the model was trained using (i) at least one training input-data set including video data representing video depicting a training object, and (ii) at least one corresponding training output-data set including object characteristic data of the training object; and

responsive to providing the video data to the trained model, receiving from the trained model, corresponding generated object characteristic data.

10 . The computing system of claim 7 , wherein using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects comprises using mapping data to map the determined object characteristic data and the determined user profile data, and the determined scene attribute data to a corresponding replacement object.

11 . The computing system of claim 7 , wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video further comprises:

obtaining a three-dimensional model of the selected replacement object;

using the obtained three-dimensional model of the selected replacement object and the determined object characteristic data, together with a time-based affine transform model, to generate a time-based two-dimensional projection of the selected replacement object;

determining object position data associated with the detected object; and

at a position indicated by the determined object position data, replacing the detected object with the corresponding time-based two-dimensional projection of the selected replacement object.

12 . The computing system of claim 7 , wherein outputting for presentation, the generated video comprises transmitting to a presentation device, video data representing the generated video for display by the presentation device.

13 . The computing system of claim 12 , wherein the presentation device is a television.

14 . The computing system of claim 7 , wherein outputting for presentation, the generated video comprises displaying the generated video.

15 . A non-transitory computer-readable medium having stored thereon program instructions that upon execution by a computing system, cause performance of a set of acts comprising:

obtaining video that depicts an object across multiple frames of the video;

detecting the object within the obtained video and determining object characteristic data associated with the detected object;

determining user profile data associated with a viewer of the video;

determining scene attribute data associated with the obtained video;

using at least the determined object characteristic data, the determined user profile data, and the determined scene attribute data as a basis to select a replacement object from among a set of multiple candidate replacement objects;

replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video, wherein replacing the detected object with the selected replacement object to generate video that is a modified version of the obtained video comprises applying a lighting normalization technique to blend the selected replacement object into the video, wherein applying the lighting normalization technique to blend the selected replacement object into the video comprises determining a shape of a shadow of the selected replacement object, determining a shape of a shadow of the detected object, and using the determined shape of the shadow of the selected replacement object and the determined shape of the shadow of the detected object as a basis to modify a shadow of the detected object, wherein the detected object and the selected replacement object differ in at least one object characteristic other than scale, and wherein the shadow of the detected object and the shadow of the selected replacement object differ in at least one characteristic other than scale; and

outputting for presentation the generated video.

16 . The method of claim 1 , wherein the scene attribute data specifies information about one or more people in the obtained video.

17 . The method of claim 1 , wherein the scene attribute data comprises scene scale data.

18 . The method of claim 17 , further comprising determining the scene scale, wherein determining the scene scale data comprises:

using a trained model to obtain scene scale data for the scene, wherein the trained model was trained with data video data and corresponding metadata specifying information about areas and/or objects in the scene as an input data set.

19 . The method of claim 1 , wherein determining object characteristic data associated with the detected object comprises:

detecting a brand and/or model of the detected object; and

using the detected brand and/or model of the detected object to look up size and/or scale data for the detected object.

20 . The method of claim 1 , wherein determining object characteristic data associated with the detected object comprises:

detecting a brand and/or or model of the detected object;

using the detected brand and/or model of the detected object to look up multiple candidate size and/or scale data sets for the detected object; and

based on an analysis of multiple objects within the obtained video, selecting a single size and/or scale data set from among the multiple candidate size and/or scale data sets.

Assignments (2)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2023
From: MERCHANT, SHASHANK C.; TANDON, PRATEEK; CUTTER, MICHAEL; RAMESH, SUNIL; LEVITIAN, KARINA
To: ROKU, INC.
Reel/Frame 062878/0282 →
Continuity (1)
Related Publication 20240298045A1 · Sep 5, 2024
References Cited (24)
US 8910201B1 · Zamiska · 2014 [cited by examiner]
US 10613726B2 · Cohen · 2020 [cited by examiner]
US 12240445B1 · Purdy · 2025 [cited by examiner]
US 20050225553A1 · Chi · 2005 [cited by examiner]
US 20090210902A1 · Slaney · 2009 [cited by examiner]
US 20130141530A1 · Zavesky · 2013 [cited by examiner]
US 20150310307A1 · Gopalan · 2015 [cited by examiner]
US 20190279681A1 · Yuan · 2019 [cited by applicant]
US 20200242367A1 · Ludwigsen · 2020 [cited by applicant]
US 20210120286A1 · Govil · 2021 [cited by examiner]
US 20210304449A1 · Mourkogiannis · 2021 [cited by applicant]
US 20210329320A1 · Triantafyllou · 2021 [cited by applicant]
US 20220012520A1 · Mok · 2022 [cited by examiner]
US 20220327320A1 · Perincherry · 2022 [cited by examiner]
US 20230262201A1 · Gilad · 2023 [cited by examiner]
US 20240233443A1 · Rush · 2024 [cited by examiner]
JP 2003006671A · 2003 [cited by examiner]
Barron et al., “Shape, Albedo, and Illumination from a Single Image of an Unknown Object”, Proceedings/CVPR, IEEE Computer Society conference on Computer Vision and Pattern Recognition, (Jun. 2012) http://www.researchga… [cited by applicant]
Great Learning Team, “Real-Time Object Detection Using TensorFlow”, (Aug. 22, 2022) https://www.mygreatlearning.com;/blog/object-detection-using-tensorflow, retrieved Apr. 27, 2023. [cited by applicant]
Guo et al., “Neural 3D Scene Reconstruction with the Manhattan-world Assumption”, arXiv:2205.02836v2 [cs.CV] May 18, 2022. [cited by applicant]
Mildenhall et al., “NeRF: Representing Scences as Neural Radiance Fields for View Synthesis”, arXiv:2003.08934v2 [cs.CV] Aug. 3, 2020. [cited by applicant]
Lopez-Moreno et al., “Multiple Light source Estimation in a Single Image”, Computer Graphics Forum, (Aug. 20, 2013) https://onlinelibrary.wiley.com/doi/abs/10.1111/cgf.12195, retrieved Apr. 27, 2023. [cited by applicant]
Elharrouss et al., “Image inpainting: A review” (Sep. 13, 2019). [cited by applicant]
Kán et al., “DeepLight: light source estimation for augmented realtiy using deep learning”, The Visual Computer, 35:873-883 (May 7, 2019). [cited by applicant]