IP Library › Granted Patent US 12,260,882
Granted Patent B2
US 12,260,882 · App. 18/666,243 · Granted Mar 25, 2025

Actor-replacement system for videos

Inventors: Sunil Ramesh (Cupertino, CA); Michael Cutter (Golden, CO); Karina Levitian (Austin, TX)
Assignee: Roku, Inc.
G11B27/036G06T7/00G06T2207/10016G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,882
App. No.
18/666,243
Granted
Mar 25, 2025
Kind
B2
Abstract

In one aspect, an example method includes (i) estimating, using a skeletal detection model, a pose of an original actor for each of multiple frames of a video; (ii) obtaining, for each of a plurality of the estimated poses, a respective image of a replacement actor; (iii) obtaining replacement speech in the replacement actor's voice that corresponds to speech of the original actor in the video; (iv) generating, using the estimated poses, the images of the replacement actor, and the replacement speech, synthetic frames corresponding to the multiple frames of the video that depict the replacement actor in place of the original actor, with the synthetic frames including facial expressions for the replacement actor that temporally align with the replacement speech; and (iv) combining the synthetic frames and the replacement speech so as to obtain a synthetic video that replaces the original actor with the replacement actor.

Claims (38)

1. A computing system comprising a processor and a non-transitory computer-readable medium having stored thereon program instructions that upon execution by the processor, cause performance of a set of acts comprising:

estimating, using a skeletal detection model, a pose of an original actor for each of multiple frames of a video;

obtaining, for each of a plurality of the estimated poses of the original actor, a respective image of a modified version of the original actor;

generating, using the estimated poses and the images of the modified version of the original actor, synthetic frames corresponding to the multiple frames of the video that depict the modified version of the original actor in place of the original actor, wherein the synthetic frames depict the modified version of the original actor in respective poses that align with the estimated poses of the original actor in corresponding frames of the video, and wherein the synthetic frames comprise facial expressions for the modified version of the original actor that temporally align with corresponding speech, wherein generating the synthetic frames comprises, for a given frame of the multiple frames, inserting, using an object insertion model, an image of the modified version of the original actor into the given frame at a location indicated by the estimated pose of the original actor so as to obtain a modified frame, and wherein generating the synthetic frames further comprises providing the corresponding speech and the modified frame as input to a temporal generative adversarial network having an ensemble of discriminators; and

combining the synthetic frames and the corresponding speech so as to obtain a synthetic video that replaces the original actor with the modified version of the original actor.

2. The computing system of claim 1 , wherein obtaining the images of the modified version of the original actor comprises estimating a pose of the modified version of the original actor for each of multiple frames of a sample video of the modified version of the original actor.

3. The computing system of claim 1 , wherein obtaining the images of the modified version of the original actor comprises:

recording a sample video of the modified version of the original actor; and

identifying poses of the modified version of the original actor during the sample video using a motion capture system.

4. The computing system of claim 1 , wherein obtaining the images of the modified version of the original actor comprises rendering an image of the modified version of the original actor using a target pose, an image of the modified version of the original actor, and a pose rendering model.

5. The computing system of claim 1 , wherein a difference between the original actor and the modified version of the original actor is in accordance with a storyline change.

6. The computing system of claim 1 , wherein generating the synthetic frames further comprises, for the given frame, extracting the original actor from the given frame using the estimated poses of the original actor.

7. The computing system of claim 1 , wherein the ensemble of discriminators comprises a frame discriminator, a sequence discriminator, and a synchronization discriminator.

8. A method comprising:

estimating, by a computing system using a skeletal detection model, a pose of an original actor for each of multiple frames of a video;

obtaining, by the computing system for each of a plurality of the estimated poses of the original actor, a respective image of a modified version of the original actor;

generating, by the computing system using the estimated poses and the images of the modified version of the original actor, synthetic frames corresponding to the multiple frames of the video that depict the modified version of the original actor in place of the original actor, wherein the synthetic frames depict the modified version of the original actor in respective poses that align with the estimated poses of the original actor in corresponding frames of the video, and wherein the synthetic frames comprise facial expressions for the modified version of the original actor that temporally align with corresponding speech, wherein generating the synthetic frames comprises, for a given frame of the multiple frames, inserting, using an object insertion model, an image of the modified version of the original actor into the given frame at a location indicated by the estimated pose of the original actor so as to obtain a modified frame, and wherein generating the synthetic frames further comprises providing the corresponding speech and the modified frame as input to a temporal generative adversarial network having an ensemble of discriminators; and

combining the synthetic frames and the corresponding speech so as to obtain a synthetic video that replaces the original actor with the modified version of the original actor.

9. The method of claim 8 , wherein obtaining the images of the modified version of the original actor comprises estimating a pose of the modified version of the original actor for each of multiple frames of a sample video of the modified version of the original actor.

10. The method of claim 8 , wherein obtaining the images of the modified version of the original actor comprises:

recording a sample video of the modified version of the original actor; and

identifying poses of the modified version of the original actor during the sample video using a motion capture system.

11. The method of claim 8 , wherein obtaining the images of the modified version of the original actor comprises rendering an image of the modified version of the original actor using a target pose, an image of the modified version of the original actor, and a pose rendering model.

12. The method of claim 8 , wherein a difference between the original actor and the modified version of the original actor is in accordance with a storyline change.

13. The method of claim 8 , wherein generating the synthetic frames further comprises, for the given frame, extracting the original actor from the given frame using the estimated poses of the original actor.

14. The method of claim 8 , wherein the ensemble of discriminators comprises a frame discriminator, a sequence discriminator, and a synchronization discriminator.

15. A non-transitory computer-readable medium having stored thereon program instructions that upon execution by a computing system, cause performance of a set of acts comprising:

estimating, using a skeletal detection model, a pose of an original actor for each of multiple frames of a video;

obtaining, for each of a plurality of the estimated poses of the original actor, a respective image of a modified version of the original actor;

generating, using the estimated poses and the images of the modified version of the original actor, synthetic frames corresponding to the multiple frames of the video that depict the modified version of the original actor in place of the original actor, wherein the synthetic frames depict the modified version of the original actor in respective poses that align with the estimated poses of the original actor in corresponding frames of the video, and wherein the synthetic frames comprise facial expressions for the modified version of the original actor that temporally align with corresponding speech, wherein generating the synthetic frames comprises, for a given frame of the multiple frames, inserting, using an object insertion model, an image of the modified version of the original actor into the given frame at a location indicated by the estimated pose of the original actor so as to obtain a modified frame, and wherein generating the synthetic frames further comprises providing the corresponding speech and the modified frame as input to a temporal generative adversarial network having an ensemble of discriminators; and

combining the synthetic frames and the corresponding speech so as to obtain a synthetic video that replaces the original actor with the modified version of the original actor.

16. The non-transitory computer-readable medium of claim 15 , wherein obtaining the images of the modified version of the original actor comprises estimating a pose of the modified version of the original actor for each of multiple frames of a sample video of the modified version of the original actor.

17. The non-transitory computer-readable medium of claim 15 , wherein obtaining the images of the modified version of the original actor comprises:

recording a sample video of the modified version of the original actor; and

identifying poses of the modified version of the original actor during the sample video using a motion capture system.

18. The non-transitory computer-readable medium of claim 15 , wherein obtaining the images of the modified version of the original actor comprises rendering an image of the modified version of the original actor using a target pose, an image of the modified version of the original actor, and a pose rendering model.

19. The non-transitory computer-readable medium of claim 15 , wherein a difference between the original actor and the modified version of the original actor is in accordance with a storyline change.

20. The non-transitory computer-readable medium of claim 15 , wherein generating the synthetic frames further comprises, for the given frame, extracting the original actor from the given frame using the estimated poses of the original actor.

Assignments (2)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2024
From: RAMESH, SUNIL; CUTTER, MICHAEL; LEVITIAN, KARINA
To: ROKU, INC.
Reel/Frame 067438/0354 →
Continuity (3)
Continuation 18349551 · Jul 10, 2023
Continuation 18062410 · Dec 6, 2022
Related Publication 20240304219A1 · Sep 12, 2024
References Cited (24)
US 11582519B1 · Bhat · 2023 [cited by applicant]
US 20130330060A1 · Seidel · 2013 [cited by applicant]
US 20180025750A1 · Smith · 2018 [cited by applicant]
US 20200084521A1 · Adler · 2020 [cited by applicant]
KR 20080002291A · 2008 [cited by applicant]
John Son; DeepBrain AI to Debut AI Studios at the 2022 NAB Show; Apr. 6, 2022, 11:00 ET; 4 pgs; https://www.prnewswire.com/news/deepbrain-ai/. [cited by applicant]
Pose Detection | ML Kit | Google Developers; Pose Detection; Sep. 20, 2022; https://developers.google.com/ml-kit/vision/ pose-detection. [cited by applicant]
Samsung's Neon ‘artificial humans’ are confusing everyone. We set the record straight; Aug. 23, 2022, 8:36 AM; pp. 1/16; https://www.cnet.com/tech/mobile/samsung-neon-artificial-humans-are-confusing-everyone-we-set-reco… [cited by applicant]
Synthesia | #1 AI Video Generation Platform; Create professional videos in 60+ languages; Aug. 23, 2022, 8:33 AM; p. 1/16; https://www.synthesia.io. [cited by applicant]
Vougioukas et al.; Realistic Speech-Driven Facial Animation with GANs; International Journal of Computer Vision (2020) 128:1398-1413; https://doi.org/10.1007/s11263-019-01251-8. [cited by applicant]
Sarkar et al., Neural Re-Rendering of Humans from a Singe Image, gvv.mpi-inf.mpg.de/projects/NHRR/. [cited by applicant]
Facebook's AI convincingly inserts people into photos | VentureBeat, Sep. 22, 2022, https://venturebeat.com/ai/facebooks-ai-convincingly-inserts-people-into-photos/. [cited by applicant]
Lee et al., Context-Aware Synthesis and Placement of Object Instances, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montréal, Canada. [cited by applicant]
Albadawy et al., Voice Conversion Using Speech-to-Speech Neuro-Style Transfer, Interspeech 2020, Oct. 25-29, 2020, Shanghai, China; pp. 4726-4730; http://dx.doi.org/10.21437/Interspeech.2020-3056. [cited by applicant]
Zhang et al., Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning, Google, arXiv:1907. Jul. 24, 2019 04448v2 [cs.CL] Jul. 24, 2019. [cited by applicant]
Ho et al., Cross-Lingual Voice Conversion With Controllable Speaker Individuality Using Variational Autoencoder and Star Generative Adversarial Network, date of publication Mar. 2, 2021, date of current version Apr. 1, … [cited by applicant]
What is Cross-Language Voice Conversion and Why it's Important, reSpeecher, May 9, 2022 10:17:37 AM, downloaded Sep. 23, 2022, https://www.respeecher.com/blog/what-is-cross-language-voice-conversion-important. [cited by applicant]
Gafni et al., Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction: CVPR 2021 open access, downloaded Dec. 6, 2022, https://openaccess.thecvf.com/content/CVPR2021/html/Gafni_Dynamic_Neural_Radian… [cited by applicant]
Elharrouss et al., “Image inpainting: A review”, Department of Computer Science and Engineering, Qatar University, Doha, Qatar (2019). [cited by applicant]
Wang et al., “Few-shot Video-to-Video Synthesis”, arXiv (Cornell University), Oct. 28, 2019, https://arixiv.org/pdf/1910.12713.pdf, retrieved Dec. 2, 2023. [cited by applicant]
Gururani et al., “SPACEx: Speech driven Portrait Animation with Controllable Expression”, arXiv (Cornell University), Nov. 17, 2022, http://arxiv.org/pdf/2211.09809v1.pdf, retrieved Dec. 2, 2023. [cited by applicant]
Mirsky et al., “The Creation and Detection of Deepfakes: A Survey”, ACM Computing Surveys, ACM, New York, NY, 54(1):1-41, Jan. 2, 2021. [cited by applicant]
Song, et al., “Talking Face Generation by Conditional Recurrent Adversarial Network”, Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, pp. 919-925, Jul. 25, 2019, http://arxiv.… [cited by applicant]
Wang, “Few-shot Video-to-Video Synthesis (NeurIPS 2019)”, Oct. 31, 2019, https://www.youtube/watch?v=8AZBuyEuDqc, retrieved Dec. 2, 2023. [cited by applicant]
Cited By (1)
US 12,739,479