IP Library › Granted Patent US 12,586,280
Granted Patent B2
US 12,586,280 · App. 18/396,578 · Granted Mar 24, 2026

Techniques for generating dubbed media content items

Inventors: Chao Pan (Sunnyvale, CA); Yiwei Zhao (San Jose, CA)
Assignee: NETFLIX, INC.
G06T13/205G06T13/40G06T15/04G06T15/506G06T19/20G06T2219/2004
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,280
App. No.
18/396,578
Granted
Mar 24, 2026
Kind
B2
Abstract

In various embodiments, a dubbing application performs three-dimensional (3D) tracking of (1) the face of an actor within video frames of a first media content item to generate 3D geometry representing the face of the actor, and (2) the face of a dubber within video frames of a second media content item to generate 3D geometry representing the face of the dubber. The dubbing application also tracks the texture and lighting of the face of the actor in the first media content item. The dubbing application aligns the 3D geometry of the face of the dubber with the 3D geometry of the face of the actor. Then, the dubbing application performs neural rendering to generate dubbed video frames using a trained machine learning model, the aligned 3D geometry of the dubber, the texture and lighting of the face of the actor, and the video frames of the first media content.

Claims (55)

1 . A computer-implemented method for generating dubbed media content items, the method comprising:

generating first three-dimensional (3D) geometry, a texture map, and a lighting map based on a face of an actor included in a first video frame of a first media content item;

generating second 3D geometry based on a face of a dubber included in a second video frame of a second media content item;

performing one or more operations that align the second 3D geometry with the first 3D geometry to generate an aligned second 3D geometry;

performing one or more operations to convert the lighting map to a neural lighting; and

performing one or more operations via one or more trained machine learning models to render a third video frame based on the aligned second 3D geometry, the texture map, the neural lighting, and the first video frame of the first media content item.

2 . The computer-implemented method of claim 1 , wherein generating the first 3D geometry comprises:

detecting a plurality of landmarks on the face of the actor included in the first video frame;

performing one or more operations to fit an intermediate 3D geometry to the plurality of landmarks; and

performing one or more optimization operations to update the intermediate 3D geometry based on the first video frame and one or more loss functions.

3 . The computer-implemented method of claim 2 , wherein the one or more loss functions penalize a difference between a mapping of the intermediate 3D geometry to a canonical space and one or more other mappings of one or more other 3D geometries associated with the face of the actor to the canonical space.

4 . The computer-implemented method of claim 2 , wherein the one or more loss functions penalize one or more differences between one or more landmarks on lips of the actor included in the first video frame and one or more corresponding landmarks on lips associated with the intermediate 3D geometry.

5 . The computer-implemented method of claim 2 , wherein the one or more loss functions penalize one or more differences between one or more landmarks on teeth of the actor included in the first video frame and one or more corresponding landmarks on teeth associated with the intermediate 3D geometry.

6 . The computer-implemented method of claim 2 , wherein the one or more loss functions penalize a difference between a degree to which a mouth associated with the intermediate 3D geometry is closed and a degree to which a detected mouth of the face of the actor included in the first video frame is closed.

7 . The computer-implemented method of claim 2 , wherein the one or more loss functions penalize a difference between the plurality of landmarks on the face of the actor included in the first video frame and a plurality of corresponding landmarks associated with the intermediate 3D geometry.

8 . The computer-implemented method of claim 1 , wherein the texture map and the lighting map are generated based on a loss function that penalizes a difference between the first video frame and a fourth video frame that has been rendered using the texture map and the lighting map.

9 . The computer-implemented method of claim 1 , wherein performing the one or more operations that align the second 3D geometry with the first 3D geometry comprises:

performing one or more operations to align a nose position and a mouth position associated with the second 3D geometry with a nose position and a mouth position associated with the first 3D geometry;

performing one or more operations to equalize a scale of one or more expressions associated with the second 3D geometry with a scale of one or more expressions associated with the first 3D geometry; and

performing one or more optimization operations to determine the one or more expressions associated with the second 3D geometry when combining a bottom portion of the second 3D geometry with a top portion of the first 3D geometry.

10 . The computer-implemented method of claim 1 , wherein performing the one or more operations to render the third video frame comprises:

performing one or more operations to convert the texture map to a neural texture;

and

processing, using a first trained machine learning model, the aligned second 3D geometry, the first video frame, a mask indicating one or more regions of the first video frame to be inpainted, and a combination of the neural texture and the neural lighting to generate the third video frame.

11 . One or more non-transitory computer-readable media storing instructions that, when executed by at least one processor, cause the at least one processor to perform steps comprising:

generating first three-dimensional (3D) geometry, a texture map, and a lighting map based on a face of an actor included in a first video frame of a first media content item;

generating second 3D geometry based on a face of a dubber included in a second video frame of a second media content item;

performing one or more operations that align the second 3D geometry with the first 3D geometry to generate an aligned second 3D geometry;

performing one or more operations to convert the lighting map to a neural lighting; and

performing one or more operations via one or more trained machine learning models to render a third video frame based on the aligned second 3D geometry, the texture map, the neural lighting, and the first video frame of the first media content item.

12 . The one or more non-transitory computer-readable media of claim 11 , wherein generating the first 3D geometry comprises:

detecting a plurality of landmarks on the face of the actor included in the first video frame;

performing one or more operations to fit an intermediate 3D geometry to the plurality of landmarks; and

performing one or more optimization operations to update the intermediate 3D geometry based on the first video frame and one or more loss functions.

13 . The one or more non-transitory computer-readable media of claim 12 , wherein the one or more loss functions penalize a difference between a mapping of the intermediate 3D geometry to a canonical space and one or more other mappings of one or more other 3D geometries associated with the face of the actor to the canonical space.

14 . The one or more non-transitory computer-readable media of claim 12 , wherein the one or more loss functions penalize one or more differences between one or more landmarks on at least one of lips or teeth of the actor included in the first video frame and one or more corresponding landmarks on at least one of lips or teeth associated with the intermediate 3D geometry.

15 . The one or more non-transitory computer-readable media of claim 11 , wherein performing the one or more operations that align the second 3D geometry with the first 3D geometry comprises:

performing one or more operations to align a nose position and a mouth position associated with the second 3D geometry with a nose position and a mouth position associated with the first 3D geometry;

performing one or more operations to equalize a scale of one or more expressions associated with the second 3D geometry with a scale of one or more expressions associated with the first 3D geometry; and

performing one or more optimization operations to determine the one or more expressions associated with the second 3D geometry when combining a bottom portion of the second 3D geometry with a top portion of the first 3D geometry.

16 . The one or more non-transitory computer-readable media of claim 11 , wherein performing the one or more operations to render the third video frame comprises:

performing one or more operations to convert the texture map to a neural texture;

and

processing, using a first trained machine learning model, the aligned second 3D geometry, the first video frame, a mask indicating one or more regions of the first video frame to be inpainted, and a combination of the neural texture and the neural lighting to generate the third video frame.

17 . The one or more non-transitory computer-readable media of claim 16 , wherein performing the one or more operations to render the third video frame further comprises performing one or more operations to crop and center the face of the actor in the first video frame based on the first 3D geometry.

18 . The one or more non-transitory computer-readable media of claim 16 , wherein the mask is applied to one or more feature spaces of the first trained machine learning model.

19 . The one or more non-transitory computer-readable media of claim 16 , wherein the first trained machine learning model comprises an encoder network and a decoder network.

20 . A system, comprising:

a memory storing instructions; and

a processor that is coupled to the memory and, when executing the instructions, is configured to perform the steps of:

generating first three-dimensional (3D) geometry, a texture map, and a lighting map based on a face of an actor included in a first video frame of a first media content item,

generating second 3D geometry based on a face of a dubber included in a second video frame of a second media content item,

performing one or more operations that align the second 3D geometry with the first 3D geometry to generate an aligned second 3D geometry,

performing one or more operations to convert the lighting map to a neural lighting, and

performing one or more operations via one or more trained machine learning models to render a third video frame based on the aligned second 3D geometry, the texture map, the neural lighting, and the first video frame of the first media content item.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2023
From: PAN, CHAO; ZHAO, YIWEI
To: NETFLIX, INC.
Reel/Frame 065958/0691 →
Continuity (1)
Related Publication 20250209707A1 · Jun 26, 2025
References Cited (59)
US 11176724B1 · Sinha et al. · 2021 [cited by applicant]
US 11222466B1 · Naruniec et al. · 2022 [cited by applicant]
US 11581020B1 · Hadap · 2023 [cited by examiner]
US 20220148335A1 · Liu · 2022 [cited by examiner]
US 20220351348A1 · Chae et al. · 2022 [cited by applicant]
US 20220358719A1 · Cao et al. · 2022 [cited by applicant]
US 20220375224A1 · Chae · 2022 [cited by applicant]
US 20230154090A1 · Bradley et al. · 2023 [cited by applicant]
US 20230343010A1 · Kwatra et al. · 2023 [cited by applicant]
US 20250014253A1 · Bharaj · 2025 [cited by examiner]
US 20250140257A1 · Cohen-Or · 2025 [cited by examiner]
CN 116740261A · 2023 [cited by examiner]
Wood, E. et al. (2022). “3D Face Reconstruction with Dense Landmarks”. In: Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T. (eds) Computer Vision—ECCV 2022. ECCV 2022. (Year: 2022). [cited by examiner]
Zheng, Wenzhuo et al. “Multi-view 3D Face Reconstruction Based on Flame”, Aug. 15, 2023, DO—10.48550/arXiv.2308.07551 (Year: 2023). [cited by examiner]
Zhang et al. machine translation of CN-116740261-A (Year: 2023). [cited by examiner]
Patel et al. “Visual dubbing pipeline with localized lip-sync and two-pass identity transfer”, Computers & Graphics, vol. 110, 2023, pp. 19-27 (Year: 2023). [cited by examiner]
Feng et al. “Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network” In: Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y. (eds) Computer Vision—ECCV 2018. ECCV 2018. Lecture Notes in C… [cited by examiner]
International Search Report for Application No. PCT/US2024/061554 dated Mar. 12, 2025. [cited by applicant]
International Search Report for Application No. PCT/US2024/061558 dated Mar. 3, 2025. [cited by applicant]
International Search Report for Application No. PCT/US2024/061560 dated Mar. 3, 2025. [cited by applicant]
Patel et al., “Visual Dubbing Pipeline With Localized Lip-Sync And Two-Pass Identity Transfer”, Computers & Graphics, vol. 110, DOI: 10.1016/j.cag.2022.11.005, Nov. 17, 2022, pp. 19-27. [cited by applicant]
Feng et al., “Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network”, European Computer Vision Association, Oct. 9, 2018, 18 pages. [cited by applicant]
Kim et al., “Neural Style-Preserving Visual Dubbing”, ACM Transactions on Graphics, https://doi.org/10.1145/3355089.3356500, vol. 38, No. 6, Article 178, Nov. 8, 2019, pp. 178:1-178:13. [cited by applicant]
Ma et al., “Real-Time Facial Expression Transformation for Monocular RGB Video”, DOI: 10.1111/CGF.13586, vol. 38, No. 1, Oct. 10, 2018, pp. 470-481. [cited by applicant]
Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, IEEE Transactions on Pattern Analysis And Machine Intelligence, DOI: 10.1109/TPAMI.2020.2970919vol. 43, No. 12, Dec. 2021, pp. 4… [cited by applicant]
Liang et al., “Expressive Talking Head Generation with Granular Audio-Visual Control”, Computer Vision and Pattern Recognition, 2022, pp. 3387-3396. [cited by applicant]
Zhou et al., “Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation”, Computer Vision and Pattern Recognition, 2021, pp. 4176-4186. [cited by applicant]
Tao, Ruijie, “Is someone talking? TalkNet: Audio-visual Active Speaker Detection Model”, retrieved from https://github.com/TaoRuijie/TalkNet-ASD on 2021, 5 pages. [cited by applicant]
Xu et al., “Simple And Effective Zero-Shot Cross-Lingual Phoneme Recognition”, arXiv:2109.11680, Sep. 23, 2021, 8 pages. [cited by applicant]
Python, “Face Landmark Detection”, Python, retrieved from dlib.net/face_landmark_detection.py.html on Jan. 15, 2024, 2 pages. [cited by applicant]
Github, “AV-HuBERT (Audio-Visual Hidden Unit BERT)”, retrieved from https://github.com/facebookresearch/av_hubert on Jan. 15, 2024, 5 pages. [cited by applicant]
“Google Keynote (Google I/O '23)”, retrieved from https://www.youtube.com/watch?v=cNfINi5CNbY&t=4490s on Jan. 15, 2024, 3 pages. [cited by applicant]
Thies et al., “Deferred Neural Rendering: Image Synthesis using Neural Textures”, arXiv:1904.12356, Apr. 28, 2019, pp. 1-12. [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks”, arXiv:1611.07004, Nov. 26, 2018, pp. 1-17. [cited by applicant]
Karras et al., “Training Generative Adversarial Networks with Limited Data”, 34th Conference on Neural Information Processing Systems, arXiv:2006.06676, Oct. 7, 2020, pp. 1-37. [cited by applicant]
Nvidia, “vid2vid: Pytorch Implementation of Our Method for High-Resolution Photorealistic Video-To-Video Translation”, retrieved from https://github.com/NVIDIA/vid2vid on Jan. 15, 2024, 7 pages. [cited by applicant]
Matthias, Niessner, “Why Neural Rendering is Super Cool!”, YouTube, Vision Seminar, retrieved from https://www.youtube.com/watch?v=-KGZmzP4P1l, May 19, 2020, 2 pages. [cited by applicant]
Gajarsky, Tomas, “Facetorch: Python library for analysing faces using PyTorch”, retrieved from https://github.com/tomas-gajarsky/facetorch on Jan. 15, 2024, 9 pages. [cited by applicant]
Github, “VNext Projects IDOL”, retrieved from https://github.com/wjf5203/VNext/blob/main/projects/IDOL/IDOL.md on Jan. 15, 2024, 3 pages. [cited by applicant]
Amir, “FECNet: Facial Expression Feature Extractor”, retrieved from https://github.com/AmirSh15/FECNet, Oct. 23, 2022, 3 pages. [cited by applicant]
Fox et al., “StyleVideoGAN: A Temporal Generative Model using a Pretrained StyleGAN”, arXiv:2107.07224, Nov. 30, 2021, 18 pages. [cited by applicant]
Kreuk et al., “Self-Supervised Contrastive Learning for Unsupervised Phoneme Segmentation (INTERSPEECH, 2020)”, retrieved from https://github.com/felixkreuk/UnsupSeg, 2020, 4 pages. [cited by applicant]
Xinwen, “AudioDVP: Photorealistic Audio-driven Video Portraits”, retrieved from https://github.com/xinwen-cs/AudioDVP on Jan. 15, 2024, 4 pages. [cited by applicant]
“Deep3DFaceRecon_pytorch: Accurate 3D Face Reconstruction with Weakly-Supervised Learning”, From Single Image to Image Set—PyTorch Implementation, retrieved from github.com/sicxu/Deep3DFaceRecon_pytorch, 2019, 6 pages. [cited by applicant]
Dong et al., “Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors”, retrieved from https://github.com/facebookresearch/supervision-by-registration on Jan. 15, 2024… [cited by applicant]
Wood et al., “3D Face Reconstruction with Dense Landmarks”, arXiv:2204.02776, Jul. 20, 2022, pp. 1-24. [cited by applicant]
Bolkart, Timo, “Convert from Basel Face Model (BFM) to FLAME”, retrieved from https://github.com/TimoBolkart/BFM_to_FLAME on Jan. 15, 2024, 4 pages. [cited by applicant]
Nagano et al., “paGAN: Real-time Avatars Using Dynamic Textures”, retrieved from https://vgl.ict.usc.edu/Research/paGAN/, 2018, 3 pages. [cited by applicant]
Zakharov et al., “Realistic-Neural-Talking-Head-Models”, retrieved from https://github.com/vincent-thevenin/Realistic-Neural-Talking-Head-Models on Jan. 15, 2024, 5 pages. [cited by applicant]
Zakharov et al., “Grey-eye Talking-heads”, retrieved from https://github.com/grey-eye/talking-heads on Jan. 15, 2024, 5 pages. [cited by applicant]
Github, “TencentARC/GFPGAN: GFPGAN aims at developing Practical Algorithms for Real-world Face Restoration”, retrieved from https://github.com/TencentARC/GFPGAN on Jan. 15, 2024, 5 pages. [cited by applicant]
Github, “VMAF—Video Multi-Method Assessment Fusion”, Netflix, retrieved from https://github.com/Netflix/vmaf on Jan. 15, 2024, 4 pages. [cited by applicant]
Grassal et al., “Neural Head Avatars from Monocular RGB Videos”, retrieved from https://github.com/philgras/neural-head-avatars, 2022, 6 pages. [cited by applicant]
Zhengyuf, “I M Avatar: Implicit Morphable Head Avatars from Videos”, retrieved from https://github.com/zhengyuf/IMavatar, 2022, 4 pages. [cited by applicant]
Chen et al., “Implicit Neural Head Synthesis via Controllable Local Deformation Fields”, retrieved from https://imaging.cs.cmu.edu/local_deformation_fields/ on Jan. 15, 2024, 3 pages. [cited by applicant]
Yang et al., “GAN Prior Embedded Network for Blind Face Restoration in the Wild”, retrieved from https://github.com/yangxy/GPEN on Jan. 15, 2024, 5 pages. [cited by applicant]
Non Final Office Action received for U.S. Appl. No. 18/396,590, dated Jul. 23, 2025, 36 pages. [cited by applicant]
Non Final Office Action received for U.S. Appl. No. 18/396,583, dated Dec. 12, 2025, 37 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 18/396,590, dated Feb. 11, 2026, 30 pages. [cited by applicant]