IP Library Granted Patent US 12,361,621
Granted Patent B2
US 12,361,621 · App. 17/967,872 · Granted Jul 15, 2025

Creating images, meshes, and talking animations from mouth shape data

Inventors: Siddarth Ravichandran (Santa Clara, CA); Anthony Sylvain Jean-Yves Liot (San Jose, CA); Dimitar Petkov Dinev (Sunnyvale, CA); Ondrej Texler (San Jose, CA); Hyun Jae Kang (Mountain View, CA); Janvi Chetan Palan (San Francisco, CA); Sajid Sadi (San Jose, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T13/40G06T7/70G06T17/20G10L15/25G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,621
App. No.
17/967,872
Granted
Jul 15, 2025
Kind
B2
Abstract

Creating images and animations of lip motion from mouth shape data includes providing, as one or more input features to a neural network model, a vector of a plurality of coefficients. Each vector of the plurality of coefficients corresponds to a different mouth shape. Using the neural network model, a data structure output specifying a visual representation of a mouth including lips having a shape corresponding to the vector is generated.

Claims (42)

1. A computer-implemented method, comprising:

providing, as an input feature to a neural network model, a vector of a plurality of coefficients, wherein each vector of the plurality of coefficients corresponds to a different mouth shape;

providing, as a further input feature to the neural network model, head pose information; and

generating, using the neural network model operating only on 3-dimensional data consisting of the vector of the plurality of coefficients and the head pose information, a data structure output specifying a visual representation of a mouth including lips having a shape corresponding to the vector.

2. The computer-implemented method of claim 1 , wherein the plurality of coefficients of each vector are viseme coefficients.

3. The computer-implemented method of claim 2 , further comprising:

generating the viseme coefficients from audio data of speech.

4. The computer-implemented method of claim 2 , further comprising:

encoding the viseme coefficients into the vector provided to the neural network model.

5. The computer-implemented method of claim 1 , wherein the data structure output is generated without using any facial landmark data.

6. The computer-implemented method of claim 1 , wherein the data structure output is a Red-Green-Blue (RGB) 2-dimensional image.

7. The computer-implemented method of claim 1 , wherein the data structure output is a mesh.

8. The computer-implemented method of claim 1 , wherein the generating the data structure output includes:

generating, using the neural network model, a plurality of images, wherein each image is generated from a different vector having a plurality of coefficients, and wherein the plurality of images collectively represent an animation of lip motion.

9. The computer-implemented method of claim 1 , further comprising:

providing to the neural network model, with the vector of the plurality of coefficients, a first mesh of a face;

wherein the data structure output is a second mesh generated by the neural network model based on the vector and the first mesh of the face, wherein the second mesh is a version of the first mesh of the face that specifies the mouth including the lips having the shape corresponding to the vector.

10. A system, comprising:

one or more processors configured to initiate operations including:

providing, as an input feature to a neural network model, a vector of a plurality of coefficients, wherein each vector of the plurality of coefficients corresponds to a different mouth shape;

providing, as a further input feature to the neural network model, head pose information; and

generating, using the neural network model operating only on 3-dimensional data consisting of the vector of the plurality of coefficients and the head pose information, a data structure output specifying a visual representation of a mouth including lips having a shape corresponding to the vector.

11. The system of claim 10 , wherein the plurality of coefficients of each vector are viseme coefficients.

12. The system of claim 11 , wherein the one or more processors are configured to initiate operations including:

generating the viseme coefficients from audio data of speech.

13. The system of claim 11 , wherein the one or more processors are configured to initiate operations including:

encoding the viseme coefficients into the vector provided to the neural network model.

14. The system of claim 10 , wherein the data structure output is generated without using any facial landmark data.

15. The system of claim 10 , wherein the data structure output is a Red-Green-Blue (RGB) 2-dimensional image.

16. The system of claim 10 , wherein the data structure output is a mesh.

17. The system of claim 10 , wherein the generating the data structure output includes:

generating, using the neural network model, a plurality of images, wherein each image is generated from a different vector having a plurality of coefficients, and wherein the plurality of images collectively represent an animation of lip motion.

18. The system of claim 10 , wherein the one or more processors are configured to initiate operations including:

providing to the neural network model, with the vector of the one or more coefficients, a first mesh of a face;

wherein the data structure output is a second mesh generated by the neural network model based on the vector and the first mesh of the face, wherein the second mesh is a version of the first mesh of the face that specifies the mouth including the lips having the shape corresponding to the vector.

19. A computer program product, comprising:

a computer readable storage medium, and program instructions collectively stored on the computer readable storage medium, wherein the program instructions are executable by one or more processors to initiate operations including:

providing, as an input feature to a neural network model, a vector of a plurality of coefficients, wherein each vector of the plurality of coefficients corresponds to a different mouth shape;

providing, as a further input feature to the neural network model, a first mesh of a head and face having a mouth region of the first mesh masked so as to specify no mouth or lip shape information;

wherein the vector of the plurality of coefficients and the first mesh information specify 3-dimensional data; and

generating, using the neural network model operating only on the 3-dimensional data, a data structure output specifying a second mesh adapted from the first mesh, wherein the second mesh includes a mouth region including lips having a shape corresponding to the vector.

20. The computer program product of claim 19 , wherein the plurality of coefficients of each vector are viseme coefficients.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2022
From: RAVICHANDRAN, SIDDARTH; LIOT, ANTHONY SYLVAIN JEAN-YVES; DINEV, DIMITAR PETKOV; TEXLER, ONDREJ; KANG, HYUN JAE; PALAN, JANVI CHETAN; SADI, SAJID
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061446/0716 →
Continuity (2)
Provisional Application 63349298 · Jun 6, 2022
Related Publication 20230394732A1 · Dec 7, 2023
References Cited (49)
US 7123262B2 · Francini et al. · 2006 [cited by applicant]
US 11158102B2 · Liu et al. · 2021 [cited by applicant]
US 11211060B2 · Li et al. · 2021 [cited by applicant]
US 11308671B2 · Chen et al. · 2022 [cited by applicant]
US 20050151743A1 · Sitrick · 2005 [cited by examiner]
US 20050286799A1 · Huang · 2005 [cited by examiner]
US 20120280974A1 · Wang · 2012 [cited by examiner]
US 20140198121A1 · Tong · 2014 [cited by applicant]
US 20180174348A1 · Bhat · 2018 [cited by examiner]
US 20190035130A1 · Hutchinson · 2019 [cited by examiner]
US 20190138096A1 · Lee et al. · 2019 [cited by applicant]
US 20200234482A1 · Krokhalev · 2020 [cited by examiner]
US 20200402284A1 · Saragih · 2020 [cited by examiner]
US 20210150793A1 · Stratton et al. · 2021 [cited by applicant]
US 20210233299A1 · Zhou et al. · 2021 [cited by applicant]
US 20210256962A1 · Liu et al. · 2021 [cited by applicant]
US 20210327404A1 · Savchenkov et al. · 2021 [cited by applicant]
US 20210350528A1 · Tang · 2021 [cited by applicant]
US 20210390748A1 · Liao · 2021 [cited by examiner]
US 20220084502A1 · Ma et al. · 2022 [cited by applicant]
US 20220108510A1 · Sagar et al. · 2022 [cited by applicant]
US 20220207262A1 · Jeong et al. · 2022 [cited by applicant]
US 20230014604A1 · Kim · 2023 [cited by examiner]
US 20230042654A1 · Zhang · 2023 [cited by applicant]
US 20230394715A1 · Texler et al. · 2023 [cited by applicant]
CN 110288682A · 2019 [cited by applicant]
CN 113554737A · 2021 [cited by applicant]
CN 113609255A · 2021 [cited by applicant]
KR 102251781B1 · 2021 [cited by applicant]
KR 20210070169A · 2021 [cited by examiner]
KR 1020210070169A · 2021 [cited by applicant]
WO 2021023869A1 · 2021 [cited by applicant]
WO 2021112365A1 · 2021 [cited by applicant]
WIPO Appln. No. PCT/KR2023/005004, International Search Report, Aug. 1, 2023, 3 pg. [cited by applicant]
WIPO Appln. No. PCT/KR2023/005004, Written Opinion,Aug. 1, 2023, 4 pg. [cited by applicant]
Aneja, D. et al., “A High-Fidelity Open Embodied Avatar with Lip Syncing and Expression Capabilities,” In 2019 International Conference on Multimodal Interaction Oct. 14, 2019 (pp. 69-73). [cited by applicant]
Prajwal, K.R. et al., “A lip sync expert is all you need for speech to lip generation in the wild,” arXiv Preprint, arXiv: 2008.10010v1, Aug. 23, 2020, 10 pg. [cited by applicant]
Ronneberger, O. et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation,” arXiv Preprint, arXiv: 1505.04597v1, May 18, 2015, 8 pg. [cited by applicant]
Siarohin, A. et al., “First Order Motion Model for Image Animation,” Advances in Neural Information Processing Systems, 2019, vol. 32, 11 pg. [cited by applicant]
“Synthesia, #1 AI Video Generation Platform,” [online] © 2023 Synthesia Limited [retrieved Nov. 2, 2023], retrieved from the Internet: <https://www.synthesia.io/>, 14 pg. [cited by applicant]
Suwajanakorn, S. et al. “Synthesizing Obama: Learning Lip Sync from Audio,” ACM Transactions on Graphics (ToG), Jul. 20, 2017, vol. 36, No. 4, pp. 1-3 [Abstract]. [cited by applicant]
Theis, J. et al., “Face2Face: Real-time Face Capture and Reenactment of RGB Videos,” arXiv Preprint, arXiv: 2007.14808v1, Jul. 29, 2020, 12 pg. [cited by applicant]
“Pinscreen: AI-Driven Virtual Avatars,” [online] © 2020 Pinscreen Inc. [retrieved Nov. 2, 2023], retrieved from the Internet: <https://www.pinscreen.com/>, 8 pg. [cited by applicant]
Zhou Y. et al., “VisemeNet: Audio-Driven Animator-Centric Speech Animation,” ACM Transactions on Graphics (ToG), vol. 37, No. 4, Aug. 2018., pp. 1-10. [cited by applicant]
Fan et al., “A Deep Bidirectional LSTM Approach for Video-Realistic Talking Head,” Multimed. Tools Appl., Springer Science + Business Media New York, 2015, 23 pg. [cited by applicant]
Yao et al., “Iterative Text-Based Editing of Talking-Heads Using Neural Retargeting,” ACM Transactions on Graphics, vol. 40, No. 3, Art. 20, Jul. 2021, 14 pg. [cited by applicant]
Guo et al., “AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis,” IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 5764-5774. [cited by applicant]
EP Appln. No. 23819978, Supplementary European Search Report, Jan. 8, 2025, 9 pg. [cited by applicant]
Li, L. et al., “Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation,” In Proc. of the AAAI Conf. on Artificial Intelligence, May 18, 2021, vol. 35, No. 3, pp. 1911-1920; [arXiv:2104.07995v2, May 7… [cited by applicant]