IP Library Granted Patent US 12,315,054
Granted Patent B2
US 12,315,054 · App. 17/422,167 · Granted May 27, 2025

Real-time generation of speech animation

Inventors: Mark Sagar (Auckland, NZ); Tim Szu-Hsien Wu (Auckland Central, NZ); Xiani Tan (Auckland, NZ); Xueyuan Zhang (Auckland, NZ)
Assignee: SOUL MACHINES LIMITED
G06T13/205G06T13/80G10L15/02G10L21/12G10L2015/025G10L2021/105
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,054
App. No.
17/422,167
Granted
May 27, 2025
Kind
B2
Abstract

To realistically animate a String (such as a sentence) a hierarchical search algorithm is provided to search for stored examples (Animation Snippets) of sub-strings of the String, in decreasing order of sub-string length, and concatenate retrieved sub-strings to complete the String of speech animation. In one embodiment, real-time generation of speech animation uses model visemes to predict the animation sequences at onsets of visemes and a look-up table based (data-driven) algorithm to predict the dynamics at transitions of visemes. Specifically posed Model Visemes may be blended with speech animation generated using another method at corresponding time points in the animation when the visemes are to be expressed. An Output Weighting Function is used to map Speech input and Expression input into Muscle-Based Descriptor weightings.

Claims (34)

1. A method for animating a communicative utterance comprising:

receiving a String to be animated, the String comprising a plurality of communicative utterance atoms;

receiving a plurality of Collections, each Collection including a plurality of Items comprising unique atom strings, each Collection storing Items of different lengths, wherein at least one of the Items is based on an occurrence of acoustic silence before or after a particular communicative utterance atom, and each Item including at least one Animation Snippet of the Item;

hierarchically searching the Collections for Items matching substrings of the String, wherein the hierarchical searching favours longer Items;

retrieving Animation Snippets for matched Items to cover the communicative utterance atoms; and

combining the retrieved Animation Snippets to animate the String.

2. The method of claim 1 wherein the communicative utterance is speech; and

wherein the at least one of the Items is pre-defined as the occurrence of acoustic silence before or after the particular communicative utterance atom.

3. The method of claim 1 wherein at least one Item includes a plurality of Animation Snippets, and an Animation Snippet is retrieved based on its duration.

4. The method of claim 1 wherein at least one Item includes a plurality of Animation Snippets, and an Animation Snippet is retrieved based on corresponding speech features.

5. The method of claim 1 wherein the Animation Snippets are associated with sound corresponding to the animation.

6. The method of claim 5 further comprising:

compressing and/or stretching the Animation Snippets to match the sound corresponding to the animation.

7. The method of claim 1 wherein one or more of the Items are included in a first Collection of Items from a plurality of different Collections of Items;

wherein the first Collection of Items comprises a Sentence Boundary Collection that includes a first type of dipho string and a second type of dipho string;

wherein the first type of dipho string represents an occurrence of acoustic silence pre-defined as being paired with a subsequent respective syllable type; and

wherein the second type of dipho string represents an occurrence of acoustic silence pre-defined as being paired with a preceding respective syllable type.

8. The method of claim 1 wherein the Animation Snippets store Muscle-Based Descriptor weights.

9. A method for animating a phoneme in context comprising:

receiving a Model Viseme;

receiving an Animation Snippet corresponding to a time series of animation weights of the phoneme being pronounced in context, the Animation Snippet of the phoneme based at least on a string that includes an occurrence of acoustic silence pre-defined as being before or after at least a particular portion of the phoneme; and

blending between the animation weights of the Model Viseme and animation weights of the Animation Snippet to animate the phoneme in context, wherein a degree of blending of the Model Viseme over time is modelled by a function, with a peak of the function being at or about a peak of the phoneme being pronounced.

10. The method of claim 9 wherein the Model Viseme is a Lip-Readable Viseme depicting the phoneme in a lip-readable manner.

11. The method of claim 9 wherein the Model Viseme is represented as an animation sequence.

12. A method for expressive speech animation including:

receiving a Priority Weighting for one or more animation inputs, wherein an Output Weighting Function is a function of the Priority Weighting, the one more animation inputs based on concatenated Animation Snippets, wherein a first Animation Snippet represents a phoneme based at least on a string that includes an occurrence of acoustic silence pre-defined as being before or after at least a particular portion of the phoneme;

receiving a first animation input associated with muscle-based descriptor information and a second animation input associated with muscle-based descriptor information;

defining at least one Muscle-Based Descriptor Class Weighting for each muscle-based descriptor;

using the first animation input and the second animation input as arguments in the Output Weighting Function configured to map the animation inputs to muscle-based descriptor weightings for animating the expressive speech animation,

wherein the Output Weighting Function is a function of the Muscle-Based Descriptor Class Weightings,

wherein the Output Weighting Function is configured to reconcile muscle-based descriptor information from the first and second animation inputs; and

animating using the mapped muscle-based descriptor weightings.

13. The method of claim 12 wherein the first animation input is for animating Speech.

14. The method of claim 12 wherein the second animation input is for animating Expression.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2026
From: SOUL MACHINES LIMITED
To: APPDIRECT, INC.
Reel/Frame 076058/0831 →
Priority Claims (1)
NZ 750233 · Jan 25, 2019 · national
Continuity (1)
Related Publication 20220108510A1 · Apr 7, 2022
References Cited (184)
US 4884972A · Gasper · 1989 [cited by examiner]
US 5111409A · Gasper · 1992 [cited by examiner]
US 5390278A · Gupta · 1995 [cited by examiner]
US 5613056A · Gasper · 1997 [cited by examiner]
US 5636325A · Farrett · 1997 [cited by examiner]
US 5657426A · Waters · 1997 [cited by examiner]
US 5687287A · Gandhi · 1997 [cited by examiner]
US 5857173A · Beard · 1999 [cited by examiner]
US 5860064A · Henton · 1999 [cited by examiner]
US 5880788A · Bregler · 1999 [cited by examiner]
US 5995119A · Cosatto · 1999 [cited by examiner]
US 6147692A · Shaw · 2000 [cited by examiner]
US 6208356B1 · Breen · 2001 [cited by examiner]
US 6249292B1 · Christian · 2001 [cited by examiner]
US 6250928B1 · Poggio · 2001 [cited by examiner]
US 6292778B1 · Sukkar · 2001 [cited by examiner]
US 6307576B1 · Rosenfeld · 2001 [cited by examiner]
US 6504546B1 · Cosatto · 2003 [cited by examiner]
US 6539354B1 · Sutton · 2003 [cited by examiner]
US 6556196B1 · Blanz · 2003 [cited by examiner]
US 6633846B1 · Bennett · 2003 [cited by examiner]
US 6654018B1 · Cosatto · 2003 [cited by examiner]
US 6661418B1 · McMillan · 2003 [cited by examiner]
US 6919892B1 · Cheiky · 2005 [cited by examiner]
US 7027054B1 · Cheiky · 2006 [cited by examiner]
US 7168953B1 · Poggio · 2007 [cited by examiner]
US 7209882B1 · Cosatto · 2007 [cited by examiner]
US 7216079B1 · Barnard · 2007 [cited by examiner]
US 7752044B2 · Lam · 2010 [cited by examiner]
US 8581911B2 · Becker · 2013 [cited by examiner]
US 8744856B1 · Ravishankar · 2014 [cited by examiner]
US 9094576B1 · Karakotsios · 2015 [cited by examiner]
US 9812151B1 · Amini · 2017 [cited by examiner]
US 9911218B2 · Theobald · 2018 [cited by examiner]
US 10360716B1 · van der Meulen · 2019 [cited by examiner]
US 10521946B1 · Roche · 2019 [cited by examiner]
US 10530928B1 · Ouimette · 2020 [cited by examiner]
US 10586369B1 · Roche · 2020 [cited by examiner]
US 10629192B1 · Streat · 2020 [cited by examiner]
US 10732708B1 · Roche · 2020 [cited by examiner]
US 10770092B1 · Adams · 2020 [cited by examiner]
US 11113859B1 · Xiao · 2021 [cited by examiner]
US 11232645B1 · Roche · 2022 [cited by examiner]
US 11270487B1 · Steptoe · 2022 [cited by examiner]
US 11386900B2 · Shillingford · 2022 [cited by examiner]
US 11468616B1 · Steptoe · 2022 [cited by examiner]
US 11551393B2 · Shang · 2023 [cited by examiner]
US 20020013707A1 · Shaw · 2002 [cited by examiner]
US 20020087329A1 · Massaro · 2002 [cited by examiner]
US 20030137515A1 · Cederwall · 2003 [cited by examiner]
US 20030184547A1 · Haratsch · 2003 [cited by examiner]
US 20040068408A1 · Qian · 2004 [cited by examiner]
US 20040098264A1 · Bowater · 2004 [cited by examiner]
US 20040107106A1 · Margaliot · 2004 [cited by examiner]
US 20040111266A1 · Coorman · 2004 [cited by examiner]
US 20050057570A1 · Cosatto · 2005 [cited by examiner]
US 20050080625A1 · Bennett · 2005 [cited by examiner]
US 20060009978A1 · Ma · 2006 [cited by examiner]
US 20060136214A1 · Sato · 2006 [cited by examiner]
US 20060149558A1 · Kahn · 2006 [cited by examiner]
US 20060221084A1 · Yeung · 2006 [cited by examiner]
US 20060290699A1 · Dimtrva et al. · 2006 [cited by applicant]
US 20070033042A1 · Marcheret · 2007 [cited by examiner]
US 20070038450A1 · Josifovski · 2007 [cited by examiner]
US 20070233492A1 · Matsumoto · 2007 [cited by examiner]
US 20080163074A1 · Tu · 2008 [cited by examiner]
US 20080259085A1 · Chen · 2008 [cited by examiner]
US 20080270129A1 · Colibro · 2008 [cited by examiner]
US 20080294433A1 · Yeung · 2008 [cited by examiner]
US 20080305454A1 · Kitching · 2008 [cited by examiner]
US 20090044112A1 · Basso et al. · 2009 [cited by applicant]
US 20090112905A1 · Mukerjee · 2009 [cited by examiner]
US 20090313016A1 · Cevik · 2009 [cited by examiner]
US 20090319270A1 · Gross · 2009 [cited by examiner]
US 20100007665A1 · Smith · 2010 [cited by examiner]
US 20100085363A1 · Smith · 2010 [cited by examiner]
US 20100145698A1 · Chen · 2010 [cited by examiner]
US 20100332229A1 · Aoyama · 2010 [cited by examiner]
US 20110106792A1 · Robertson · 2011 [cited by examiner]
US 20110131041A1 · Cortez · 2011 [cited by examiner]
US 20110175921A1 · Havaldar · 2011 [cited by examiner]
US 20120130717A1 · Xu · 2012 [cited by examiner]
US 20130006629A1 · Honda · 2013 [cited by examiner]
US 20130065205A1 · Park · 2013 [cited by applicant]
US 20130191129A1 · Kurata · 2013 [cited by examiner]
US 20130218568A1 · Tamura · 2013 [cited by examiner]
US 20130304587A1 · Ralston · 2013 [cited by examiner]
US 20140035929A1 · Matthews · 2014 [cited by examiner]
US 20140141392A1 · Yoon · 2014 [cited by examiner]
US 20140372100A1 · Jeong · 2014 [cited by examiner]
US 20150052084A1 · Kolluru · 2015 [cited by examiner]
US 20150287403A1 · Holzer Zaslansky · 2015 [cited by examiner]
US 20160030744A1 · Hubert-Brierre · 2016 [cited by examiner]
US 20160180568A1 · Bullivant · 2016 [cited by examiner]
US 20160203827A1 · Leff · 2016 [cited by examiner]
US 20160328875A1 · Fang · 2016 [cited by examiner]
US 20170154457A1 · Theobald · 2017 [cited by examiner]
US 20170178623A1 · Shamir · 2017 [cited by examiner]
US 20170243387A1 · Li · 2017 [cited by examiner]
US 20170294188A1 · Hayakawa · 2017 [cited by examiner]
US 20180027123A1 · Cartwright · 2018 [cited by examiner]
US 20180027351A1 · Cartwright · 2018 [cited by examiner]
US 20180047385A1 · Jiang · 2018 [cited by examiner]
US 20180068661A1 · Printz · 2018 [cited by examiner]
US 20180075843A1 · Hayakawa · 2018 [cited by examiner]
US 20180095636A1 · Valdivia · 2018 [cited by examiner]
US 20180096507A1 · Valdivia · 2018 [cited by examiner]
US 20180098059A1 · Valdivia · 2018 [cited by examiner]
US 20180157901A1 · Arbatman · 2018 [cited by examiner]
US 20180182396A1 · An · 2018 [cited by examiner]
US 20180197322A1 · Sagar · 2018 [cited by examiner]
US 20180253881A1 · Edwards · 2018 [cited by examiner]
US 20180277145A1 · Yamaya · 2018 [cited by examiner]
US 20180279063A1 · Sun · 2018 [cited by examiner]
US 20180295240A1 · Dickins · 2018 [cited by examiner]
US 20180336902A1 · Cartwright · 2018 [cited by examiner]
US 20180350388A1 · Jain · 2018 [cited by examiner]
US 20190013008A1 · Kunitake · 2019 [cited by examiner]
US 20190057533A1 · Habra · 2019 [cited by examiner]
US 20190147838A1 · Serletic, II · 2019 [cited by examiner]
US 20190287515A1 · Li · 2019 [cited by examiner]
US 20190392823A1 · Li · 2019 [cited by examiner]
US 20200106708A1 · Sleevi · 2020 [cited by examiner]
US 20200126283A1 · Van Vuuren · 2020 [cited by examiner]
US 20200135226A1 · Mittal · 2020 [cited by examiner]
US 20200160581A1 · Heller · 2020 [cited by examiner]
US 20200211248A1 · Baker · 2020 [cited by examiner]
US 20200279553A1 · McDuff · 2020 [cited by examiner]
US 20200302667A1 · del val Santos · 2020 [cited by examiner]
US 20210050031A1 · Hancock · 2021 [cited by examiner]
US 20210110831A1 · Shillingford · 2021 [cited by examiner]
US 20210158812A1 · Wooters · 2021 [cited by examiner]
US 20210248801A1 · Li · 2021 [cited by examiner]
US 20210248804A1 · Hussen Abdelaziz · 2021 [cited by examiner]
US 20210327431A1 · Stewart · 2021 [cited by examiner]
US 20210375260A1 · Yu · 2021 [cited by examiner]
US 20210390949A1 · Wang · 2021 [cited by examiner]
US 20220075820A1 · Walker · 2022 [cited by examiner]
US 20220076025A1 · Shin · 2022 [cited by examiner]
US 20220076026A1 · Walker · 2022 [cited by examiner]
US 20220076424A1 · Shin · 2022 [cited by examiner]
US 20220076705A1 · Walker · 2022 [cited by examiner]
US 20220076707A1 · Walker · 2022 [cited by examiner]
US 20220084273A1 · Pan · 2022 [cited by examiner]
US 20220108510A1 · Sagar · 2022 [cited by examiner]
US 20220191429A1 · Astarabadi · 2022 [cited by examiner]
US 20220215830A1 · Jawahar · 2022 [cited by examiner]
US 20220247973A1 · Astarabadi · 2022 [cited by examiner]
US 20220392430A1 · Kilgore · 2022 [cited by examiner]
US 20230111633A1 · Paruchuri · 2023 [cited by examiner]
US 20230117787A1 · Wu · 2023 [cited by examiner]
US 20230130287A1 · Zhao · 2023 [cited by examiner]
US 20230237987A1 · Fukuda · 2023 [cited by examiner]
US 20230306959A1 · Lin · 2023 [cited by examiner]
US 20230353707A1 · Astarabadi · 2023 [cited by examiner]
US 20230377238A1 · Hutton · 2023 [cited by examiner]
US 20240013802A1 · Federov · 2024 [cited by examiner]
US 20240087557A1 · Levine · 2024 [cited by examiner]
US 20240135973A1 · Bai · 2024 [cited by examiner]
US 20240177391A1 · Pan · 2024 [cited by examiner]
EP 2849087A1 · 2015 [cited by applicant]
JP 2009087328A · 2009 [cited by applicant]
KR 20060031449A · 2006 [cited by applicant]
KR 1020060090687A · 2006 [cited by applicant]
KR 100813034B1 · 2008 [cited by applicant]
KR 20120130627A · 2012 [cited by applicant]
WO WO2004100128A1 · 2004 [cited by applicant]
WO WO2012154618A2 · 2012 [cited by applicant]
WO 2015016723A1 · 2015 [cited by applicant]
WO 2017044499A1 · 2017 [cited by applicant]
WO WO2017075452A1 · 2017 [cited by applicant]
S. Okita, Y. Mitsukura and N. Hamada, “Augmented classification of Japanese visemes and hierarchical weighted discrimination for visual speech recognition,” 2013 IEEE Conference on Systems, Process & Control (ICSPC), Ku… [cited by examiner]
R. Amini, C. Lisetti and G. Ruiz, “HapFACS 3.0: FACS-Based Facial Expression Generator for 3D Speaking Virtual Characters,” in IEEE Transactions on Affective Computing, vol. 6, No. 4, pp. 348-360, Oct.-Dec. 1, 2015, doi… [cited by examiner]
Zhilin Wu and P. S. Aleksic, “Inner lip feature extraction for MPEG-4 facial animation,” 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing, Montreal, QC, Canada, 2004, pp. iii-633, doi: 10.1… [cited by examiner]
Zhigang Deng, U. Neumann, J. P. Lewis, Tae-Yong Kim, M. Bulut and S. Narayanan, “Expressive Facial Animation Synthesis by Learning Speech Coarticulation and Expression Spaces,” in IEEE Transactions on Visualization and … [cited by examiner]
E. Bozkurt, C. E. Erdem, E. Erzin, T. Erdem and M. Ozkan, “Comparison of Phoneme and Viseme Based Acoustic Units for Speech Driven Realistic lip Animation,” 2007 3DTV Conference, Kos, Greece, 2007, pp. 1-4, doi: 10.1109… [cited by examiner]
L. Dong, Y. Wang, K. Ni and K. Lu, “Facial animation system based on image warping algorithm,” 2011 International Conference on Electronics, Communications and Control (ICECC), Ningbo, China, 2011, pp. 2648-2653, doi: 1… [cited by examiner]
S. A. King and R. E. Parent, “Creating speech-synchronized animation,” in IEEE Transactions on Visualization and Computer Graphics, vol. 11, No. 3, pp. 341-352, May-Jun. 2005, doi: 10.1109/TVCG.2005.43. (Year: 2005). [cited by examiner]
A. Verma, L. V. Subramaniam, N. Rajput, C. Neti and T. A. Faruquie, “Animating expressive faces across languages, ” in IEEE Transactions on Multimedia, vol. 6, No. 6, pp. 791-800, Dec. 2004, doi: 10.1109/TMM.2004.837256… [cited by examiner]
PCT Application No. PCT/IB2020/050620 International Search Report dated Jun. 25, 2020. [cited by applicant]
Office Action issued in Korean Patent Application No. 10-2021-7026491 dated Mar. 22, 2024. [cited by applicant]
Office Action, Australian Patent Application No. 2020211809, mailing date Jan. 5, 2024. [cited by applicant]
European Search Report and Opinion, European Application No. 20744394.6, mailing date Aug. 1, 2022. [cited by applicant]
Written Opinion, Signapore Patent Application No. 11202107022Y, mailing date May 11, 2023. [cited by applicant]