IP Library Granted Patent US 12,450,806
Granted Patent B2
US 12,450,806 · App. 17/873,437 · Granted Oct 21, 2025

System and method for generating emotionally-aware virtual facial expressions

Inventors: Subham Biswas (Maharashtra, IN); Saurabh Tahiliani (Uttar Pradesh, IN)
Assignee: Verizon Patent and Licensing Inc.
G06T13/00G06F40/20G10L15/16G10L15/22G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,806
App. No.
17/873,437
Granted
Oct 21, 2025
Kind
B2
Abstract

Techniques for generating emotionally-aware digital content are disclosed. In one embodiment, a method is disclosed comprising obtaining audio input, obtaining a textual representation of the audio input; using the textual representation of the audio input to identify an emotion corresponding to the audio input; generating an emotionally-aware facial representation in accordance with the textual representation and the identified emotion; using the emotionally-aware facial representation to generate one or more images comprising at least one facial expression corresponding to the identified emotion; and providing digital content comprising the one or more images.

Claims (50)

1. A method comprising:

obtaining, by a computing device, audio input;

obtaining, by the computing device, a textual representation of the audio input;

using, by the computing device, the textual representation of the audio input to identify an emotion corresponding to the audio input;

generating, by the computing device, an emotionally-aware facial representation in accordance with the textual representation and the identified emotion, the generation of the emotionally-aware facial representation comprising:

determining, by the computing device, a set of image embeddings using a first neural network trained to generate the set of image embeddings using the identified emotion;

determining, by the computing device, a set of text embeddings using a second neural network trained to generate the set of text embeddings using a phonemic representation of the audio input; and

using, by the computing device, the set of image embeddings and the set of text embeddings and a third neural network trained to generate the emotionally-aware facial representation in accordance with the textual representation and the identified emotion;

using, by the computing device, the emotionally-aware facial representation to generate one or more images comprising at least one facial expression corresponding to the identified emotion; and

providing, by the computing device, at a time corresponding to the obtaining of the audio input and the identification of the emotion, digital content comprising the one or more images, such that a digital representation of a user is modified to visually depict the emotionally-aware facial representation.

2. The method of claim 1 , wherein a previously-generated image is used with the set of image embeddings and the set of text embeddings, by the third neural network, to generate the emotionally-aware facial representation.

3. The method of claim 1 , wherein the first neural network comprises a Stacked-CNN-LSTM neural network comprising a Convolutional Neural Network and a Long-Short-Term Memory (LSTM) neural network, the second neural network comprises an attention-based neural network and the third neural network comprises an attention-based encoder-decoder neural network.

4. The method of claim 1 , further comprising:

using, by the computing device, the textual representation of the audio input to determine the phonemic representation of the audio input.

5. The method of claim 1 , wherein the emotionally-aware facial representation comprises a representation of the at least one facial expression incorporated into a facial structure.

6. The method of claim 5 , wherein an image corresponding to the audio input comprises the facial structure into which the at least one facial expression is incorporated.

7. The method of claim 1 , wherein using the textual representation of the audio input to identify an emotion corresponding to the audio input further comprises:

using, by the computing device, a trained emotion classifier and the textual representation of the audio input to determine the identified emotion.

8. The method of claim 1 , wherein the audio input corresponds to a figure having a face depicted in an image corresponding to the audio input, and the one or more images comprise the face of a character depicted with the at least one facial expression corresponding to the identified emotion.

9. The method of claim 1 , wherein the digital content comprises a video comprising a number of frames generated using a number of images, each image comprising at least one facial expression corresponding to a respective emotion.

10. A non-transitory computer-readable storage medium tangibly encoded with computer-executable instructions that when executed by a processor associated with a computing device perform a method comprising:

obtaining audio input;

obtaining a textual representation of the audio input;

using the textual representation of the audio input to identify an emotion corresponding to the audio input;

generating an emotionally-aware facial representation in accordance with the textual representation and the identified emotion, the generation of the emotionally-aware facial representation comprising:

determining, by the computing device, a set of image embeddings using a first neural network trained to generate the set of image embeddings using the identified emotion;

determining, by the computing device, a set of text embeddings using a second neural network trained to generate the set of text embeddings using a phonemic representation of the audio input; and

using, by the computing device, the set of image embeddings and the set of text embeddings and a third neural network trained to generate the emotionally-aware facial representation in accordance with the textual representation and the identified emotion;

using the emotionally-aware facial representation to generate one or more images comprising at least one facial expression corresponding to the identified emotion; and

providing, at a time corresponding to the obtaining of the audio input and the identification of the emotion, digital content comprising the one or more images, such that a digital representation of a user is modified to visually depict the emotionally-aware facial representation.

11. The non-transitory computer-readable storage medium of claim 10 , wherein a previously-generated image is used with the set of image embedding and the set of text embeddings, by the third neural network, to generate the emotionally-aware facial representation.

12. The non-transitory computer-readable storage medium of claim 10 , wherein the first neural network comprises a Stacked-CNN-LSTM neural network comprising a Convolutional Neural Network and a Long-Short-Term Memory (LSTM) neural network, the second neural network comprises an attention-based neural network and the third neural network comprises an attention-based encoder-decoder neural network.

13. The non-transitory computer-readable storage medium of claim 10 , the method further comprising:

using the textual representation of the audio input to determine the phonemic representation of the audio input.

14. The non-transitory computer-readable storage medium of claim 10 , wherein the emotionally-aware facial representation comprises a representation of the at least one facial expression incorporated into a facial structure.

15. The non-transitory computer-readable storage medium of claim 14 , wherein an image corresponding to the audio input comprises the facial structure into which the at least one facial expression is incorporated.

16. The non-transitory computer-readable storage medium of claim 10 , wherein using the textual representation of the audio input to identify an emotion corresponding to the audio input further comprises:

using a trained emotion classifier and the textual representation of the audio input to determine the identified emotion.

17. The non-transitory computer-readable storage medium of claim 10 , wherein the audio input corresponds to a figure having a face depicted in an image corresponding to the audio input, and the one or more images comprise the face of a character depicted with the at least one facial expression corresponding to the identified emotion.

18. A computing device comprising:

a processor, configured to:

obtain audio input;

obtain a textual representation of the audio input;

use the textual representation of the audio input to identify an emotion corresponding to the audio input;

generate an emotionally-aware facial representation in accordance with the textual representation and the identified emotion, the generation of the emotionally aware facial representation comprising:

determining, by the computing device, a set of image embeddings using a first neural network trained to generate the set of image embeddings using the identified emotion;

determining, by the computing device, a set of text embeddings using a second neural network trained to generate the set of text embeddings using a phonemic representation of the audio input; and

using, by the computing device, the set of image embeddings and the set of text embeddings and a third neural network trained to generate the emotionally-aware facial representation in accordance with the textual representation and the identified emotion;

use the emotionally-aware facial representation to generate one or more images comprising at least one facial expression corresponding to the identified emotion; and

provide, at a time corresponding to the obtaining of the audio input and the identification of the emotion, digital content comprising the one or more images, such that a digital representation of a user is modified to visually depict the emotionally-aware facial representation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2022
From: BISWAS, SUBHAM; TAHILIANI, SAURABH
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 060623/0164 →
Continuity (1)
Related Publication 20240037824A1 · Feb 1, 2024
References Cited (20)
US 10204625B2 · Mishra · 2019 [cited by examiner]
US 11462207B1 · Tao · 2022 [cited by examiner]
US 11749279B2 · Gharpure · 2023 [cited by examiner]
US 11756244B1 · Bhunia · 2023 [cited by examiner]
US 11816434B2 · Williams · 2023 [cited by examiner]
US 20170352361A1 · Thörn · 2017 [cited by examiner]
US 20180136615A1 · Kim · 2018 [cited by examiner]
US 20190045270A1 · Vats · 2019 [cited by examiner]
US 20200051583A1 · Wu · 2020 [cited by examiner]
US 20200410739A1 · Shin · 2020 [cited by examiner]
US 20210004543A1 · Kobayashi · 2021 [cited by examiner]
US 20220137992A1 · Arth · 2022 [cited by examiner]
US 20220139376A1 · Buesser · 2022 [cited by examiner]
US 20220171960A1 · Nelson · 2022 [cited by examiner]
US 20230169990A1 · Biswas · 2023 [cited by examiner]
US 20230237417A1 · Kumar · 2023 [cited by examiner]
US 20230260536A1 · Xu · 2023 [cited by examiner]
US 20230360437A1 · Shin · 2023 [cited by examiner]
US 20240037824A1 · Biswas · 2024 [cited by examiner]
CN 114091930A · 2022 [cited by examiner]