IP Library › Granted Patent US 12,333,258
Granted Patent B2
US 12,333,258 · App. 17/894,967 · Granted Jun 17, 2025

Multi-level emotional enhancement of dialogue

Inventors: Sanchita Tiwari (Cumming, GA); Justin Ali Kennedy (Norwell, MA); Dirk Van Dall (Shelter Island, NY); Xiuyang Yu (Unionville, CT); Daniel Cahall (Philadelphia, PA); Brian Kazmierczak (Hamden, CT)
Assignee: Disney Enterprises, Inc.
G06F40/35G06F40/284G06F40/289G06N20/00G10L13/08G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,258
App. No.
17/894,967
Granted
Jun 17, 2025
Kind
B2
Abstract

A system for emotionally enhancing dialogue includes a computing platform having processing hardware and a system memory storing a software code including a predictive model. The processing hardware is configured to execute the software code to receive dialogue data identifying an utterance for use by a digital character in a conversation, analyze, using the dialogue data, an emotionality of the utterance at multiple structural levels of the utterance, and supplement the utterance with one or more emotional attributions, using the predictive model and the emotionality of the utterance at the multiple structural levels, to provide one or more candidate emotionally enhanced utterance(s). The processing hardware further executes the software code to perform an audio validation of the candidate emotionally enhanced utterance(s) to provide a validated emotionally enhanced utterance, and output an emotionally attributed dialogue data providing the validated emotionally enhanced utterance for use by the digital character in the conversation.

Claims (48)

1. A system comprising:

a computing platform having a speech synthesizer, a processing hardware and a system memory storing a software code, the software code including at least one of a trained machine learning (ML) model configured to function as an autoregressive generator or a stochastic model trained using unsupervised or semi-supervised learning;

the processing hardware configured to execute the software code to:

receive dialogue data identifying an utterance for use by a digital character in a conversation;

analyze, using the dialogue data, an emotionality of the utterance at a plurality of structural levels of the utterance;

supplement the utterance with one or more emotional attributions, using the at least one of the trained ML model or the stochastic model and the emotionality of the utterance at the plurality of structural levels, to provide a plurality of candidate emotionally enhanced utterances;

perform an audio validation of at least some of the plurality of candidate emotionally enhanced utterances to provide a validated emotionally enhanced utterance including a non-verbal vocalization, wherein the audio validation identifies the validated emotionally enhanced utterance as having a best audio quality of the at least some of the plurality of candidate emotionally enhanced utterances;

output an emotionally attributed dialogue data providing the validated emotionally enhanced utterance for use by the digital character in the conversation; and

synthesize, using the speech synthesizer and the emotionally attributed dialogue data, the validated emotionally enhanced utterance to generate a synthesized speech for utterance by the digital character.

2. The system of claim 1 , wherein the digital character is configured to utter the synthesized speech.

3. The system of claim 1 , wherein the plurality of structural levels of the utterance comprise at least two of token level, a phrase level, or an entire utterance level.

4. The system of claim 1 , wherein the predictive model comprises the trained ML model, and wherein the trained ML model comprises transformer-based token insertion ML model.

5. The system of claim 4 , wherein the transformer-based token insertion ML model takes as input one or more of: a global sentiment of the dialogue data, a global mood of the dialogue data, a sentiment by phrase within the dialogue data, a sentiment by token in the dialogue data, a given character embedding, or extracted entities from the dialogue data.

6. The system of claim 1 , wherein the one or more emotional attributions identify one or more of a prosodic variation, or a word rate.

7. The system of claim 1 , wherein the conversation is between the digital character and a user of the system, and wherein the processing hardware is further configured to execute the software code to:

obtain a user profile of the user, the user profile including a user history of the user; and

analyze the emotionality of the utterance further using the user profile.

8. The system of claim 1 , wherein the digital character is associated with a character persona, and wherein the processing hardware is further configured to execute the software code to:

obtain the character persona; and

analyze the emotionality of the utterance further using the character persona.

9. The system of claim 1 , wherein the processing hardware is further configured to execute the software code to:

determine, using another predictive model of the software code, a quality score for each of the plurality of candidate emotionally enhanced utterances to provide a plurality of quality scores corresponding respectively to the plurality of candidate emotionally enhanced utterances; and

identify, based on the plurality of quality scores, the at least some of the plurality of candidate emotionally enhanced utterances for the audio validation.

10. The system of claim 1 , wherein the dialogue data and the emotionally attributed dialogue data comprise text.

11. The system of claim 1 , wherein the validated emotionally enhanced utterance further include a non-verbal filler.

12. A method for use by a system including a computing platform having a speech synthesizer, a processing hardware and a system memory storing a software code, the software code including at least one of a trained machine learning (ML) model configured to function as an autoregressive generator or a stochastic model trained using unsupervised or semi-supervised learning:

receiving, by the software code executed by the processing hardware, dialogue data identifying an utterance for use by a digital character in a conversation;

analyzing, by the software code executed by the processing hardware and using the dialogue data, an emotionality of the utterance at a plurality of structural levels of the utterance;

supplementing the utterance with one or more emotional attributions, by the software code executed by the processing hardware and using the at least one of the trained ML model or the stochastic model and the emotionality of the utterance at the plurality of structural levels, to provide a plurality of candidate emotionally enhanced utterances;

performing, by the software code executed by the processing hardware, an audio validation of at least some of the plurality of candidate emotionally enhanced utterances to provide a validated emotionally enhanced utterance including a non-verbal vocalization, wherein the audio validation identifies the validated emotionally enhanced utterance as having a best audio quality of the at least some of the plurality of candidate emotionally enhanced utterances;

outputting, by the software code executed by the processing hardware, an emotionally attributed dialogue data providing the validated emotionally enhanced utterance for use by the digital character in the conversation; and

synthesizing, by the software code executed by the processing hardware using the speech synthesizer and the emotionally attributed dialogue data, the validated emotionally enhanced utterance to generate a synthesized speech for utterance by the digital character.

13. The method of claim 12 , further comprising:

the digital character uttering the synthesized speech.

14. The method of claim 12 , wherein the plurality of structural levels of the utterance comprise at least two of token level, a phrase level, or an entire utterance level.

15. The method of claim 12 , wherein the predictive model comprises the trained ML model, and wherein the trained ML model comprises a transformer-based token insertion ML model.

16. The method of claim 12 , wherein the one or more emotional attributions identify one or more of a prosodic variation, or a word rate.

17. The method of claim 12 , wherein the conversation is between the digital character and a user of the system, the method further comprising:

obtaining, by the software code executed by the processing hardware, a user profile of the user, the user profile including a user history of the user; and

analyzing the emotionality of the utterance further using the user profile.

18. The method of claim 12 , wherein the digital character is associated with a character persona, the method further comprising:

obtaining, by the software code executed by the processing hardware, the character persona; and

analyzing the emotionality of the utterance further using the character persona.

19. The method of claim 12 , further comprising:

determining, by the software code executed by the processing hardware and using another predictive model of the software code, a quality score for each of the plurality of candidate emotionally enhanced utterances to provide a plurality of quality scores corresponding respectively to the plurality of candidate emotionally enhanced utterances; and

identifying, by the software code executed by the processing hardware based on the plurality of quality scores, the at least some of the plurality of candidate emotionally enhanced utterances for the audio validation.

20. The method of claim 12 , wherein the dialogue data and the emotionally attributed dialogue data comprise text.

21. The method of claim 12 , wherein the validated emotionally enhanced utterance further include a non-verbal filler.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2022
From: TIWARI, SANCHITA; KENNEDY, JUSTIN ALI; VAN DALL, DIRK; YU, XIUYANG; CAHALL, DANIEL; KAZMIERCZAK, BRIAN
To: DISNEY ENTERPRISES, INC.
Reel/Frame 060900/0691 →
Continuity (1)
Related Publication 20240070399A1 · Feb 29, 2024
References Cited (76)
US 11122240B2 · Peters · 2021 [cited by examiner]
US 11170175B1 · Kohli · 2021 [cited by examiner]
US 11238519B1 · Ravi · 2022 [cited by examiner]
US 11526541B1 · Chadwick · 2022 [cited by examiner]
US 11562744B1 · Gao · 2023 [cited by examiner]
US 11721357B2 · Togawa · 2023 [cited by examiner]
US 11756567B2 · Wilson · 2023 [cited by examiner]
US 11854538B1 · Rozgic · 2023 [cited by examiner]
US 11900914B2 · Fernandez Guajardo · 2024 [cited by examiner]
US 12106746B2 · Lin · 2024 [cited by examiner]
US 12288560B2 · Reece · 2025 [cited by examiner]
US 20030023443A1 · Shizuka · 2003 [cited by examiner]
US 20080077387A1 · Ariu · 2008 [cited by examiner]
US 20110208522A1 · Pereg · 2011 [cited by examiner]
US 20140303957A1 · Lee · 2014 [cited by examiner]
US 20150038806A1 · Kaleal, III · 2015 [cited by examiner]
US 20160329043A1 · Kim · 2016 [cited by examiner]
US 20170125008A1 · Maisonnier · 2017 [cited by examiner]
US 20170162186A1 · Tamura · 2017 [cited by examiner]
US 20190220505A1 · Shinohara · 2019 [cited by examiner]
US 20190251152A1 · Leydon · 2019 [cited by examiner]
US 20200042285A1 · Choi · 2020 [cited by examiner]
US 20200066264A1 · Kwatra · 2020 [cited by examiner]
US 20200286506A1 · Deshpande · 2020 [cited by examiner]
US 20210011545A1 · Min · 2021 [cited by examiner]
US 20210097976A1 · Chicote · 2021 [cited by examiner]
US 20210097980A1 · Lezzoum · 2021 [cited by examiner]
US 20210151034A1 · Hasan · 2021 [cited by examiner]
US 20210151046A1 · Nicholson · 2021 [cited by examiner]
US 20210183358A1 · Mao · 2021 [cited by examiner]
US 20210201162A1 · Bhan · 2021 [cited by examiner]
US 20210233031A1 · Preuss · 2021 [cited by examiner]
US 20210264900A1 · Reece · 2021 [cited by examiner]
US 20210264921A1 · Reece · 2021 [cited by examiner]
US 20210264929A1 · Osebe · 2021 [cited by examiner]
US 20210287657A1 · Deng · 2021 [cited by examiner]
US 20210295820A1 · Hirvonen · 2021 [cited by examiner]
US 20210335367A1 · Graff · 2021 [cited by examiner]
US 20210343270A1 · Zhang · 2021 [cited by examiner]
US 20210365962A1 · Tolentino · 2021 [cited by examiner]
US 20220059225A1 · Shah · 2022 [cited by examiner]
US 20220132218A1 · Aher · 2022 [cited by examiner]
US 20220165254A1 · Decker · 2022 [cited by examiner]
US 20220180893A1 · Lihan · 2022 [cited by examiner]
US 20220223064A1 · Chauhan · 2022 [cited by examiner]
US 20220392428A1 · Fernandez Guajardo · 2022 [cited by examiner]
US 20230007359A1 · Aher · 2023 [cited by examiner]
US 20230058259A1 · Suneja · 2023 [cited by examiner]
US 20230096543A1 · Moy · 2023 [cited by examiner]
US 20230102789A1 · Nigul · 2023 [cited by examiner]
US 20230111824A1 · Mukherjee · 2023 [cited by examiner]
US 20230114150A1 · Lillelund · 2023 [cited by examiner]
US 20230138741A1 · Patel · 2023 [cited by examiner]
US 20230222293A1 · Key · 2023 [cited by examiner]
US 20230229934A1 · Iwama · 2023 [cited by examiner]
US 20230245651A1 · Wang · 2023 [cited by examiner]
US 20230259437A1 · White · 2023 [cited by examiner]
US 20230260536A1 · Xu · 2023 [cited by examiner]
US 20230298616A1 · Brownlee · 2023 [cited by examiner]
US 20230335123A1 · Liu · 2023 [cited by examiner]
US 20230350847A1 · Cluff · 2023 [cited by examiner]
US 20230368146A1 · Pande · 2023 [cited by examiner]
US 20230379543A1 · Singh · 2023 [cited by examiner]
US 20230379544A1 · Singh · 2023 [cited by examiner]
US 20230412686A1 · Mansaray · 2023 [cited by examiner]
US 20240005082A1 · Yin · 2024 [cited by examiner]
US 20240005335A1 · Foy · 2024 [cited by examiner]
US 20240005905A1 · Chen · 2024 [cited by examiner]
US 20240070397A1 · Yuan · 2024 [cited by examiner]
US 20240105207A1 · Kruk · 2024 [cited by examiner]
US 20240249740A1 · Nesfield · 2024 [cited by examiner]
Yi Lei, Shan Yang, Lei Xie “Fine-Grained Emotion Strength Transfer, Control and Prediction for Emotional Speech Synthesis” IEEE Spoken Language Technology Workshop (SLT), Shenzhen, China, 2021, pp. 423-430. [cited by applicant]
Kun Zhou, Berrak Sisman, Rajib Rana, Bjorn W. Schuller, Haizhou Li “Emotion Intensity and it's Control for Emotional Voice Conversion” IEEE Transaction on Affective Computing May 2022 18 Pgs. [cited by applicant]
Kaichun Yao, Libo Zhang, Tiejian Luo, Dawei Du, Yanjun Wu “Non-deterministic and emotional chatting machine: learning emotional conversion generation using conditional variational autoencoders” Neural Computing and Appl… [cited by applicant]
Jong Yoon Lim, Inkyu Sa, Ho Seok Ahn, Norina Gasteiger, Sanghyub John Lee, and Bruce MacDonald “Substance Extraction from Text Using Cpverage-Based Deep Learning Language Models” MDPI Apr. 12, 2021 pp. 1-27. [cited by applicant]
Kai Yang, Raymond Y. K. Lau, and Ahmed Abbasi “Getting Personal: A Deep Learning Artifact for Text-based Measurement of Personality” Forthcoming in Information Systems Research (ISR), 2022 48 Pgs. [cited by applicant]
Cited By (2)
US 12,530,545 US 12,554,934