IP Library Granted Patent US 12,205,577
Granted Patent B1
US 12,205,577 · App. 17/217,031 · Granted Jan 21, 2025

Virtual conversational companion

Inventors: Taehwan Kim (La Crescenta, CA); Sanqiang Zhao (Pittsburgh, PA); Robinson Piramuthu (Oakland, CA); Seokhwan Kim (San Jose, CA); Yang Liu (Los Altos, CA); Gokhan Tur (Los Altos, CA); Eshan Bhatnagar (Santa Clara, CA)
Assignee: Amazon Technologies, Inc.
G10L15/1815G06T13/205G06T13/40G06T13/80G10L13/08G10L15/22G10L25/57G10L2015/227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,577
App. No.
17/217,031
Granted
Jan 21, 2025
Kind
B1
Abstract

Techniques for rendering visual content, in response to one or more utterances, are described. A device receives one or more utterances that define a parameter(s) for desired output content. A system (or the device) identifies natural language data corresponding to the desired content, and uses natural language generation processes to update the natural language data based on the parameter(s). The system (or the device) then generates an image based on the updated natural language data. The system (or the device) also generates video data of an avatar. The device displays the image and the avatar, and synchronizes movements of the avatar with output of synthesized speech of the updated natural language data. The device may also display subtitles of the updated natural language data, and cause a word of the subtitles to be emphasized when synthesized speech of the word is being output.

Claims (101)

1. A computer-implemented method comprising:

receiving first input audio data corresponding to a first spoken natural language input;

determining the first spoken natural language input includes:

a first portion requesting first content be output; and

a second portion representing a first parameter to be used instead of a second parameter included in the first content;

generating, by a content generator component, first natural language data corresponding to the first content updated to include the first parameter instead of the second parameter;

generating, by an image generator component, first image data representing the first natural language data;

generating, by an avatar generator component, first video data including a representation of an avatar corresponding to the first natural language data; and

generating, by a response generator component, first output data including the first image data, first output audio data including first synthesized speech corresponding to the first natural language data, and the first video data, wherein the first output data synchronizes display of the representation of the avatar with output of the first synthesized speech.

2. The computer-implemented method of claim 1 , further comprising:

determining, by the content generator component, second natural language data corresponding to the first content;

determining, by the content generator component, a portion of the second natural language data corresponding to the first second parameter; and

generating, by the content generator component, the first natural language data by performing natural language generation processing to replace the portion with the first parameter.

3. The computer-implemented method of claim 1 , further comprising:

determining, by a question and answering (Q&A) component, a dialog identifier corresponding to an ongoing dialog;

determining, by the Q&A component, second natural language data associated with the dialog identifier, the second natural language data being represented in previous output data;

generating, by the Q&A component, third natural language data corresponding to a question related to the second natural language data;

generating, by the avatar generator component, first data including the representation of the avatar corresponding to the third natural language data;

receiving, by the response generator component, the third natural language data and the first data;

sending, by the response generator component, the third natural language data to a text-to-speech (TTS) component;

receiving, by the response generator component, second output audio data from the TTS component, the second output audio data including second synthesized speech corresponding to the third natural language data; and

generating, by the response generator component, second output data including the first data and the second output audio data, the second output data synchronizing display of the representation of the avatar with output of the second synthesized speech.

4. The computer-implemented method of claim 1 , further comprising:

determining a user identifier corresponding to the first input audio data;

querying, by the content generator component, a profile storage for a parameter associated with the user identifier;

receiving, by the content generator component and from the profile storage, first data representing a third parameter; and

generating, by the content generator component, the first natural language data further based at least in part on the third parameter.

5. The computer-implemented method of claim 1 , further comprising:

receiving second input audio data corresponding to a second spoken natural language input;

determining the second spoken natural language input requests a next segment of content be output;

determining a content identifier corresponding to the first content;

receiving, by a content navigator component, the content identifier and a request for a next segment of the first content;

determining, by the content generator component and based at least in part on the content navigator component receiving the content identifier and the request, second natural language data associated with the content identifier and corresponding to the next segment;

generating, by the content generator component, third natural language data by replacing at least a first portion of the second natural language data with the first parameter;

generating, by the image generator component, second image data representing the third natural language data; and

generating, by the response generator component, second output data including the second image data and second output audio data including second synthesized speech corresponding to the third natural language data.

6. The computer-implemented method of claim 1 , further comprising:

receiving, by the response generator component, the first output audio data from a text-to-speech component, the first output audio data including a start token and an end token corresponding to a portion of the first synthesized speech; and

generating, by the response generator component, the first output data to:

correspond a first portion of the first video data, corresponding to a beginning of a facial expression of the avatar, with a first portion of the first synthesized speech corresponding to the start token; and

correspond a second portion of the first video data, corresponding to an end of the facial expression, with a second portion of the first synthesized speech corresponding to the end token.

7. The computer-implemented method of claim 1 , further comprising, by the response generator component:

configuring the first output audio data to correspond to a first output time duration; and

configuring the first video data to correspond to the first output time duration.

8. The computer-implemented method of claim 1 , wherein the avatar corresponds to a first character and the method further comprises:

determining second natural language data corresponding to the first content;

determining the second natural language data corresponds to a second character; and

generating, by the avatar generator component, second video data including a representation of a second avatar different from the avatar.

9. The computer-implemented method of claim 8 , wherein the first synthesized speech corresponds to first audio characteristics and the method further comprises:

generating, by the response generator component, second output audio data including second synthesized speech corresponding to the second natural language data, the second synthesized speech further corresponding to second audio characteristics different from the first audio characteristics.

10. A computing system comprising:

a first component configured to determine a first spoken natural language input includes:

a first portion requesting first content be output; and

a second portion representing a first parameter to be used instead of a second parameter included in the first content;

a content generator component configured to generate first natural language data corresponding to the first content updated to include the first parameter instead of the second parameter;

an image generator component configured to generate first image data representing the first natural language data;

an avatar generator component configured to generate first video data including a representation of an avatar corresponding to the first natural language data; and

a response generator component configured to generate first output data including the first image data, first output audio data including first synthesized speech corresponding to the first natural language data, and the first video data, the first output data synchronizing display of the representation of the avatar with output of the first synthesized speech.

11. The computing system of claim 10 , wherein the content generator component is further configured to:

determine second natural language data corresponding to the first content;

determine a portion of the second natural language data corresponding to the second parameter; and

generate the first natural language data by performing natural language generation processing to replace the portion with the first parameter.

12. The computing system of claim 10 , further comprising:

a question and answering component configured to:

determine a dialog identifier corresponding to an ongoing dialog;

determine second natural language data associated with the dialog identifier, the second natural language data being represented in previous output data; and

generate third natural language data corresponding to a question related to the second natural language data; and

the avatar generator component configured to generate first data including the representation of the avatar corresponding to the third natural language data,

wherein the response generator component is further configured to:

receive the third natural language data;

receive the first data;

send the third natural language data to a text-to-speech (TTS) component;

receive, from the TTS component, second output audio data including second synthesized speech corresponding to the third natural language data; and

generate second output data including the first data and the second output audio data, the second output data synchronizing display of the representation of the avatar with output of the second synthesized speech.

13. The computing system of claim 10 , wherein the content generator component is further configured to:

determine a user identifier corresponding to the first spoken natural language input;

query a profile storage for a parameter associated with the user identifier;

receive, from the profile storage, first data representing a third parameter; and

generate the first natural language data further based at least in part on the third parameter.

14. The computing system of claim 10 , wherein:

the first component is further configured to determine a second spoken natural language input requests a next segment of content be output;

a content generator component is further configured to:

determine a content identifier corresponding to the first content;

determine second natural language data associated with the content identifier and corresponding to the next segment; and

generate third natural language data by replacing at least a first portion of the second natural language data with the first parameter;

the image generator component is further configured to generate second image data representing the third natural language data; and

the response generator component is further configured to generate second output data including the second image data and second output audio data including second synthesized speech corresponding to the third natural language data.

15. The computing system of claim 10 , wherein the response generator component is further configured to:

receive the first output audio data from a text-to-speech component, the first output audio data including a start token and an end token corresponding to a portion of the first synthesized speech; and

generate the first output data to:

correspond a first portion of the first video data, corresponding to a beginning of a facial expression of the avatar, with a first portion of the first synthesized speech corresponding to the start token; and

correspond a second portion of the first video data, corresponding to an end of the facial expression, with a second portion of the first synthesized speech corresponding to the end token.

16. The computing system of claim 10 , wherein the response generator component is further configured to:

configure the first output audio data to correspond to a first output time duration; and

configure the first video data to correspond to the first output time duration.

17. The computing system of claim 10 , wherein the system is further configured to:

determine second natural language data corresponding to the first content;

determine the second natural language data corresponds to a second character; and

generate, by the avatar generator component, second video data including a representation of a second avatar different from the avatar.

18. The computing system of claim 17 , wherein the first synthesized speech corresponds to first audio characteristics and wherein the system is further configured to:

generate, by the response generator component, second output audio data including second synthesized speech corresponding to the second natural language data, the second synthesized speech further corresponding to second audio characteristics different from the first audio characteristics.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2021
From: KIM, TAEHWAN; ZHAO, SANQIANG; PIRAMUTHU, ROBINSON; KIM, SEOKHWAN; LIU, YANG; TUR, GOKHAN; BHATNAGAR, ESHAN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 055767/0053 →
References Cited (18)
US 10318586B1 · Rose · 2019 [cited by examiner]
US 20190295545A1 · Andreas · 2019 [cited by examiner]
US 20200134103A1 · Mankovskii · 2020 [cited by examiner]
Breazeal, et al., (2016) “Social Robotics,” In: Siciliano B., Khatib O. (eds) Springer Handbook of Robotics, https://doi.org/10.1007/978-3-31932552-1_72, pp. 1935-1971. [cited by applicant]
Turing, “Computing Machinery and Intelligence,” Mind Journal, Oct. 1950, vol. LIX. No. 236, (retrieved from https://academic.oup.com/mind/article/LIX/236/433/986238), pp. 433-460. [cited by applicant]
Smith et al., “The Development of Embodied Cognition: Six Lessons from Babies,” 2005 Massachusetts Institute of Technonolgy, Artificial Life, vol. 11, pp. 13-29. [cited by applicant]
Chen et al., “UNITER: UNiversal Image-TExt Representation Learning,” 2020 European Conference on Computer Vision, arXiv:1909.11740, pp. 1-26. [cited by applicant]
Brown et al., “Language Models are Few-Shot Learners,” (NeurIPS 2020), arXiv:2005.14165, pp. 1-25. [cited by applicant]
Breazeal, “Toward sociable robots,” 2003, Robotics and Autonomous Systems, 42, pp. 167-175. [cited by applicant]
Greer, “Eight Ways to Help Improve Your Child's Vocabulary,” https://lifehacker.com/eight-ways-to-help-improve-your-childs-vocabulary-1645796717, 2014, 9 pages. [cited by applicant]
Lan et al., “ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations,” International Conference on Learning Representations (ICLR 2020), arXiv:1909.11942, pp. 1-17. [cited by applicant]
Rajpurkar et al., “Squad: 100,000+ Questions for Machine Comprehension of Text,” arXiv:1606.05250, 2016, 10 pages. [cited by applicant]
Gratch et al., “Can virtual humans be more engaging than real ones?”, 12th International Conference on Human-Computer Interaction, Beijing, China, 2007, 10 pages. [cited by applicant]
Van Pinxteren et al., “Human-like communication in conversational agents: a literature review and research agenda,” Journal of Service Management, vol. 31 No. 2, 2020, pp. 203-225. [cited by applicant]
Rasipuram et al., “Automatic multimodal assessment of soft skills in social interactions: a review,” Multimedial Tools and Applications, 2020, 25 pages. [cited by applicant]
Price, “Ask Alexa or Google Home to Read your Child a Personalized Bedtime Story with this Skill,” https://lifehacker.com/ask-alexa-or-google-home-to-read-your-child-a-personali-1829249795, 2018, 3 pages. [cited by applicant]
Briones, “How This Digital Avatar is Elevating AI Technology,” ForbesLife, https://www.forbes.com/sites/isisbriones/2020/09/28/how-this-digital-avatar-is-elevating-ai-technology/?sh=696520a33a8a, 2020, 8 pages. [cited by applicant]
Zakharov et al., Few-Shot Adversarial Learning of Realistic Neural Talking Head Models, arXiv:1905.08233, 2019, pp. 1-21. [cited by applicant]
Cited By (7)
US 12,361,933 US 12,482,163 US 12,591,735 US 12,626,691 US 12,651,122 US 12,652,352 US 12,682,533