IP Library Granted Patent US 12,579,721
Granted Patent B2
US 12,579,721 · App. 18/081,076 · Granted Mar 17, 2026

Generating video content from user input data

Inventors: Robinson Piramuthu (Oakland, CA); Sanqiang Zhao (Santa Clara, CA); Yadunandana Rao (Sunnyvale, CA); Zhiyuan Fang (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G06T13/00G06F40/166G06F40/279G06F40/40G06T11/00G10L13/02G10L15/18G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,721
App. No.
18/081,076
Granted
Mar 17, 2026
Kind
B2
Abstract

Techniques for generating content associated with a user input/system generated response are described. Natural language data associated with a user input may be generated. For each portion of the natural language data, ambiguous references to entities in the portion may be replaced with the corresponding entity. Entities included in the portion may be extracted, and image data representing the entity may be determined. Background image data associated with the entities and the portion may be determined, and attributes which modify the entities in the natural language sentence may be extracted. Spatial relationships between two or more of the entities may further be extracted. Image data representing the natural language data may be generated based on the background image data, the entities, the attributes, and the spatial relationships. Video data may be generated based on the image data, where the video data includes animations of the entities moving.

Claims (117)

1 . A computer-implemented method comprising:

receiving first input audio data corresponding to a first spoken natural language input requesting a narrative be output, the first spoken natural language input indicating a first narrative parameter;

generating parameter data including the first narrative parameter included in the first spoken natural language input;

generating outline data based on the parameter data, wherein the outline data includes the first narrative parameter and a second narrative parameter not included in the parameter data;

processing, using a first trained machine learning (ML) component, the outline data to generate natural language data corresponding to the narrative requested in the first spoken natural language input, wherein the natural language data includes more words than the outline data, and the natural language data comprises:

a first portion corresponding to a first scene of the narrative, and

a second portion corresponding to a second scene of the narrative;

processing, using a second trained ML component, the first portion of the natural language data to determine a first entity represented in the first portion of the natural language data;

determining, using the second trained ML component, first image data corresponding to the first entity;

processing, using a third trained ML component, the first portion of the natural language data to determine first background image data corresponding to the first scene of the narrative;

processing, using a fourth trained ML component, the first portion of the natural language data and the first entity to determine an attribute corresponding to the first entity, wherein the attribute represents how the first entity is to be presented;

processing, using a fifth trained ML component, the first image data, the first background image data, and the attribute to generate first scene data, wherein the first scene data indicates how the first image data is to be rendered with the first background image data based on the attribute;

based on the first scene data, generating first output image data including the first image data and the first background image data, the first output image data corresponding to the first scene of the narrative;

processing, using the second trained ML component, the second portion of the natural language data to determine the first entity is represented in the second portion of the natural language data;

determining the first image data is to be used to render the first entity in the second scene of the narrative based on the first image data being used to represent the first entity in the first scene data;

generating second output image data including the first image data, the second output image data corresponding to the second scene of the narrative;

causing presentation of the first output image data; and

causing presentation of the second output image data.

2 . The computer-implemented method of claim 1 , further comprising:

storing a representation of the first entity, the first image data, the first background image data, and the attribute in association with a content request identifier corresponding to the first spoken natural language input;

retrieving the representation of the first entity;

based at least in part on determining that the first entity is represented in the second portion, retrieving the first image data;

processing, using the second trained ML component, the second portion of the natural language data to determine a second entity represented in the second portion of the natural language data;

determining, using the second trained ML component, second image data corresponding to the second entity;

processing, using the third trained ML component, the second portion of the natural language data to determine second background image data for the second scene of the narrative; and

processing, using the fifth trained ML component, the first image data, the second image data, and the second background image data to generate second scene data, wherein the second scene data indicates how the first image data and the second image data are to be rendered with the second background image data.

3 . The computer-implemented method of claim 1 , further comprising:

processing, using a sixth ML component, the natural language data to determine a third portion of the natural language data that corresponds to the first narrative parameter; and

replacing the third portion of the natural language data with the first narrative parameter.

4 . The computer-implemented method of claim 1 , further comprising:

processing, using a sixth ML component, the natural language data and the first entity to determine a portion of the first background image data where the first image data is to be located when it is rendered with the first background image data,

wherein the first scene data indicates the portion of the first background image data where the first image data is to be located when it is rendered with the first background image data, and

wherein the second output image data is generated to include the first image data located with respect to the first background image data as indicated in the first scene data.

5 . A computer-implemented method comprising:

receiving first input data requesting content be output, the first input data indicating a first parameter;

generating outline data based on the first parameter, wherein the outline data includes the first parameter and a second parameter not included in the first input data;

based on the outline data, generating natural language data corresponding to the content requested in the first input data, wherein the natural language data includes more words than the outline data;

processing, using a trained machine learning (ML) component, a first portion of the natural language data to determine a first entity included in the first portion of the natural language data;

determining, using the trained ML component, first image data corresponding to the first entity;

determining first background image data representing the natural language data;

generating first scene data indicating how the first image data is to be rendered with the first background image data;

based on the first scene data, generating first output image data including the first image data and the first background image data, the first output image data representing at least the first portion of the natural language data;

determining, by the trained ML component, the first entity is included in a second portion of the natural language data;

based on the first entity being included in the second portion of the natural language data and the first image data being used to represent the first entity in the first scene data, generating second output image data including the first image data, the second output image data representing at least the second portion of the natural language data;

causing presentation of the first output image data; and

causing presentation of the second output image data.

6 . The computer-implemented method of claim 5 , further comprising:

determining, in the first portion of the natural language data, an attribute corresponding to the first entity, wherein the attribute represents how the first entity is to be presented; and

generating the first scene data to indicate how the first image data is to be rendered using the attribute.

7 . The computer-implemented method of claim 5 , further comprising:

determining a spatial relationship between the first entity and a second entity included in the first portion of the natural language data, wherein the spatial relationship represents how the first entity is to be rendered with the second entity.

8 . The computer-implemented method of claim 5 , further comprising:

determining a second entity included in the second portion of the natural language data;

determining second image data corresponding to the second entity;

determining second background image data representing the second portion of the natural language data; and

generating second scene data indicating how the first image data and the second image data are to be rendered with the second background image data.

9 . The computer-implemented method of claim 5 , further comprising:

determining a third portion of the natural language data corresponding to a second entity;

determining the second entity corresponds to the first entity; and

based on the second entity corresponding to the first entity, replacing the second entity with the first entity in the third portion of the natural language data.

10 . The computer-implemented method of claim 5 , further comprising:

performing text-to-speech (TTS) processing using the natural language data to generate first output audio data comprising:

a first portion corresponding to the first portion of the natural language data, and

a second portion corresponding to the second portion of the natural language data;

causing the first output audio data to be presented;

while the first portion of the first output audio data is being presented, causing the first output image data to be presented; and

after the first portion of the first output audio data is presented, and while the second portion of the first output audio data is being presented, causing the second output image data to be presented.

11 . The computer-implemented method of claim 5 , further comprising:

generating third output image data corresponding to the first output image data, wherein the first entity is represented differently in the first output image data than in the third output image data; and

generating output video data using the first output image data and the third output image data.

12 . The computer-implemented method of claim 5 , wherein:

the first input data requests a narrative be output,

the first output image data corresponds to a first portion of the narrative, and

the second output image data corresponds to a second portion of the narrative.

13 . A computing system comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the computing system to:

receive first input data requesting content be output, the first input data indicating a first parameter;

generate outline data based on the first parameter, wherein the outline data includes the first parameter and a second parameter not included in the first input data;

based on the outline data, generate natural language data corresponding to the content requested in the first input data, wherein the natural language data includes more words than the outline data;

process, using a trained machine learning (ML) component, a first portion of the natural language data to determine a first entity included in the first portion of the natural language data;

determine, using the trained ML component, first image data corresponding to the first entity;

determine first background image data representing the natural language data;

generate first scene data indicating how the first image data is to be rendered with the first background image data;

based on the first scene data, generate first output image data including the first image data and the first background image data, the first output image data representing at least the first portion of the natural language data;

determine, by the trained ML component, the first entity is included in a second portion of the natural language data;

based on the first entity being included in the second portion of the natural language data and the first image data being used to represent the first entity in the first scene data, generate second output image data including the first image data, the second output image data representing at least the second portion of the natural language data;

cause presentation of the first output image data; and

cause presentation of the second output image data.

14 . The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

determine, in the first portion of the natural language data, an attribute corresponding to the first entity, wherein the attribute represents how the first entity is to be presented; and

generate the first scene data to indicate how the first image data is to be rendered using the attribute.

15 . The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

determine a spatial relationship between the first entity and a second entity included in the first portion of the natural language data, wherein the spatial relationship represents how the first entity is to be rendered with the second entity.

16 . The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

determine a second entity included in the second portion of the natural language data;

determine second image data corresponding to the second entity;

determine second background image data representing the second portion of the natural language data; and

generate second scene data indicating how the first image data and the second image data are to be rendered with the second background image data.

17 . The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

determine a third portion of the natural language data corresponding to a second entity;

determine the second entity corresponds to the first entity; and

based on the second entity corresponding to the first entity, replace the second entity with the first entity in the third portion of the natural language data.

18 . The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

perform text-to-speech (TTS) processing using the natural language data to generate first output audio data comprising:

a first portion corresponding to the first portion of the natural language data, and

a second portion corresponding to the second portion of the natural language data;

cause the first output audio data to be presented;

while the first portion of the first output audio data is being presented, cause the first output image data to be presented; and

after the first portion of the first output audio data is presented, and while the second portion of the first output audio data is being presented, cause the second output image data to be presented.

19 . The computing system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the computing system to:

generate third output image data corresponding to the first output image data, wherein the first entity is represented differently in the first output image data than in the third output image data; and

generate output video data using the first output image data and the third output image data.

20 . The computing system of claim 13 , wherein:

the first input data requests a narrative be output,

the first output image data corresponds to a first portion of the narrative, and

the second output image data corresponds to a second portion of the narrative.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: PIRAMUTHU, ROBINSON; ZHAO, SANQIANG; RAO, YADUNANDANA; FANG, ZHIYUAN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 062088/0968 →
Continuity (2)
Provisional Application 63407944 · Sep 19, 2022
Related Publication 20240095987A1 · Mar 21, 2024
References Cited (37)
US 7664313B1 · Sproat · 2010 [cited by applicant]
US 10922049B2 · Bolden · 2021 [cited by examiner]
US 10999566B1 · Mahyar et al. · 2021 [cited by applicant]
US 11651537B2 · Duffy · 2023 [cited by examiner]
US 11863829B2 · Gao et al. · 2024 [cited by applicant]
US 20060217979A1 · Pahud · 2006 [cited by examiner]
US 20070147654A1 · Clatworthy · 2007 [cited by examiner]
US 20150331941A1 · Defouw et al. · 2015 [cited by applicant]
US 20190108219A1 · Barrett et al. · 2019 [cited by applicant]
US 20190304104A1 · Amer · 2019 [cited by examiner]
US 20190304156A1 · Amer · 2019 [cited by examiner]
US 20240194193A1 · Rathnam et al. · 2024 [cited by applicant]
US 20240331434A1 · Sabapathy et al. · 2024 [cited by applicant]
US 20240420404A1 · Kasap et al. · 2024 [cited by applicant]
US 20250021833A1 · Mukherjee et al. · 2025 [cited by applicant]
Yao et al. (NPL—“Plan-And-Write: Towards Better Automatic Storytelling”) Feb. 19, 2019 Association for the Advancement of AI 2019 (Year: 2019). [cited by examiner]
Avrahami et al. (NPL—“Blended Diffusion for Text-driven Editing of Natural Images”) Mar. 28, 2022 The Hebrew University of Jerusalem & Reichman University (Year: 2022). [cited by examiner]
International Search Report and Written Opinion mailed Nov. 30, 2023 for International Patent Application No. PCT/US2023/073510. [cited by applicant]
Coyne et al., “WordsEye: an automatic text-to-scene conversion system”, Proceedings of the 2001 International ACM SIGGroup Conference on Supporting Group Work, ACM, New York, NY, US, Aug. 1, 2002, pp. 487-496, XP0591481… [cited by applicant]
Yao, et al., “Plan-and-Write: Towards Better Automatic Storytelling”, The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), 2019, 7378-7385, https://doi.org/10.1609/aaai.v33i01.33017378. [cited by applicant]
Ghazarian, et al. “Plot-guided Adversarial Example Construction for Evaluating Open-domain Story Generation”, NAACL 2021, https://arxiv.org/abs/2104.05801. [cited by applicant]
Engel, et al. “GANSynth: Adversarial Neural Audio Synthesis”, ICLR 2019, https://arxiv.org/abs/1902.08710. [cited by applicant]
Huang, et al., “Music Transformer: Generating Music With Long-Term Structure”, Sep. 12, 2018, https://arxiv.org/abs/1809.04281. [cited by applicant]
Dhariwal, et al. “Jukebox: A Generative Model for Music”, Apr. 30, 2020, https://arxiv.org/abs/2005.00341. [cited by applicant]
Turney, “Thumbs Up or Thumbs Down? Semantic Orientation Applied to Unsupervised Classification of Reviews”, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 2002, 417-424, https:/… [cited by applicant]
Pennington, et al. “GloVe: Global Vectors for Word Representation”, 2014, https://nlp.stanford.edu/projects/glove/. [cited by applicant]
Waite, “Generating Long-Term Structure in Songs and Stories”, Magenta, Jul. 15, 2016, https://magenta.tensorflow.org/2016/07/15/lookback-rnn-attention-rnn. [cited by applicant]
Rombach, et al. “High-Resolution Image Synthesis with Latent Diffusion Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 10684-10695, https://arxiv.org/abs/2112.107… [cited by applicant]
Nichol, et al., “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models”, Dec. 20, 2021, https://arxiv.org/abs/2112.10741. [cited by applicant]
Liu, et al., “Paint Transformer: Feed Forward Neural Painting with Stroke Prediction”, In Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6598-6607, https://arxiv.org/abs/2108.03798. [cited by applicant]
“Text2Scene: GeneratingCompositional Scenes from Textual Descriptions”, CVPR, Rice University, 2019, https://www.vislang.ai/text2scene. [cited by applicant]
Ramesh, et al. “Hierarchical Text-Conditional Image Generation with CLIP Latents”, 2022, https://arxiv.org/abs/2204.06125. [cited by applicant]
Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, 2022, https://arxiv.org/abs/2205.11487. [cited by applicant]
Lee, et al., “Higher-order Coreference Resolution with Coarse-to-fine Inference”, NAACL-HLT, 2018, https://arxiv.org/abs/1804.05392. [cited by applicant]
Honnibal, et al., “A Non-Monotonic Arc-Eager Transition System for Dependency Parsing”, Proceedings of the Seventeenth Conference on Computational Natural Language Learning, 2013, pp. 163-172, https://aclanthology.org/W… [cited by applicant]
Khashabi, et al., “UnifiedQA: Crossing Format Boundaries with a Single QA System”, EMNLP Findings, 2020, https://arxiv.org/abs/2005.00700. [cited by applicant]
FitzGerald, et al., “Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Endcoders for Natural Language Understanding Systems”, Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery an… [cited by applicant]