IP Library Granted Patent US 10,789,755
Granted Patent B2
US 10,789,755 · App. 16/230,945 · Granted Sep 29, 2020

Artificial intelligence in interactive storytelling

Inventors: Mohamed R. Amer (Brooklyn, NY); Timothy J. Meo (Kendall Park, NJ); Aswin Nadamuni Raghavan (Princeton, NJ); Alex C. Tozzo (Philadelphia, PA); Amir Tamrakar (Princeton, NJ); David A. Salter (Basking Ridge, PA); Kyung-Yoon Kim (North York, CA)
Assignee: SRI International
G06T13/80G06F16/345G06F16/738G06F16/9024G06F40/205G06K9/00335G06K9/00342G06K9/00718G06K9/00744G06K9/6215G06N3/0445G06N3/0454G06N3/08G06N3/084G06N3/088G06N20/00G06T7/251G06T13/40G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,789,755
App. No.
16/230,945
Granted
Sep 29, 2020
Kind
B2
Abstract

This disclosure describes techniques that include generating, based on a description of a scene, a movie or animation that represents at least one possible version of a story corresponding to the description of the scene. This disclosure also describes techniques for training a machine learning model to generate predefined data structures from textual information, visual information, and/or other information about a story, an event, a scene, or a sequence of events or scenes within a story. This disclosure also describes techniques for using GANs to generate, from input, an animation of motion (e.g., an animation or a video clip). This disclosure also describes techniques for implementing an explainable artificial intelligence system that may provide end users with information (e.g., through a user interface) that enables an understanding of at least some of the decisions made by the AI system.

Claims (86)

1. A method comprising:

receiving, by a computing system, information about a sequence of events involving a plurality of objects, wherein the information about the sequence of events includes textual information identifying the plurality of objects, at least one attribute associated with each of the plurality of objects, and at least one relationship between two or more of the plurality of objects;

generating, by a machine learning system configured to execute on the computing system and based on the information about the sequence of events, a data structure specifying the attributes associated with the objects and further specifying relationships between at least some of the plurality of objects;

translating, by the computing system, the data structure into a plurality of commands understood by an animation generation system, wherein the commands include a command to configure an animation to include the attributes associated with the objects, and a command to configure the animation to reflect the relationships between the at least some of the plurality of objects; and

automatically generating, by outputting the plurality of commands to the animation generation system, the animation illustrating the sequence of events.

2. The method of claim 1 , further comprising:

prior to generating the data structure, training the machine learning system with a training data set.

3. The method of claim 1 , wherein the information about the sequence of events includes story information and user input,

wherein the story information includes at least one of a script, dialogue information, slug line information, audio information, or video information; and

wherein the user input includes a user command, information about a movement, information about a gesture, information about a gaze, and information about a pose.

4. The method of claim 1 ,

wherein the plurality of objects include a plurality of actors and a plurality of props;

wherein each of the attributes identify, for each object that the attribute is associated with, at least one of: an object type, a location, or a description; and

wherein the relationship between the two or more of the plurality of objects includes at least one of: a spatial relationship, a temporal relationship, or an action.

5. The method of claim 3 , wherein the user input includes at least one of a plurality of user commands including:

a command to create one of the plurality of actors;

a command to specify one or more attributes of one of the plurality of objects;

a command to identify an action performed by one of the plurality of actors;

a command to specify a spatial relationship between two or more of the plurality of objects;

a command to specify a temporal relationship between two or more of the plurality of objects; and

a command confirming a change to the data structure proposed by either the computing system or a user.

6. The method of claim 3 , wherein the plurality of objects includes a plurality of actors, the method further comprising at least one of:

detecting, by the computing system, a user's movement to determine the information about the movement, and associating the information about the movement with at least one of the plurality of actors;

detecting, by the computing system, a user's gesture to determine the information about the gesture, and associating the information about the gesture with at least one of the plurality of actors;

detecting, by the computing system, a user's gaze to determine the information about the gaze, and associating the information about the gaze with at least one of the plurality of actors; and

detecting, by the computing system, a user's pose to determine the information about the pose, and associating the information about the pose with at least one of the plurality of actors.

7. The method of claim 1 , further comprising:

receiving, by the computing system, further information about the sequence of events;

modifying, by the computing system, the data structure to generate an updated data structure reflecting the further information about the sequence of events;

updating, based on the updated data structure, the plurality of commands; and

automatically generating, by outputting the updated plurality of commands to the animation system, an updated animation.

8. The method of claim 7 , wherein receiving the further information about the sequence of events includes:

identifying an ambiguity associated with the data structure;

prompting a user for information about the ambiguity; and

resolving, by the computing system and based on input received after prompting the user for information about the ambiguity, the ambiguity.

9. The method of claim 1 , wherein the data structure is a spatio-temporal composition graph.

10. The method of claim 9 , wherein the spatio-temporal composition graph includes nodes identifying actors, attributes of actors, actions performed by actors, and physical objects in the scene.

11. A method of identifying content comprising:

receiving, by a computing system, information about a sequence of events involving a plurality of objects and actions, wherein the information about the sequence of events includes textual information identifying the plurality of objects, at least one attribute associated with each of the plurality of objects, and at least one relationship between one object and an associated action and a second object and its associated action;

generating, by a machine learning system configured to execute on the computing system and based on the information about the sequence of events, a spatio-temporal composition graph specifying the attributes associated with the objects and further specifying relationships between at least some of the plurality of objects;

translating, by the computing system, the spatio-temporal composition graph into animation information understood by an animation generation system, wherein the animation information includes a command to configure an animation to include the attributes associated with the objects, and a command to configure the animation to reflect the relationships between the at least some of the plurality of objects; and

detecting, by the computing system, input that includes a query about the content;

determining a response to the query by analyzing, by the computing system, the animation information; and

outputting, by the computing system, the response to satisfy the query.

12. A system comprising:

a storage device; and

processing circuitry having access to the storage device and configured to:

receive information about a sequence of events involving a plurality of objects, wherein the information about the sequence of events includes textual information identifying the plurality of objects, at least one attribute associated with each of the plurality of objects, and at least one relationship between two or more of the plurality of objects,

generate, based on the information about the sequence of events, a data structure specifying the attributes associated with the objects and further specifying relationships between at least some of the plurality of objects, and

translate the data structure into a plurality of commands understood by an animation generation system, wherein the commands include a command to configure an animation to include the attributes associated with the objects, and a command to configure the animation to reflect the relationships between the at least some of the plurality of objects; and

automatically generate, by outputting the plurality of commands to the animation generation system, the animation illustrating the sequence of events.

13. The system of claim 12 , wherein the processing circuitry is further configured to:

prior to generating the data structure, train the machine learning system with a training data set.

14. The system of claim 12 , wherein the information about the sequence of events includes story information and user input,

wherein the story information includes at least one of a script, dialogue information, slug line information, audio information, or video information; and

wherein the user input includes a user command, information about a movement, information about a gesture, information about a gaze, and information about a pose.

15. The system of claim 12 ,

wherein the plurality of objects include a plurality of actors and a plurality of props;

wherein each of the attributes identify, for each object that the attribute is associated with, at least one of: an object type, a location, or a description; and

wherein the relationship between the two or more of the plurality of objects includes at least one of: a spatial relationship, a temporal relationship, or an action.

16. The system of claim 14 , wherein the user input includes a plurality of user commands including:

a command to create one of the plurality of actors;

a command to specify one or more attributes of one of the plurality of objects;

a command to identify an action performed by one of the plurality of actors;

a command to specify a spatial relationship between two or more of the plurality of objects; and

a command to specify a temporal relationship between two or more of the plurality of objects.

17. The system of claim 14 , wherein the plurality of objects includes a plurality of actors, and wherein the processing circuitry is further configured to:

detect a user's movement to determine the information about the movement, and associating the information about the movement with at least one of the plurality of actors;

detect a user's gesture to determine the information about the gesture, and associating the information about the gesture with at least one of the plurality of actors;

detect a user's gaze to determine the information about the gaze, and associating the information about the gaze with at least one of the plurality of actors; and

detect a user's pose to determine the information about the pose, and associating the information about the pose with at least one of the plurality of actors.

18. The system of claim 12 , wherein the processing circuitry is further configured to:

receive further information about the sequence of events;

modify the data structure to generate an updated data structure reflecting the further information about the sequence of events;

update, based on the updated data structure, the plurality of commands; and

automatically generate, by outputting the updated plurality of commands, an updated animation.

19. The system of claim 18 , wherein to receive the further information about the sequence of events, the processing circuitry is further configured to:

identify an ambiguity associated with the data structure;

prompt a user for information about the ambiguity; and

resolve, based on input received after prompting the user for information about the ambiguity, the ambiguity.

20. The system of claim 12 , wherein the data structure is a spatio-temporal composition graph that includes nodes identifying actors, attributes of actors, actions performed by actors, and physical objects in the scene.

21. A non-transitory computer-readable storage medium comprising instructions that, when executed, configure processing circuitry of a computing system to:

receive information about a sequence of events involving a plurality of objects, wherein the information about the sequence of events includes textual information identifying the plurality of objects, at least one attribute associated with each of the plurality of objects, and at least one relationship between two or more of the plurality of objects;

generate, based on the information about the sequence of events, a data structure specifying the attributes associated with the objects and further specifying relationships between at least some of the plurality of objects;

translate the data structure into a plurality of commands understood by an animation generation system, wherein the commands include a command to configure an animation to include the attributes associated with the objects, and a command to configure the animation to reflect the relationships between the at least some of the plurality of objects; and

automatically generate, by outputting the plurality of commands to the animation generation system, the animation illustrating the sequence of events.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2025
From: SRI INTERNATIONAL
To: GLENEAGLE INNOVATIONS LP
Reel/Frame 071968/0850 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 5, 2019
From: AMER, MOHAMED R.; MEO, TIMOTHY J.; NADAMUNI RAGHAVAN, ASWIN; TOZZO, ALEX C.; TAMRAKAR, AMIR; SALTER, DAVID A.; KIM, KYUNG-YOON
To: SRI INTERNATIONAL
Reel/Frame 048509/0094 →
Continuity (3)
Provisional Application 62652195 · Apr 3, 2018
Provisional Application 62778754 · Dec 12, 2018
Related Publication 20190304157A1 · Oct 3, 2019
Cited By (12)
US 12,205,028 US 12,248,898 US 12,307,799 US 12,314,679 US 12,321,454 US 12,333,794 US 12,347,283 US 12,361,741 US 12,394,283 US 12,499,887 US 12,561,522 US 12,650,866