IP Library Granted Patent US 11,210,836
Granted Patent B2
US 11,210,836 · App. 16/231,020 · Granted Dec 28, 2021

Applying artificial intelligence to generate motion information

Inventors: Mohamed R. Amer (Brooklyn, NY); Xiao Lin (Princeton, NJ)
Assignee: SRI International
G06T13/80G06F16/345G06F16/738G06F16/9024G06F40/205G06K9/00335G06K9/00342G06K9/00718G06K9/00744G06K9/6215G06N3/0445G06N3/0454G06N3/08G06N3/084G06N3/088G06N20/00G06T7/251G06T13/40G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,210,836
App. No.
16/231,020
Granted
Dec 28, 2021
Kind
B2
Abstract

This disclosure describes techniques that include generating, based on a description of a scene, a movie or animation that represents at least one possible version of a story corresponding to the description of the scene. This disclosure also describes techniques for training a machine learning model to generate predefined data structures from textual information, visual information, and/or other information about a story, an event, a scene, or a sequence of events or scenes within a story. This disclosure also describes techniques for using GANs to generate, from input, an animation of motion (e.g., an animation or a video clip). This disclosure also describes techniques for implementing an explainable artificial intelligence system that may provide end users with information (e.g., through a user interface) that enables an understanding of at least some of the decisions made by the AI system.

Claims (44)

1. A method for generating information about motion, comprising:

receiving, by a computing system, text input indicating motion by an actor;

applying, by the computing system, a generative adversarial network (GAN) model trained using a training set that includes abstracted animations of motion that include less data than pixel-level representations of motion, to the text input to generate motion information, including motion sequence, without conditioning, the motion information representing the motion by the actor, wherein the GAN model is trained using less computational resources than a GAN model trained using pixel-level representations of motion; and

outputting, by the computing system, the motion information.

2. The method of claim 1 , wherein motion information includes at least one of: skeleton movement, joint angles between body parts, a mesh of a human body, a facial expression, movement of body landmarks, movement of facial feature landmarks.

3. The method of claim 1 ,

wherein applying the GAN model includes applying a GAN model trained to generate a predicted abstracted animation of motion from the text input indicating motion; and

wherein outputting the motion information includes outputting an abstracted animation of motion representing the motion by the actor.

4. The method of claim 3 ,

wherein the GAN uses at least one of a convolutional neural network (CNN) and a recurrent neural network (RNN) to generate the motion information; and

wherein the GAN combines the CNN and the RNN for motion completion.

5. The method of claim 4 , wherein the GAN includes a gradient penalty.

6. The method of claim 5 , wherein the generative adversarial network is a dense validation generative adversarial network (DVGAN).

7. The method of claim 6 , wherein training the GAN includes:

training the GAN using a discriminator that produces scores at a dense frequency of time resolutions and at a dense frequency of frames.

8. The method of claim 1 ,

wherein applying the GAN model includes applying a GAN model trained by evaluating each of the animations over a set of time resolutions and over a set of frame frequencies.

9. The method of claim 1 ,

wherein training the GAN includes producing a GAN validation score at each of the time resolutions and at each of the frame frequencies.

10. The method of claim 1 ,

wherein the abstracted animations of motion are derived from video clips drawn from movies.

11. A system comprising:

a storage device; and

processing circuitry having access to the storage device and configured to:

receive text input indicating motion by an actor,

apply a generative adversarial network (GAN) model, trained using a training set that includes abstracted animations of motion that include less data than pixel-level representations of motion, to the text input to generate motion information including motion sequence representing the motion by the actor, wherein the GAN model is trained using less computational resources than a GAN model trained using pixel-level representations of motion, and

output the motion information.

12. The system of claim 11 , wherein motion information includes at least one of: skeleton movement, joint angles between body parts, a mesh of a human body, a facial expression, movement of body landmarks, movement of facial feature landmarks.

13. The system of claim 11 , wherein to apply the GAN, the processing circuitry is further configured to:

apply a GAN model trained by evaluating each of the animations over a set of time resolutions and over a set of frame frequencies.

14. The system of claim 11 , wherein to apply the GAN model, the processing circuitry is further configured to:

apply a GAN model trained to generate a predicted abstracted animation of motion from the text input indicating motion; and

wherein outputting the motion information includes outputting an abstracted animation of motion representing the motion by the actor.

15. The system of claim 14 ,

wherein the GAN uses at least one of a convolutional neural network (CNN) and a recurrent neural network (RNN) to generate the motion information; and

wherein the GAN combines the CNN and the RNN for motion completion.

16. The system of claim 15 , wherein the GAN includes a Wasserstein GAN with a gradient penalty.

17. The system of claim 16 , wherein the generative adversarial network is a dense validation generative adversarial network (DVGAN).

18. The system of claim 16 , wherein to train the GAN, the processing circuitry is further configured to:

use a discriminator that produces scores at a dense frequency of time resolutions and at a dense frequency of frames.

19. A non-transitory computer-readable storage medium comprising instructions that, when executed, configure processing circuitry of a computing system to:

receive text input indicating motion by an actor;

apply a generative adversarial network (GAN) model, trained using a training set that includes abstracted animations of motion and trained by evaluating each of the abstracted animations of motion over a set of time resolutions and over a set of frame frequencies, to the text input to generate motion information including motion sequence representing the motion by the actor; and

output the motion information.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2025
From: SRI INTERNATIONAL
To: GLENEAGLE INNOVATIONS LP
Reel/Frame 071968/0850 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2019
From: AMER, MOHAMED; LIN, XIAO
To: SRI INTERNATIONAL
Reel/Frame 048256/0363 →
Continuity (3)
Provisional Application 62652195 · Apr 3, 2018
Provisional Application 62778754 · Dec 12, 2018
Related Publication 20190304104A1 · Oct 3, 2019
Cited By (3)
US 12,314,306 US 12,576,873 US 12,693,930