IP Library › Granted Patent US 12,361,623
Granted Patent B2
US 12,361,623 · App. 18/095,501 · Granted Jul 15, 2025

Avatar generation and augmentation with auto-adjusted physics for avatar motion

Inventor: Michael Taylor (San Bruno, CA)
Assignee: Sony Interactive Entertainment Inc.
G06T13/40A63F13/57A63F13/63G06F3/04845G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,623
App. No.
18/095,501
Granted
Jul 15, 2025
Kind
B2
Abstract

A method including accessing an avatar including a plurality of physiological characteristics, wherein the avatar includes a plurality of motion profiles. The method including receiving editing of a motion profile of the avatar via a user interface. The method including modifying each of the plurality of motion profiles to be consistent with the motion profile that edited. The method including automatically generating a prompt to apply the motion profile that is edited to the avatar, wherein the prompt is provided as input for an image generation artificial intelligence system configured to implement latent diffusion. The method including generating a sequence of video frames using the image generation artificial intelligence system showing the avatar in motion based on the motion profile that is edited. The method including presenting the sequence of video frames in the user interface.

Claims (84)

1. A method, comprising:

accessing an avatar including a plurality of physiological characteristics and a plurality of motion profiles, wherein each motion profile is predefined and defines a plurality of distinct motions performed by the avatar in a virtual environment, and wherein each motion profile is represented as a sequence of images having a length comprising a starting key frame including a first pose of the avatar and an ending key frame including a second pose of the avatar;

receiving editing of a motion profile in the plurality of motions profiles of the avatar via a user interface, wherein editing of the motion profile comprises an adjustment to a parameter associated with the length of the sequence of images;

modifying each of the plurality of motion profiles to be consistent with the motion profile that is edited by automatically generating a prompt to apply, for each motion profile of the plurality of motion profiles, the motion profile that is edited to the starting key frame and the ending key frame, wherein the prompt is provided as input for an image generation artificial intelligence system configured to implement latent diffusion;

generating a sequence of video frames using the image generation artificial intelligence system showing the avatar in motion based on the prompt, the motion profile that is edited, and the starting key frame and the ending key frame; and

presenting the sequence of video frames in the user interface.

2. The method of claim 1 , further comprising:

receiving, via the user interface, feedback regarding the motion profile that is edited;

providing further editing of the motion profile that is edited based on the feedback;

generating another prompt to apply the motion profile that is further edited to the avatar; and

generating another sequence of video frames using the image generation artificial intelligence system showing the avatar in another motion based on the motion profile that is further edited.

3. The method of claim 1 , further comprising:

receiving editing of a physiological characteristic of the avatar; and

editing each of the plurality of motion profiles of the avatar to be consistent with the physiological characteristic that is edited.

4. The method of claim 3 , further comprising:

displaying the avatar in the user interface; and

receiving manipulation of the avatar to generate the editing of the physiological characteristic.

5. The method of claim 1 , further comprising:

receiving editing of a physiological characteristic of the avatar via the user interface;

editing the motion profile of the avatar to be consistent with the physiological characteristic that is edited; and

modifying each of the plurality of motion profiles to be consistent with the motion profile that edited.

6. The method of claim 1 , wherein the editing a motion profile includes at least one of:

receiving editing of a frequency of the motion profile;

receiving editing of an amplitude of the motion profile; and

receiving editing of a speed of the motion profile.

7. The method of claim 1 ,

wherein the length of the sequence of images defines a cycle of complete motion of the avatar, and wherein the adjustment to the parameter associated with the length comprises:

increasing or decreasing the length of the cycle, wherein increasing the length of the cycle increases a speed of motion of the avatar, and wherein decreasing the length of the cycle decreases the speed of motion of the avatar.

8. A non-transitory computer-readable medium storing a computer program for performing a method, the computer-readable medium comprising:

program instructions for accessing an avatar including a plurality of physiological characteristics and a plurality of motion profiles, wherein each motion profile is predefined and defines a plurality of distinct motions performed by the avatar in a virtual environment, and wherein each motion profile is represented as a sequence of images having a length comprising a starting key frame including a first pose of the avatar and an ending key frame including a second pose of the avatar;

program instructions for receiving editing of a motion profile in the plurality of motions profiles of the avatar via a user interface, wherein editing of the motion profile comprises an adjustment to a parameter associated with the length of the sequence of images;

program instructions for modifying each of the plurality of motion profiles to be consistent with the motion profile that is edited by automatically generating a prompt to apply, for each motion profile of the plurality of motion profiles, the motion profile that is edited to the starting key frame and the ending key frame, wherein the prompt is provided as input for an image generation artificial intelligence system configured to implement latent diffusion;

program instructions for generating a sequence of video frames using the image generation artificial intelligence system showing the avatar in motion based on the prompt, the motion profile that is edited, and the starting key frame and the ending key frame; and

program instructions for presenting the sequence of video frames in the user interface.

9. The non-transitory computer-readable medium of claim 8 , further comprising:

program instructions for receiving, via the user interface, feedback regarding the motion profile that is edited;

program instructions for providing further editing of the motion profile that is edited based on the feedback;

program instructions for generating another prompt to apply the motion profile that is further edited to the avatar; and

program instructions for generating another sequence of video frames using the image generation artificial intelligence system showing the avatar in another motion based on the motion profile that is further edited.

10. The non-transitory computer-readable medium of claim 8 , further comprising:

program instructions for receiving editing of a physiological characteristic of the avatar; and

program instructions for editing each of the plurality of motion profiles of the avatar to be consistent with the physiological characteristic that is edited.

11. The non-transitory computer-readable medium of claim 10 , further comprising:

program instructions for displaying the avatar in the user interface; and

program instructions for receiving manipulation of the avatar to generate the editing of the physiological characteristic.

12. The non-transitory computer-readable medium of claim 8 , further comprising:

program instructions for receiving editing of a physiological characteristic of the avatar via the user interface;

program instructions for editing the motion profile of the avatar to be consistent with the physiological characteristic that is edited; and

program instructions for modifying each of the plurality of motion profiles to be consistent with the motion profile that edited.

13. The non-transitory computer-readable medium of claim 8 ,

wherein the editing a motion profile includes at least one of:

program instructions for receiving editing of a frequency of the motion profile;

program instructions for receiving editing of an amplitude of the motion profile; and

program instructions for receiving editing of a speed of the motion profile.

14. The non-transitory computer-readable medium of claim 8 , further comprising:

wherein the length of the sequence of images defines a cycle of complete motion of the avatar, and wherein the adjustment to the parameter associated with the length comprises:

program instructions for increasing or decreasing the length of the cycle, wherein increasing the length of the cycle increases a speed of motion of the avatar, and wherein decreasing the length of the cycle decreases the speed of motion of the avatar.

15. A computer system comprising:

a processor;

a memory coupled to the processor and having stored therein instructions that, if executed by the computer system, cause the computer system to execute a method, comprising:

accessing an avatar including a plurality of physiological characteristics and a plurality of motion profiles, wherein each motion profile is predefined and defines a plurality of distinct motions performed by the avatar in a virtual environment, and wherein each motion profile is represented as a sequence of images having a length comprising a starting key frame including a first pose of the avatar and an ending key frame including a second pose of the avatar;

receiving editing of a motion profile in the plurality of motions profiles of the avatar via a user interface, wherein editing of the motion profile comprises an adjustment to a parameter associated with the length of the sequence of images;

modifying each of the plurality of motion profiles to be consistent with the motion profile that is edited by automatically generating a prompt to apply, for each motion profile of the plurality of motion profiles, the motion profile that is edited to the starting key frame and the ending key frame, wherein the prompt is provided as input for an image generation artificial intelligence system configured to implement latent diffusion;

generating a sequence of video frames using the image generation artificial intelligence system showing the avatar in motion based on the prompt, the motion profile that is edited, and the starting key frame and the ending key frame; and

presenting the sequence of video frames in the user interface.

16. The computer system of claim 15 , the method further comprising:

receiving, via the user interface, feedback regarding the motion profile that is edited;

providing further editing of the motion profile that is edited based on the feedback;

generating another prompt to apply the motion profile that is further edited to the avatar; and

generating another sequence of video frames using the image generation artificial intelligence system showing the avatar in another motion based on the motion profile that is further edited.

17. The computer system of claim 15 , the method further comprising:

receiving editing of a physiological characteristic of the avatar; and

editing each of the plurality of motion profiles of the avatar to be consistent with the physiological characteristic that is edited.

18. The computer system of claim 15 , the method further comprising:

receiving editing of a physiological characteristic of the avatar via the user interface;

editing the motion profile of the avatar to be consistent with the physiological characteristic that is edited; and

modifying each of the plurality of motion profiles to be consistent with the motion profile that edited.

19. The computer system of claim 15 , wherein in the method the editing a motion profile includes at least one of:

receiving editing of a frequency of the motion profile;

receiving editing of an amplitude of the motion profile; and

receiving editing of a speed of the motion profile.

20. The computer system of claim 15 ,

wherein the length of the sequence of images defines a cycle of complete motion of the avatar, and wherein the adjustment to the parameter associated with the length comprises:

increasing or decreasing the length of the cycle, wherein increasing the length of the cycle increases a speed of motion of the avatar, and wherein decreasing the length of the cycle decreases the speed of motion of the avatar.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2023
From: TAYLOR, MICHAEL
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 063631/0120 →
Continuity (1)
Related Publication 20240233231A1 · Jul 11, 2024
References Cited (22)
US 10325417B1 · Scapel · 2019 [cited by examiner]
US 10388071B2 · Rico · 2019 [cited by examiner]
US 20060143569A1 · Kinsella · 2006 [cited by examiner]
US 20160086500A1 · Kaleal, III · 2016 [cited by examiner]
US 20160274662A1 · Rimon · 2016 [cited by examiner]
US 20170278306A1 · Rico · 2017 [cited by examiner]
US 20200356590A1 · Clarke · 2020 [cited by examiner]
US 20210233317A1 · Son · 2021 [cited by examiner]
US 20220379170A1 · Menaker · 2022 [cited by examiner]
US 20230130555A1 · Wang · 2023 [cited by examiner]
US 20230260182A1 · Saito · 2023 [cited by examiner]
US 20230386522A1 · Buzinover · 2023 [cited by examiner]
US 20240104789A1 · Ghosh · 2024 [cited by examiner]
US 20240135514A1 · Pakhomov · 2024 [cited by examiner]
Uriel Singer et al: “Make-A-Video: Text-to-Video Generation without Text-Video Data”, arxiv.org, Cornell University Library,201O Linlibrary Cornell University Ithaca, NY 14853, Sep. 29, 2022 (Sep. 29, 2022) (Year: 2022). [cited by examiner]
Jonathan Ho et al:“Imagen Video: High Definition Models”, rxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 5, 2022 (Year: 2022). [cited by examiner]
Uriel Singeret al: “Make-A-Video: Text-to-Video Generation Without Text-Video Data”, arxiv.org, Cornell University Library, 2010 Linlibrary Cornell University Ithaca, NY 14853, Sep. 29, 2022 (Sep. 29, 2022) (Year: 2022). [cited by examiner]
ISR WO PCT/US2023/085238, dated Apr. 19, 2024, Total 12 pages. [cited by applicant]
Uriel Singer et al: “Make-A-Video: Text-to-Video Generation without Text-Video Data”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 29, 2022 (Sep. 29, 2022), XP0913299… [cited by applicant]
Kaiduo Zhang et al: HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation,arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, … [cited by applicant]
Jonathan Ho et al: “Imagen Video: High Definition Models”, rxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 5, 2022 (Oct. 5, 2022), XP091334501, * section 1, last paragrap… [cited by applicant]
Starke Sebastian et al: “DeepPhase”, ACM Transactions on Graphics, ACM, NY, US, vol. 41, No. 4, Jul. 22, 2022 (Jul. 22, 2022), pp. 1-13, XP059192210, ISSN: 0730-0301, DOI: 10.1145/3528223.3530178 * p. 136:2, [1]. [cited by applicant]