IP Library › Granted Patent US 12,597,175
Granted Patent B2
US 12,597,175 · App. 18/590,707 · Granted Apr 7, 2026

Avatar creation from natural language description

Inventors: Kun Jin (San Jose, CA); Siva Penke (San Jose, CA)
Assignee: Samsung Electronics Co., Ltd.
G06T11/00G06F40/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,175
App. No.
18/590,707
Granted
Apr 7, 2026
Kind
B2
Abstract

In one embodiment, a method includes accessing, by a computing device, a natural-language input comprising a description of an avatar and generating, from the natural-language input and by a trained avatar-creation model, a set of avatar-attribute feature vectors. The method further includes determining, by a trained avatar-attribute classifier, one or more avatar attributes from the set of avatar-attribute feature vectors; and generating the avatar based on the determined one or more avatar attributes for presentation on a display of a computing device.

Claims (62)

1 . A method comprising:

accessing, by a computing device, a natural-language input comprising a description of an avatar;

determining whether the accessed natural-language input identifies a required attribute for the avatar, and either:

in response to a determination that the accessed natural-language input does not identify the required attribute, then (1) submitting a request for a user to identify the required attribute for the avatar (2) receiving a user input identifying the required attribute and (3) providing the natural language input to a trained avatar-creation model; or

in response to a determination that the accessed natural-language input identifies the required attribute, then providing the natural-language input to a trained avatar-creation model;

generating, from the natural-language input and by the trained avatar-creation model, a set of avatar-attribute feature vectors;

determining, by a trained avatar-attribute classifier, one or more avatar attributes from the set of avatar-attribute feature vectors; and

generating the avatar based on the determined one or more avatar attributes for presentation on a display of a computing device.

2 . The method of claim 1 , further comprising converting the accessed natural-language input from a speech input to a text input.

3 . The method of claim 1 , further comprising converting the accessed natural-language input from a first language to a default language.

4 . The method of claim 1 , wherein the user input identifying the required attribute explicitly identifies the required attribute, and the user input explicitly identifying the required attribute is included in the natural-language input provided to the trained avatar-creation model.

5 . The method of claim 1 , wherein the required attribute comprises a gender of the avatar.

6 . The method of claim 1 , wherein the trained avatar creation model comprises an encoder of a transformer model.

7 . The method of claim 1 , further comprising:

determining, by a scenario detection model of the avatar-creation model, whether the natural-language input identifies one or more scenarios; and

in response to a determination that the natural-language input identifies one or more scenarios, then determining one or more avatar attributes corresponding to the identified one or more scenarios and generating the avatar based at least in part on the one or more avatar attributes corresponding to the identified one or more scenarios.

8 . The method of claim 7 , further comprising weighting the one or more avatar attributes corresponding to the identified one or more scenarios relatively higher than the one or more avatar attributes determined from the set of avatar-attribute feature vectors.

9 . The method of claim 1 , further comprising:

accessing user feedback made in response to display of the avatar; and

combining a description of the user feedback with the accessed natural-language input to generate one or more revised avatar attributes and a revised avatar for presentation on the display of the computing device.

10 . The method of claim 1 , wherein the trained avatar-attribute classifier comprises a classifier trained by:

selecting a set of predefined attributes;

generating, by a text-generation language model, a training dataset comprising one or more sentences, each sentence containing at least one attribute from the set of predefined attributes; and

training the classifier based on the training dataset and the set of predefined attributes.

11 . One or more non-transitory computer readable storage media storing software that is operable when executed by one or more processors to:

access a natural-language input comprising a description of an avatar;

determine whether the accessed natural-language input identifies a required attribute for the avatar, and:

in response to a determination that the accessed natural-language input does not identify the required attribute, then (1) submit a request for a user to identify the required attribute for the avatar (2) receive a user input identifying the required attribute and (3) provide the natural language input to a trained avatar-creation model; or

in response to a determination that the accessed natural-language input identifies the required attribute, then provide the natural-language input to a trained avatar-creation model;

generate from the natural-language input and by a trained avatar-creation model, a set of avatar-attribute feature vectors;

determine, by a trained avatar-attribute classifier, one or more avatar attributes from the set of avatar-attribute feature vectors; and

generate the avatar based on the determined one or more avatar attributes for presentation on a display of a computing device.

12 . The media of claim 11 , wherein the user input identifying the required attribute explicitly identifies the required attribute, and the user input explicitly identifying the required attribute is included in the natural-language input provided to the trained avatar-creation model.

13 . The media of claim 11 , wherein the software is further operable when executed by one or more processors to:

determine, by a scenario detection model of the avatar-creation model, whether the natural-language input identifies one or more scenarios; and

in response to a determination that the natural-language input identifies one or more scenarios, then determine one or more avatar attributes corresponding to the identified one or more scenarios and generate the avatar based at least in part on the one or more avatar attributes corresponding to the identified one or more scenarios.

14 . The media of claim 11 , wherein the software is further operable when executed by one or more processors to:

access user feedback made in response to display of the avatar; and

combine a description of the user feedback with the accessed natural-language input to generate one or more revised avatar attributes and a revised avatar for presentation on the display of the computing device.

15 . The media of claim 11 , wherein the trained avatar-attribute classifier comprises a classifier trained by:

selecting a set of predefined attributes;

generating, by a text-generation language model, a training dataset comprising one or more sentences, each sentence containing at least one attribute from the set of predefined attributes; and

training the classifier based on the training dataset and the set of predefined attributes.

16 . A system comprising one or more non-transitory computer readable storage media storing instructions; and one or more processors coupled to the non-transitory computer readable storage media, the one or more processors operable to execute the instructions to:

access a natural-language input comprising a description of an avatar;

determine whether the accessed natural-language input identifies a required attribute for the avatar, and:

in response to a determination that the accessed natural-language input does not identify the required attribute, then (1) submit a request for a user to identify the required attribute for the avatar (2) receive a user input identifying the required attribute and (3) provide the natural language input to a trained avatar-creation model; or

in response to a determination that the accessed natural-language input identifies the required attribute, then provide the natural-language input to a trained avatar-creation model;

generate from the natural-language input and by a trained avatar-creation model, a set of avatar-attribute feature vectors;

determine, by a trained avatar-attribute classifier, one or more avatar attributes from the set of avatar-attribute feature vectors; and

generate the avatar based at least in part on the determined one or more avatar attributes for presentation on a display of a computing device.

17 . The system of claim 16 , wherein the user input identifying the required attribute explicitly identifies the required attribute, and the user input explicitly identifying the required attribute is included in the natural-language input provided to the trained avatar-creation model.

18 . The system of claim 16 , wherein the one or more processors are further operable to execute the instructions to:

determine, by a scenario detection model of the avatar-creation model, whether the natural-language input identifies one or more scenarios; and

in response to a determination that the natural-language input identifies one or more scenarios, then determine one or more avatar attributes corresponding to the identified one or more scenarios and generate the avatar based on the one or more avatar attributes corresponding to the identified one or more scenarios.

19 . The system of claim 16 , wherein the one or more processors are further operable to execute the instructions to:

access user feedback made in response to display of the avatar; and

combine a description of the user feedback with the accessed natural-language input to generate one or more revised avatar attributes and a revised avatar for presentation on the display of the computing device.

20 . The system of claim 16 , wherein the trained avatar-attribute classifier comprises a classifier trained by:

selecting a set of predefined attributes;

generating, by a text-generation language model, a training dataset comprising one or more sentences, each sentence containing at least one attribute from the set of predefined attributes; and

training the classifier based on the training dataset and the set of predefined attributes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 29, 2024
From: JIN, KUN; PENKE, SIVA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 066601/0562 →
Continuity (2)
Provisional Application 63525810 · Jul 10, 2023
Related Publication 20250022187A1 · Jan 16, 2025
References Cited (25)
US 9665563B2 · Min · 2017 [cited by applicant]
US 9992556B1 · Price · 2018 [cited by examiner]
US 10042536B2 · Goossens · 2018 [cited by applicant]
US 10304439B2 · Okaniwa · 2019 [cited by applicant]
US 11514634B2 · Liao · 2022 [cited by applicant]
US 20070113181A1 · Blattner · 2007 [cited by applicant]
US 20150187112A1 · Rozen · 2015 [cited by applicant]
US 20210125389A1 · Tommy · 2021 [cited by examiner]
US 20210248804A1 · Hussen Abdelaziz · 2021 [cited by applicant]
US 20210327404A1 · Savchenkov · 2021 [cited by applicant]
US 20220084273A1 · Pan · 2022 [cited by applicant]
US 20230177878A1 · Sekar · 2023 [cited by examiner]
US 20240029330A1 · Donnell · 2024 [cited by examiner]
US 20240304177A1 · Wu · 2024 [cited by examiner]
US 20240404225A1 · Ghosh · 2024 [cited by examiner]
CN 113609255A · 2021 [cited by applicant]
CN 115409923A · 2022 [cited by applicant]
CN 112990302B · 2023 [cited by applicant]
CN 115392216B · 2023 [cited by applicant]
Zehranaz Canfes and M. Furkan Atasoy and Alara Dirik and Pinar Yanardag, “Text and Image Guided 3D Avatar Generation and Manipulation”, 2202.06079, https://arxiv.org/abs/2202.06079 (Year: 2022). [cited by examiner]
Zhao et al. “Zero-Shot Text-to-Parameter Translation for Game Character Auto-Creation”, Mar. 2, 2023, https://doi.org/10.48550/ arXiv.2303.01311 (Year: 2023). [cited by examiner]
“10 Best Free AI Avatar Generator Online Websites [Updated]” by Himanshu Tyagi, available at <https://www.codeitbro.com/free-ai-avatar-generator-websites/>, Jan. 16, 2023. [cited by applicant]
“Text and Image Guided 3D Avatar Generation and Manipulation” by Zehranaz Canfes, M. Furkan Atasoy, Alara Dirik, Pinar Yanardag, Available at <https://catlab-team.github.io/latent3D/>, Accessed on Dec. 20, 2023. [cited by applicant]
“10 Best AI Image Generator Review” by Harley Wayne, available at <https://topten.ai/image-generator-review/>, Nov. 7, 2023. [cited by applicant]
Synthesia platform, <https://www.synthesia.io/home?utm_term=synthesia&utm_campaign=Synthesia&utm_source=google&utm_medium=cpc&hsa_acc=5132031546&hsa_cam=20686432010&hsa_grp=156404489922&hsa_ad=677786397346&hsa_src=g&hsa… [cited by applicant]