IP Library › Granted Patent US 12,561,878
Granted Patent B2
US 12,561,878 · App. 18/440,889 · Granted Feb 24, 2026

Text-driven motion recommendation and neural mesh stylization system and a method for producing human mesh animation using the same

Inventors: Tae Hyun Oh (Pohang-si, KR); You Wang Kim (Pohang-si, KR); Ji Yeon Kim (Pohang-si, KR)
Assignee: POSTECH RESEARCH AND BUSINESS DEVELOPMENT FOUNDATION
G06T13/40G06F40/30G06T15/04G06T17/20G06T19/20G06T2219/2012G06T2219/2024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,878
App. No.
18/440,889
Granted
Feb 24, 2026
Kind
B2
Abstract

The present disclosure provides a text-driven motion recommendation and neural mesh stylization system and a method producing human mesh animation using the same. The system comprises at least one instruction stored in a memory, and a processor that executes the at least one instruction, wherein the at least one instruction, when executed by the processor, causes the processor to find raw action labels matching a query given as a text prompt in a human motion dataset stored in a database, encode the raw action labels and the query for vectorizing the raw action labels and the query, and measure similarity between the raw action labels and the query based on the vectorized vectors.

Claims (49)

1 . A text-driven motion recommendation and neural mesh stylization system comprising:

at least one instruction stored in a memory; and

a processor that executes the at least one instruction,

wherein the at least one instruction, when executed by the processor, causes the processor to:

find raw action labels matching a query given as a text prompt in a human motion dataset stored in a database;

encode the raw action labels and the query for vectorizing the raw action labels and the query;

measure similarity between the raw action labels and the query based on vectorized vectors to obtain content meshes;

obtain style attributes comprising color and displacement from a decoupled neural style field (DNSF) network that takes a template human mesh and learn text-driven style attributes; and

apply the style attributes to the content meshes to obtain a human mesh sequence in motion.

2 . The system of claim 1 , wherein the at least one instruction, when executed by the processor, further causes the processor to:

select a plurality of indices of the raw action labels based on the measured similarity; and

retrieve top-k action labels corresponding the plurality of indices by a top-k filter from encoded motion datasets with the raw action labels.

3 . The system of claim 2 , wherein the at least one instruction, when executed by the processor, further causes the processor to:

vectorize the query and the top-k action labels; and

retrieve a highest-scored raw action label as a final matched result for the input text prompt.

4 . The system of claim 3 , wherein the at least one instruction, when executed by the processor, further causes the processor to:

find a best semantically matched motion sequence from a motion database based on the highest-scored raw action label; and

sample the content meshes in multi-modal context corresponding to the best semantically matched motion sequence.

5 . The system of claim 1 , wherein the at least one instruction, when executed by the processor, further causes the processor to:

map the style attributes from the template human mesh and merge the style attributes mapped from the template human mesh with the content meshes by the DNSF network.

6 . The system of claim 5 , wherein the at least one instruction, when executed by the processor, further causes the processor to:

achieve a same mesh stylization as a basic neural style field while decoupling a style from a content mesh.

7 . The system of claim 1 , wherein the at least one instruction, when executed by the processor, further causes the processor to:

detailize and texturize the human mesh sequence by optimizing the DNSF network in a temporally-consistent and pose-agnostic manner.

8 . The system of claim 7 , wherein the at least one instruction, when executed by the processor, further causes the processor to:

compute a semantic loss between the text prompt and a text obtained by encoding the detailized and texturized human mesh sequence for optimizing the DNSF network.

9 . A method for producing human mesh animation performed by a processor, the method comprising:

finding raw action labels matching a query given as a text prompt in a human motion dataset stored in a database;

encoding the raw action labels and the query for vectorizing the raw action labels and the query;

measuring similarity between the raw action labels and the query based on vectorized vectors to obtain content meshes;

obtaining style attributes comprising color and displacement from a decoupled neural style field (DNSF) network that takes a template human mesh and learn text-driven style attributes; and

applying the style attributes to the content meshes to obtain a human mesh sequence in motion.

10 . The method of claim 9 , further comprising:

selecting a plurality of indices of the raw action labels based on the measured similarity; and

retrieving top-k action labels corresponding the plurality of indices by a top-k filter from encoded motion datasets with the raw action labels.

11 . The method of claim 10 , further comprising:

vectorizing the query and the top-k action labels; and

retrieving a highest-scored raw action label as a final matched result for the text prompt.

12 . The method of claim 11 , further comprising:

finding a best semantically matched motion sequence from a motion database based on the highest-scored raw action label; and

sampling the content meshes in multi-modal context corresponding to the best semantically matched motion sequence.

13 . The method of claim 9 , further comprising:

mapping the style attributes from the template human mesh and merge the style attributes mapped from the template human mesh with the content meshes by the DNSF network.

14 . The method of claim 13 , further comprising:

achieving a same mesh stylization as a basic neural style field while decoupling a style from a content mesh.

15 . The method of claim 9 , further comprising:

detailizing and texturizing the human mesh sequence by optimizing the DNSF network in a temporally-consistent and pose-agnostic manner.

16 . The method of claim 15 , further comprising:

computing a semantic loss between the text prompt and a text obtained by encoding the detailized and texturized human mesh sequence for optimizing the DNSF network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2024
From: OH, TAE HYUN; KIM, YOU WANG; KIM, JI YEON
To: POSTECH RESEARCH AND BUSINESS DEVELOPMENT FOUNDATION
Reel/Frame 066464/0972 →
Priority Claims (2)
KR 10-2023-0018505 · Feb 13, 2023 · national
KR 10-2023-0139071 · Oct 17, 2023 · national
Continuity (1)
Related Publication 20240273798A1 · Aug 15, 2024
References Cited (15)
US 12340480B2 · Zhi · 2025 [cited by examiner]
US 20240193891A1 · Markhasin · 2024 [cited by examiner]
US 20240242452A1 · Zhi · 2024 [cited by examiner]
US 20240273798A1 · Oh · 2024 [cited by examiner]
US 20250157114A1 · Yuan · 2025 [cited by examiner]
US 20250166664A1 · Karim · 2025 [cited by examiner]
EP 4379666A1 · 2024 [cited by examiner]
Hong F, Zhang M, Pan L, Cai Z, Yang L, Liu Z. Avatarclip: Zero-shot text-driven generation and animation of 3d avatars. arXiv preprint arXiv:2205.08535. May 17, 2022. (Year: 2022) (Year: 2022) (Year: 2022) (Year: 2022) … [cited by examiner]
Patashnik O, Wu Z, Shechtman E, Cohen-Or D, Lischinski D. Styleclip: Text-driven manipulation of stylegan imagery. InProceedings of the IEEE/CVF international conference on computer vision 2021 (pp. 2085-2094). [cited by examiner]
Youwang K, Ji-Yeon K, Oh TH. CLIP-Actor: Text-Driven Recommendation and Stylization for Animating Human Meshes. arXiv preprint arXiv:2206.04382. Jun. 9, 2022. [cited by examiner]
Hong, Fangzhou, et al. “AvatarCLIP: Zero-shot text-driven generation and animation of 3D avatars.” arXiv preprint arXiv:2205.08535 (May 17, 2022.). [cited by applicant]
Korean Office Action in Application No. 10-2023-0139071, dated Jul. 28, 2025, 6 pages. [cited by applicant]
Radford, Alec et al, Learning Transferable Visual Models From Natural Language Supervision, arXiv preprint, Feb. 26, 2021, arXiv:2103.00020, arXiv. [cited by applicant]
Song, Kaitao, et al, MPNet: Masked and Permuted Pre-training for Language Understanding, arXiv preprint, Nov. 2, 2020, arXiv:2004.09297, arXiv. [cited by applicant]
Michel, Oscar, et al, Text2Mesh: Text-Driven Neural Stylization for Meshes, arXiv preprint, Dec. 6, 2021., arXiv:2112.03221, arXiv. [cited by applicant]