IP Library › Granted Patent US 12,573,140
Granted Patent B2
US 12,573,140 · App. 18/444,337 · Granted Mar 10, 2026

Machine learning-based generation of three-dimensional models

Inventors: Chinmay Talegaonkar (La Jolla, CA); Peng Liu (San Diego, CA); Lei Wang (San Diego, CA); Junkang Zhang (San Diego, CA); Ning Bi (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06T17/00G06T3/18G06T7/75G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,140
App. No.
18/444,337
Granted
Mar 10, 2026
Kind
B2
Abstract

Systems and techniques are disclosed for generating a three-dimensional (3D) model. For example, a process can include estimating a plurality of features associated with at least a portion of images; inverse warping the plurality of features into reference pose features having a reference pose; generating filtered reference pose features by selecting features from the reference pose features based on a distance of the selected features from corresponding features from the reference pose; generating modified reference pose features by modifying the filtered reference pose features based on a feature grid associated with a reference model associated with the reference pose; projecting the filtered reference pose features into one or more two dimensional (2D) planes; identifying first features associated with the person from the one or more 2D planes; and generating a 3D model of the person having a pose using the first features, the modified reference pose features, and pose information.

Claims (53)

1 . An apparatus for generating a mutable three-dimensional (3D) model, the apparatus comprising:

at least one memory; and

at least one processor coupled to the at least one memory and configured to:

estimate a plurality of features associated with images of a person;

inverse warp the plurality of features into reference pose features having a reference pose;

generate filtered reference pose features by selecting features from the reference pose features based on a distance of the selected features from corresponding features from the reference pose;

generate modified reference pose features by modifying the filtered reference pose features based on corresponding features of a reference model associated with the reference pose, wherein the reference model is mapped to a feature grid;

project the filtered reference pose features into one or more two dimensional (2D) planes;

identify first features associated with the person from the one or more 2D planes; and

generate a 3D model of the person having a pose using the first features from the one or more 2D planes, the modified reference pose features, and pose information corresponding to the pose.

2 . The apparatus of claim 1 , wherein the at least one processor is configured to:

select one or more features from the filtered reference pose features corresponding to a first reference feature of the reference model; and

interpolate the one or more features to represent the first reference feature in the modified reference pose features.

3 . The apparatus of claim 1 , wherein the at least one processor is configured to:

estimate a position of a plurality of bones in an image using a first model; and

estimate positions of the plurality of features based on the position of the plurality of bones using a second model, wherein the plurality of features corresponds to a surface of the person in the image.

4 . The apparatus of claim 1 , wherein the reference pose features comprise a T-pose.

5 . The apparatus of claim 1 , wherein the feature grid associated with the reference model comprises features anchored on a skinned multi-person linear model in the reference pose.

6 . The apparatus of claim 1 , wherein the at least one processor is configured to:

add a positional encoding to each of the modified reference pose features.

7 . The apparatus of claim 1 , wherein the at least one processor is configured to:

fine tune parameters of a machine learning model based on the first features extracted from one or more of the images.

8 . The apparatus of claim 1 , wherein the reference model includes a machine learning model trained from a dataset including images of people having different physical characteristics.

9 . The apparatus of claim 8 , wherein the machine learning model includes a multilayer perceptron configured to generate the 3D model based on the first features, the modified reference pose features, and the pose information.

10 . The apparatus of claim 8 , wherein the at least one processor is configured to:

warp, use the machine learning model, points in the modified reference pose features and the first features based on the pose information.

11 . The apparatus of claim 1 , wherein the at least one processor is configured to:

query each plane of the one or more 2D planes for information related to a feature in the filtered reference pose features to identify the first features.

12 . The apparatus of claim 1 , wherein the at least one processor is configured to:

combine the filtered reference pose features and the first features into combined features; and

provide the combined features to a machine learning model to generate the 3D model of the person.

13 . A method for generating a mutable three-dimensional (3D) model, comprising:

estimating a plurality of features associated with images of a person;

inverse warping the plurality of features into reference pose features having a reference pose;

generating filtered reference pose features by selecting features from the reference pose features based on a distance of the selected features from corresponding features from the reference pose;

generating modified reference pose features by modifying the filtered reference pose features based on corresponding features of a reference model associated with the reference pose, wherein the reference model is mapped to a feature grid;

projecting the filtered reference pose features into one or more two dimensional (2D) planes;

identifying first features associated with the person from the one or more 2D planes; and

generating a 3D model of the person having a pose using the first features from the one or more 2D planes, the modified reference pose features, and pose information corresponding to the pose.

14 . The method of claim 13 , wherein modifying the filtered reference pose features based on the feature grid comprises:

selecting one or more features from the filtered reference pose features corresponding to a first reference feature of the reference model; and

interpolating the one or more features to represent the first reference feature in the modified reference pose features.

15 . The method of claim 13 , wherein estimating the plurality of features comprises:

estimating a position of a plurality of bones in an image using a first model; and

estimating positions of the plurality of features based on the position of the plurality of bones using a second model, wherein the plurality of features corresponds to a surface of the person in the image.

16 . The method of claim 13 , wherein the feature grid associated with the reference model comprises features anchored on a skinned multi-person linear model in the reference pose.

17 . The method of claim 13 , further comprising:

adding a positional encoding to each of the modified reference pose features.

18 . The method of claim 13 , wherein identifying the first features comprises:

fine tuning parameters of a machine learning model based on the first features extracted from one or more of the images.

19 . The method of claim 13 , wherein the reference model includes a machine learning model trained from a dataset including images of people having different physical characteristics.

20 . The method of claim 19 , further comprising:

warping, using the machine learning model, points in the modified reference pose features and the first features based on the pose information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2024
From: TALEGAONKAR, CHINMAY; LIU, PENG; WANG, LEI; ZHANG, JUNKANG; BI, NING
To: QUALCOMM INCORPORATED
Reel/Frame 066916/0569 →
Continuity (2)
Provisional Application 63591969 · Oct 20, 2023
Related Publication 20250131647A1 · Apr 24, 2025
References Cited (13)
US 20180350105A1 · Taylor · 2018 [cited by examiner]
US 20240047225A1 · Wei · 2024 [cited by examiner]
US 20240303910A1 · Xiao · 2024 [cited by examiner]
US 20240340398A1 · Varekamp · 2024 [cited by examiner]
US 20240406363A1 · Varekamp · 2024 [cited by examiner]
US 20250181056A1 · Urick · 2025 [cited by examiner]
CN 114998514A · 2022 [cited by examiner]
Chen Y, Zheng Z, Li Z, Xu C, Liu Y. Meshavatar: Learning high-quality triangular human avatars from multi-view videos. InEuropean Conference on Computer Vision Sep. 2, 20249 (pp. 250-269). Cham: Springer Nature Switzerl… [cited by examiner]
Guo C, Li J, Kant Y, Sheikh Y, Saito S, Cao C. Vid2avatar-pro: Authentic avatar from videos in the wild via universal prior. InProceedings of the Computer Vision and Pattern Recognition Conference 2025 (pp. 5559-5570). [cited by examiner]
Vuran O, Ho Hl. ReMu: Reconstructing Multi-layer 3D Clothed Human from Image Layers. arXiv preprint arXiv:2508.01381. Aug. 2, 2025. [cited by examiner]
Chen X., et al., “VeRi3D: Generative Vertex-based Radiance Fields for 3D Controllable Human Image Synthesis”, Arxiv.Org, Cornell University, 201 Olin Library Cornell University Ithaca, NY 14853, arXiv:2309.04800v1 [cs.C… [cited by applicant]
International Search Report and Written Opinion—PCT/US2024/047225—ISA/EPO—Jan. 8, 2025. [cited by applicant]
Xu Z., et al., “Relightable and Animatable Neural Avatar from Sparse-View Video”, Arxiv.Org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, arXiv:2308.07903v1 [cs.CV], Aug. 15, 2023, p… [cited by applicant]