IP Library Granted Patent US 11,354,860
Granted Patent B1
US 11,354,860 · App. 17/158,909 · Granted Jun 7, 2022

Object reconstruction using media data

Inventors: Yan Deng (La Jolla, CA); Michel Adib Sarkis (San Diego, CA); Ning Bi (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06T17/205G06T13/40G06T19/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,354,860
App. No.
17/158,909
Granted
Jun 7, 2022
Kind
B1
Abstract

Systems and techniques are provided for performing video-based activity recognition. For example, a process can include generating a three-dimensional (3D) model of a first portion of an object based on one or more frames depicting the object. The process can also include generating a mask for the one or more frames, the mask including an indication of one or more of regions of the object. The process can further include generating a 3D base model based on the 3D model of the first portion of the object and the mask, the 3D base model representing the first portion of the object and a second portion of the object. The process can include generating, based on the mask and the 3D base model, a 3D model of the second portion of the object.

Claims (60)

1. An apparatus for generating one or more models, comprising:

a memory; and

one or more processors coupled to the memory, the one or more processors configured to:

generate a three-dimensional (3D) model of a first portion of an object based on one or more frames depicting the object;

generate at least one mask for the one or more frames, the at least one mask including an indication of one or more of regions of the object;

generate a 3D base model based on a union between a representation of the 3D model of the first portion of the object and the at least one mask, the 3D base model representing the first portion of the object and a second portion of the object; and

generate, based on the at least one mask and the 3D base model, a 3D model of the second portion of the object.

2. The apparatus of claim 1 , wherein the 3D model of the second portion corresponds to an item that is part of the object.

3. The apparatus of claim 1 , wherein the object is a person, the first portion of the object corresponds to a head of the person, and the second portion of the object corresponds to hair on the head of the person.

4. The apparatus of claim 1 , wherein the 3D model of the second portion corresponds to an item that is at least one of separable from the object and movable relative to the object.

5. The apparatus of claim 1 , wherein the object is a person, the first portion of the object corresponds to a body region of the person, and the second portion of the object corresponds to an accessory or clothing worn by the person.

6. The apparatus of claim 1 , wherein the 3D model of the second portion of the object abuts at least a portion of the 3D model of the first portion of the object.

7. The apparatus of claim 1 , wherein the one or more processors are configured to:

segment each frame of the one or more frames into one or more regions; and

generate a mask for each frame of the one or more frames, wherein the mask for each frame includes an indication of the one or more regions.

8. The apparatus of claim 1 , wherein the representation of the 3D model of the first portion of the object is a rasterization of the 3D model of the first portion of the object for a frame, wherein the at least one mask includes a first mask for a first region of the object and a second mask for a second region of the object, and wherein the one or more processors are configured to:

determine the union between the rasterization of the 3D model of the first portion of the object for the frame, the first mask for the first region of the object, and the second mask for the second region of the object.

9. The apparatus of claim 8 , wherein the first region is a face region of the object and the second region is a hair region of the object.

10. The apparatus of claim 1 , wherein the one or more processors are configured to:

project each vertex of an initial 3D model to a mask associated with a frame of the one or more frames based on pose information associated with the frame;

determine whether each vertex of the 3D model of the first portion is located within a first region of the mask associated with the frame; and

extract the 3D base model based on vertices of the 3D model of the first portion being within the first region of the mask associated with the frame.

11. The apparatus of claim 10 , wherein the object is a person and the first region corresponds to a facial region of the person and a hair region of the person.

12. The apparatus of claim 10 , wherein the object is a person and the first region corresponds to a body region of the person and a dress region worn by the person.

13. The apparatus of claim 10 , wherein the one or more processors are configured to:

remove one or more vertices from the 3D base model based on a probability that each vertex of the one or more vertices is within a region of the one or more regions of a frame from the one or more frames.

14. The apparatus of claim 1 , wherein the one or more processors are configured to:

generate an animation in an application using the 3D model of the first portion and the 3D model of the second portion, wherein the object comprises a person, the 3D model of the first portion corresponds to a head of the person, and the 3D model of the second portion corresponds to hair of the person.

15. The apparatus of claim 14 , wherein the application includes functions to transmit and receive at least one of audio and text.

16. The apparatus of claim 14 , wherein the 3D model of the first portion and the 3D model of the second portion depict a user of the application.

17. The apparatus of claim 1 , wherein the one or more processors are configured to:

receive input corresponding to selection of at least one graphical control for modifying the 3D model of the second portion; and

modify the 3D model of the second portion based on the received input.

18. A method of generating one or more models, comprising:

generating a three-dimensional (3D) model of a first portion of an object based on one or more frames depicting the object;

generating at least one mask for the one or more frames, the at least one mask including an indication of one or more regions of the object;

generating a 3D base model based on a union between a representation of the 3D model of the first portion of the object and the at least one mask, the 3D base model representing the first portion of the object and a second portion of the object; and

generating, based on the at least one mask and the 3D base model, a 3D model of the second portion of the object.

19. The method of claim 18 , wherein the 3D model of the second portion corresponds to an item that is part of the object.

20. The method of claim 18 , wherein the object is a person, the first portion of the object corresponds to a head of the person, and the second portion of the object corresponds to hair on the head of the person.

21. The method of claim 18 , wherein the 3D model of the second portion corresponds to an item that is at least one of separable from the object and movable relative to the object.

22. The method of claim 18 , wherein the object is a person, the first portion of the object corresponds to a body region of the person, and the second portion of the object corresponds to an accessory or clothing worn by the person.

23. The method of claim 18 , wherein the 3D model of the second portion of the object abuts at least a portion of the 3D model of the first portion of the object.

24. The method of claim 18 , wherein generating the at least one mask for the one or more frames comprises:

segmenting each frame of the one or more frames into one or more regions; and

generating a mask for each frame of the one or more frames, wherein the mask for each frame includes an indication of the one or more regions.

25. The method of claim 18 , wherein the representation of the 3D model of the first portion of the object is a rasterization of the 3D model of the first portion of the object for a frame, wherein the at least one mask includes a first mask for a first region of the object and a second mask for a second region of the object, the method further comprising:

determining the union between the rasterization of the 3D model of the first portion of the object for the frame, the first mask for the first region of the object, and the second mask for the second region of the object.

26. The method of claim 18 , wherein generating the 3D base model comprises:

projecting each vertex of an initial 3D model to a mask associated with a frame of the one or more frames based on pose information associated with the frame;

determining whether each vertex of the 3D model of the first portion is located within a first region of the mask associated with the frame; and

extracting the 3D base model based on vertices of the 3D model of the first portion being within the first region of the mask associated with the frame.

27. The method of claim 26 , further comprising:

removing one or more vertices from the 3D base model based on a probability that each vertex of the one or more vertices is within a region of the one or more regions of a frame from the one or more frames.

28. The method of claim 18 , further comprising:

generating an animation in an application using the 3D model of the first portion and the 3D model of the second portion, wherein the object comprises a person, the 3D model of the first portion corresponds to a head of the person, and the 3D model of the second portion corresponds to hair of the person.

29. The method of claim 28 , wherein the 3D model of the first portion and the 3D model of the second portion depict a user of the application.

30. The method of claim 18 , further comprising:

receiving input corresponding to selection of at least one graphical control for modifying the 3D model of the second portion; and

modifying the 3D model of the second portion based on the received input.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE SPELLING OF INVENTOR NUMBER 2'S MIDDLE NAME TO ADIB PREVIOUSLY RECORDED AT REEL: 055658 FRAME: 0051. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 23, 2022
From: DENG, YAN; SARKIS, MICHEL ADIB; BI, NING
To: QUALCOMM INCORPORATED
Reel/Frame 059485/0532 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 18, 2021
From: DENG, YAN; SARKIS, MICHEL ABID; BI, NING
To: QUALCOMM INCORPORATED
Reel/Frame 055658/0051 →
Cited By (1)
US 12,406,375