IP Library Granted Patent US 11,640,687
Granted Patent B2
US 11,640,687 · App. 17/219,698 · Granted May 2, 2023

Volumetric capture and mesh-tracking based machine learning 4D face/body deformation training

Inventors: Kenji Tashiro (San Jose, CA); Qing Zhang (San Jose, CA)
Assignee: Sony Group Corporation
G06T13/40G06N20/00G06T17/20G06V40/103G06V40/23
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,640,687
App. No.
17/219,698
Filed
Mar 31, 2021
Granted
May 2, 2023
Kind
B2
Examiner
GUO, XILIN
Art Unit
2616
USPC
345/474
Abstract

Mesh-tracking based dynamic 4D modeling for machine learning deformation training includes: using a volumetric capture system for high-quality 4D scanning, using mesh-tracking to establish temporal correspondences across a 4D scanned human face and full-body mesh sequence, using mesh registration to establish spatial correspondences between a 4D scanned human face and full-body mesh and a 3D CG physical simulator, and training surface deformation as a delta from the physical simulator using machine learning. The deformation for natural animation is able to be predicted and synthesized using the standard MoCAP animation workflow. Machine learning based deformation synthesis and animation using standard MoCAP animation workflow includes using single-view or multi-view 2D videos of MoCAP actors as input, solving 3D model parameters (3D solving) for animation (deformation not included), and given 3D model parameters solved by 3D solving, predicting 4D surface deformation from ML training.

Claims (41)

1. A method programmed in a non-transitory memory of a device comprising:

using a volumetric capture system for high-quality 4D scanning, wherein the volumetric capture system is configured for capturing high quality photos and video simultaneously;

using mesh-tracking to establish temporal correspondences across a 4D scanned human face and full-body mesh sequence;

using mesh registration to establish spatial correspondences between the 4D scanned human face and full-body mesh sequence and a 3D computer graphics physical simulator, wherein the 3D computer graphics physical simulator receives a surface representation and a hierarchical set of interconnected parts from a rigging implementation; and

training surface deformation as a delta from the 3D computer graphics physical simulator using machine learning including combining detail-rich implicit functions and parametric representations to construct 3D models.

2. The method of claim 1 further comprising acquiring multiple, separate 3D scans.

3. The method of claim 1 further comprising predicting and synthesizing the deformation for natural animation using a standard motion capture animation.

4. The method of claim 1 further comprising:

using single-view or multi-view 2D videos of motion capture actors as input;

solving 3D model parameters for animation; and

based on the 3D model parameters solved by 3D solving, predicting 4D surface deformation from machine learning training.

5. An apparatus comprising:

a non-transitory memory for storing an application, the application for:

using a volumetric capture system for high-quality 4D scanning, wherein the volumetric capture system is configured for capturing high quality photos and video simultaneously;

using mesh-tracking to establish temporal correspondences across a 4D scanned human face and full-body mesh sequence;

using mesh registration to establish spatial correspondences between the 4D scanned human face and full-body mesh sequence and a 3D computer graphics physical simulator, wherein the 3D computer graphics physical simulator receives a surface representation and a hierarchical set of interconnected parts from a rigging implementation; and

training surface deformation as a delta from the 3D computer graphics physical simulator using machine learning including combining detail-rich implicit functions and parametric representations to construct 3D models; and

a processor coupled to the memory, the processor configured for processing the application.

6. The apparatus of claim 5 wherein the application is further configured for acquiring multiple, separate 3D scans.

7. The apparatus of claim 5 wherein the application is further configured for predicting and synthesizing the deformation for natural animation using a standard motion capture animation.

8. The apparatus of claim 5 wherein the application is further configured for:

using single-view or multi-view 2D videos of motion capture actors as input;

solving 3D model parameters for animation; and

based on the 3D model parameters solved by 3D solving, predicting 4D surface deformation from machine learning training.

9. A system comprising:

a volumetric capture system for high-quality 4D scanning, wherein the volumetric capture system is configured for capturing high quality photos and video simultaneously; and

a computing device configured for:

using mesh-tracking to establish temporal correspondences across a 4D scanned human face and full-body mesh sequence;

using mesh registration to establish spatial correspondences between the 4D scanned human face and full-body mesh sequence and a 3D computer graphics physical simulator, wherein the 3D computer graphics physical simulator receives a surface representation and a hierarchical set of interconnected parts from a rigging implementation; and

training surface deformation as a delta from the 3D computer graphics physical simulator using machine learning including combining detail-rich implicit functions and parametric representations to construct 3D models.

10. The system of claim 9 wherein the computing device is further configured for predicting and synthesizing the deformation for natural animation using a standard motion capture animation.

11. The system of claim 9 wherein the computing device is further configured for:

using single-view or multi-view 2D videos of motion capture actors as input;

solving 3D model parameters for animation; and

based on the 3D model parameters solved by 3D solving, predicting 4D surface deformation from machine learning training.

12. A method programmed in a non-transitory memory of a device comprising:

using single-view or multi-view 2D videos of motion capture actors as input;

solving 3D model parameters for animation; and

based on the 3D model parameters solved by 3D solving, predicting 4D surface deformation from machine learning training using mesh-tracking to establish temporal correspondences across a 4D scanned human face and full-body mesh sequence;

using mesh registration to establish spatial correspondences between the 4D scanned human face and full-body mesh sequence and a 3D computer graphics physical simulator, wherein the 3D computer graphics physical simulator receives a surface representation and a hierarchical set of interconnected parts from a rigging implementation; and

training surface deformation as a delta from the 3D computer graphics physical simulator using the machine learning including combining detail-rich implicit functions and parametric representations to construct 3D models.

Assignments (2)
CHANGE OF NAME Recorded May 16, 2023
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 063664/0744 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2021
From: TASHIRO, KENJI; ZHANG, QING
To: SONY CORPORATION
Reel/Frame 055789/0779 →
Continuity (2)
Provisional Application 63003097 · Mar 31, 2020
Related Publication 20210304478A1 · Sep 30, 2021