IP Library › Granted Patent US 12,749,337
Granted Patent B2
US 12,749,337 · App. 18/260,126 · Granted Sep 29, 2026

Face tracking method and electronic device

Inventors: Wenyu Chen (Singapore, CN); Gengdai Liu (Guangzhou, CN)
Assignee: BIGO TECHNOLOGY PTE. LTD.
G06V40/172G06T7/248G06T7/60G06T17/20G06V10/7715G06V20/41G06V20/46G06V40/165G06V40/171G06V40/174G06T2207/10016G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,749,337
App. No.
18/260,126
Granted
Sep 29, 2026
Kind
B2
Abstract

Provided is a face tracking method. The method includes: in a process of tracking a face in a video frame, determining whether an optimization thread is running; in response to the optimization thread running and the video frame being a key frame, updating a second keyframe data set based on the video frame; in response to receiving a clear instruction from the optimization thread, clearing the video frame in the second keyframe data set, and updating the second keyframe data set to a first keyframe data set; in response to the optimization thread not running and the video frame being the key frame, updating the first keyframe data set based on the video frame and the second keyframe data set; and making the optimization thread optimize a facial identity based on the first keyframe data set by invoking the optimization thread upon updating the first keyframe data set.

Claims (105)

1 . A face tracking method, applicable to an optimization thread, the method comprising:

determining a current facial identity vector used by a tracking thread as an initial facial identity vector upon invoking the optimization thread;

acquiring a first keyframe data set, wherein the first keyframe data set is a data set updated by the tracking thread upon performing face tracking, and the first keyframe data set comprises face tracking data;

acquiring optimized face tracking data by optimizing the face tracking data in the first keyframe data set based on the initial facial identity vector;

acquiring an optimized facial identity vector by performing iterative optimization on the initial facial identity vector based on the optimized face tracking data;

upon each iteration, determining, based on the optimized facial identity vector and the initial facial identity vector, whether an iteration stop condition is satisfied;

in response to the iteration stop condition being satisfied, making the tracking thread determine, in response to receiving the optimized facial identity vector, a received facial identity vector as the current facial identity vector by sending the optimized facial identity vector to the tracking thread;

in response to the iteration stop condition being not satisfied, making the tracking thread update, in response to a second keyframe set of a second keyframe data set being a non-empty set, the second keyframe data set to the first keyframe data set upon receiving a clear instruction by sending the clear instruction for clearing video frame in the second keyframe data set to the tracking thread; and

determining the optimized facial identity vector as the initial facial identity vector, and returning to the process of acquiring the first keyframe data set.

2 . The method according to claim 1 , wherein the face tracking data comprises a face keypoint, posture data, and expression data; and acquiring the optimized face tracking data by optimizing the face tracking data in the first keyframe data set based on the initial facial identity vector comprises:

constructing a three-dimensional face model based on the initial facial identity vector and the expression data;

acquiring a face keypoint of the three-dimensional face model; and

acquiring optimized posture data and optimized expression data as the optimized face tracking data by solving optimal posture data and optimal expression data based on the face keypoint of the three-dimensional face model and the face keypoint in the face tracking data.

3 . The method according to claim 2 , wherein acquiring the optimized posture data and the optimized expression data as the optimized face tracking data by solving the optimal posture data and the optimal expression data based on the face keypoint of the three-dimensional face model and the face keypoint in the face tracking data comprises:

solving the optimized face tracking data according to the following formula:

(

P

i

k

,

δ

i

k

)

=

arg

⁢

min

⁡

(

∑

j

∏

P

i

(

C

0

k

-

1

+

C

exp

k

-

1

δ

i

)

j

-

Q

i

j

+

γ

⁢

P

i

,

δ

i

)

wherein k represents a k th iteration;

C

0

k

-

1

 represents a neutral face used in the k th iteration; C_exp{circumflex over ( )}(k−1) represents an expression shape fusion deformer used in the k th iteration; ΠP i (⋅) represents j face keypoints acquired by projecting the three-dimensional face model (C 0 +C exp δ i ) j ; Q i represents a face keypoint of the face tracking data in the first keyframe data set; γ represents a parameter; δ i represents the expression data; P i represents the posture data; and i represents an i th keyframe.

4 . The method according to claim 1 , wherein the face tracking data comprises a face keypoint, posture data, and expression data; and acquiring the optimized facial identity vector by performing the iterative optimization on the initial facial identity vector based on the optimized face tracking data comprises:

calculating a face size of a tracked face based on the face keypoint;

calculating an expression weight of each keyframe based on the expression data of each keyframe; and

acquiring the optimized facial identity vector by performing iterative solving based on the face tracking data, the face size, the expression weight of each keyframe, the current facial identity vector, and the initial facial identity vector.

5 . The method according to claim 4 , wherein calculating the expression weight of each keyframe based on the expression data of each keyframe comprises:

determining minimum expression data from the expression data of all keyframes; and

calculating the expression weight of the keyframe based on a predetermined constant term, the minimum expression data, and the expression data of the keyframe, wherein the expression weight of the keyframe is negatively related to the expression data of the keyframe.

6 . The method according to claim 4 , wherein acquiring the optimized facial identity vector by performing the iterative solving based on the face tracking data, the face size, the expression weight of each keyframe, the current facial identity vector, and the initial facial identity vector comprises:

constructing a three-dimensional face model based on the current facial identity vector and the expression data of each keyframe;

acquiring a plurality of projected face keypoints by projecting the three-dimensional face model to a two-dimensional plane;

calculating a distance sum between the plurality of projected face keypoints and the face keypoints; and

acquiring the optimized facial identity vector by performing the iterative solving based on the expression data, the distance sum, the expression weight of each keyframe, the face size, the current facial identity vector, and the initial facial identity vector.

7 . The method according to claim 1 , wherein determining, based on the optimized facial identity vector and the initial facial identity vector, whether the iteration stop condition is satisfied comprises:

calculating a face change rate based on the optimized facial identity vector and the initial facial identity vector;

determining whether the face change rate is less than a predetermined change rate threshold;

determining that the iteration stop condition is satisfied in response to the face change rate being less than the predetermined change rate threshold; and

determining that the iteration stop condition is not satisfied in response to the face change rate being not less than the predetermined change rate threshold.

8 . The method according to claim 7 , wherein calculating the face change rate based on the optimized facial identity vector and the initial facial identity vector comprises:

acquiring a face size of an average face;

calculating a distance between a face mesh corresponding to the optimized facial identity vector and a face mesh corresponding to the initial facial identity vector; and

calculating a ratio of the distance to the face size of the average face as the face change rate.

9 . The method according to claim 1 , wherein upon sending the optimized facial identity vector to the tracking thread, the method further comprises:

updating a frame vector of each keyframe based on expression data and posture data of optimized keyframes in the first keyframe data set; and

updating a first PCA subspace based on the frame vector of each keyframe.

10 . The method according to claim 1 , wherein prior to determining the current facial identity vector used by the tracking thread as the initial facial identity vector, the method further comprises:

assigning a first PCA subspace to a second PCA subspace.

11 . An electronic device for tracking a face, comprising:

at least one processor; and

a storage apparatus configured to store at least one program;

wherein the at least one program, when run by the at least one processor, causes the at least one processor to perform the face tracking method as defined in claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2023
From: CHEN, WENYU; LIU, GENGDAI
To: BIGO TECHNOLOGY PTE. LTD.
Reel/Frame 064127/0928 →
Priority Claims (1)
CN 202110007729.1 · Jan 5, 2021 · national
Continuity (1)
Related Publication 20240062579A1 · Feb 22, 2024
References Cited (12)
US 20120321134A1 · Shen et al. · 2012 [cited by applicant]
CN 103646391A · 2014 [cited by applicant]
CN 108345821A · 2018 [cited by applicant]
CN 111914613A · 2020 [cited by applicant]
CN 112712044A · 2021 [cited by applicant]
Qi, X., Liu, C., & Schuckers, S. (2018). CNN based key frame extraction for face in video recognition. In 2018 IEEE 4th International Conference on Identity, Security, and Behavior Analysis (ISBA) (pp. 1-8). [cited by examiner]
Wang, Z., Ling, J., Feng, C., Lu, M., & Xu, F. (2022). Emotion-Preserving Blendshape Update With Real-Time Face Tracking. IEEE Transactions on Visualization and Computer Graphics, 28(6), 2364-2375. [cited by examiner]
Cao, C., Hou, Q., & Zhou, K. (2014). Displaced dynamic expression regression for real-time facial tracking and animation. ACM Trans. Graph., 33(4). [cited by examiner]
International Search Report of the International Searching Authority for State Intellectual Property Office of the People's Republic of China in PCT application No. PCT/CN2022/070133 issued on Mar. 30, 2022, which is an… [cited by applicant]
The State Intellectual Property Office of People's Republic of China, First Office Action in Patent Application No. 202110007729.1 issued on May 31, 2023, which is a foreign counterpart application corresponding to this… [cited by applicant]
Cao, Chen et al., “3D Shape Regression for Real-time Facial Animation”, ACM Transactions on Graphics (TOG), vol. 32, No. 4, Article 41, Publication Date: Jul. 2013. [cited by applicant]
Cao, Chen et al., “Displaced Dynamic Expression Regression for Real-time Facial Tracking and Animation”, ACM Transactions on graphics (TOG), vol. 33, No. 4, Article 43, Publication Date: Jul. 2014. [cited by applicant]