IP Library › Granted Patent US 12,620,260
Granted Patent B2
US 12,620,260 · App. 18/396,971 · Granted May 5, 2026

Information processing method, computer device, and storage medium

Inventors: Chun Wang (Chongqing, CN); Dingheng Zeng (Chongqing, CN); Xunyi Zhou (Chongqing, CN); Ning Jiang (Chongqing, CN)
Assignee: MASHANG CONSUMER FINANCE CO., LTD.
G06V40/168G06T17/00G06V10/54G06V40/174
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,260
App. No.
18/396,971
Granted
May 5, 2026
Kind
B2
Abstract

An information processing method is provided. The method includes obtaining a target video. A first target image feature corresponding to a face image of each frame is obtained. A target identity coefficient and a target texture coefficient corresponding to the target image feature in the face image of a different frame in the target video are obtained. A first target identity feature is obtained according to the target identity coefficient, and a first target texture feature is obtained according to the target texture coefficient. Once a first target feature is obtained by splicing the target image feature, the first target identity feature and the first target texture feature, a first target expression coefficient is obtained based on the first target feature.

Claims (114)

1 . An information processing method, comprising:

obtaining a target video comprising a plurality of frames, each of the plurality of frames comprising a face image corresponding to a same object;

obtaining a first target image feature corresponding to the face image of each frame;

determining a first target identity coefficient and a first target texture coefficient;

determining a first target identity feature based on the first target identity coefficient;

determining a first target texture feature based on the first target texture coefficient;

splicing the first target image feature, the first target identity feature, and the first target texture feature, and obtaining a first target feature; and

determining a first target expression coefficient based on the first target feature.

2 . The information processing method according to claim 1 , wherein determining the first target identity coefficient and the first target texture coefficient comprises:

obtaining a first identity coefficient and a first texture coefficient of the face image corresponding to the first target image feature in a previous frame of the target video;

obtaining a second identity coefficient and a second texture coefficient corresponding to the first target image feature;

performing a weighted summation on the first identity coefficient and the second identity coefficient, and obtaining the first target identity coefficient corresponding to the first target image feature;

performing the weighted summation on the first texture coefficient and the second texture coefficient, and obtaining the first target texture coefficient corresponding to the first target image feature.

3 . The information processing method according to claim 2 , wherein after determining the first target expression coefficient, the method further comprises:

using the first target identity coefficient to replace the second identity coefficient corresponding to the first target image feature in the face image of a current frame in the target video;

using the first target texture coefficient to replace the second texture coefficient corresponding to the first target image feature in the face image of the current frame in the target video.

4 . The information processing method according to claim 3 , wherein obtaining the target video comprises:

acquiring an initial video;

extracting the face image of each frame in the initial video;

determining the same object by analyzing the face image of each frame, and determining one or more video segments from the initial video, each of the one or more video segments comprising at least two frames and the same object comprised in each of the at least two frames;

determining one of the one or more video segments with a number of frames greater than a preset threshold as the target video.

5 . The information processing method according to claim 4 , wherein determining one of the one or more video segments with the number of frames greater than the preset threshold as the target video comprises:

determining the one of the one or more video segments with the number of frames greater than the preset threshold as a first target video segment;

obtaining a second target video segment by performing a style transformation on the first target video segment; and

determining each of the first target video segment and the second target video segment as the target video.

6 . The information processing method according to claim 5 , wherein after determining the first target identity coefficient and the first target texture coefficient, further comprises:

generating a first target loss function, comprising:

inputting the first target identity coefficient into a second preset backbone model, and outputting a first identity feature;

inputting the first target texture coefficient into a third preset backbone model, and outputting a first texture feature;

splicing the first target image feature, the first identity feature, and the first texture feature, and obtaining a first feature;

inputting the first feature into a preset head network model, and outputting a first predicted expression coefficient;

generating a first predicted face three-dimensional model according to a label identity coefficient, a label texture coefficient, the first predicted expression coefficient, a label posture coefficient, and a label lighting coefficient;

obtaining a first difference between a first face estimated value corresponding to the first predicted face three-dimensional model and an un-occluded area in the face image;

obtaining a second difference between first predicted face three-dimensional key points corresponding to the first predicted face three-dimensional model and face three-dimensional key points in the face image; and

establishing the first target loss function based on the first difference and the second difference.

7 . The information processing method according to claim 6 , further comprising:

performing an optimization on a first network parameters of the second preset backbone model, the third preset backbone model, and the preset head network model according to the first target loss function;

returning to repeatedly generate the first target loss function, iteratively optimizing the first network parameters of the second preset backbone model, the third preset backbone model, and the preset head network model through the first target loss function that is generated, until the first target loss function converges, and obtaining a second target preset backbone model, a third target preset backbone model, and a target preset head network model that have been trained.

8 . A computer device comprising:

a storage device;

at least one processor; and

the storage device storing one or more programs, which when executed by the at least one processor, cause the at least one processor to:

obtain a target video comprising a plurality of frames, each of the plurality of frames comprising a face image corresponding to a same object;

obtain a first target image feature corresponding to the face image of each frame;

determine a first target identity coefficient and a first target texture coefficient;

determine a first target identity feature based on the first target identity coefficient;

determine a first target texture feature based on the first target texture coefficient;

splice the first target image feature, the first target identity feature, and the first target texture feature, and obtain a first target feature; and

determine a first target expression coefficient based on the first target feature.

9 . The computer device according to claim 8 , wherein the at least one processor determines the first target identity coefficient and the first target texture coefficient by:

obtaining a first identity coefficient and a first texture coefficient of the face image corresponding to the first target image feature in a previous frame of the target video;

obtaining a second identity coefficient and a second texture coefficient corresponding to the first target image feature;

performing a weighted summation on the first identity coefficient and the second identity coefficient, and obtaining the first target identity coefficient corresponding to the first target image feature;

performing the weighted summation on the first texture coefficient and the second texture coefficient, and obtaining the first target texture coefficient corresponding to the first target image feature.

10 . The computer device according to claim 9 , wherein after determining the first target expression coefficient, the at least one processor is further caused to:

use the first target identity coefficient to replace the second identity coefficient corresponding to the first target image feature in the face image of a current frame in the target video;

use the first target texture coefficient to replace the second texture coefficient corresponding to the first target image feature in the face image of the current frame in the target video.

11 . The computer device according to claim 10 , wherein the at least one processor obtains the target video by:

acquiring an initial video;

extracting the face image of each frame in the initial video;

determining the same object by analyzing the face image of each frame, and determining one or more video segments from the initial video, each of the one or more video segments comprising at least two frames and the same object comprised in each of the at least two frames;

determining one of the one or more video segments with a number of frames greater than a preset threshold as the target video.

12 . The computer device according to claim 11 , wherein the at least one processor determines one of the one or more video segments with the number of frames greater than the preset threshold as the target video by:

determining the one of the one or more video segments with the number of frames greater than the preset threshold as a first target video segment;

obtaining a second target video segment by performing a style transformation on the first target video segment; and

determining each of the first target video segment and the second target video segment as the target video.

13 . The computer device according to claim 12 , wherein after determining the first target identity coefficient and the first target texture coefficient, the at least one processor is further caused to:

generate a first target loss function, comprising:

input the first target identity coefficient into a second preset backbone model, and output a first identity feature;

input the first target texture coefficient into a third preset backbone model, and output a first texture feature;

splice the first target image feature, the first identity feature, and the first texture feature, and obtain a first feature;

input the first feature into a preset head network model, and output a first predicted expression coefficient;

generate a first predicted face three-dimensional model according to a label identity coefficient, a label texture coefficient, the first predicted expression coefficient, a label posture coefficient, and a label lighting coefficient;

obtain a first difference between a first face estimated value corresponding to the first predicted face three-dimensional model and an un-occluded area in the face image;

obtain a second difference between first predicted face three-dimensional key points corresponding to the first predicted face three-dimensional model and face three-dimensional key points in the face image; and

establish the first target loss function based on the first difference and the second difference.

14 . The computer device according to claim 13 , the at least one processor is further caused to:

perform an optimization on a first network parameters of the second preset backbone model, the third preset backbone model, and the preset head network model according to the first target loss function;

return to repeatedly generate the first target loss function, iteratively optimizing the first network parameters of the second preset backbone model, the third preset backbone model, and the preset head network model through the first target loss function that is generated, until the first target loss function converges, and obtaining a second target preset backbone model, a third target preset backbone model, and a target preset head network model that have been trained.

15 . A non-transitory storage medium having instructions stored thereon, when the instructions are executed by a processor of a computer device, the processor is caused to perform an information processing method, wherein the method comprises:

obtaining a target video comprising a plurality of frames, each of the plurality of frames comprising a face image corresponding to a same object;

obtaining a first target image feature corresponding to the face image of each frame;

determining a first target identity coefficient and a first target texture coefficient;

determining a first target identity feature based on the first target identity coefficient;

determining a first target texture feature based on the first target texture coefficient;

splicing the first target image feature, the first target identity feature, and the first target texture feature, and obtaining a first target feature; and

determining a first target expression coefficient based on the first target feature.

16 . The non-transitory storage medium according to claim 15 , wherein determining the first target identity coefficient and the first target texture coefficient comprises:

obtaining a first identity coefficient and a first texture coefficient of the face image corresponding to the first target image feature in a previous frame of the target video;

obtaining a second identity coefficient and a second texture coefficient corresponding to the first target image feature;

performing a weighted summation on the first identity coefficient and the second identity coefficient, and obtaining the first target identity coefficient corresponding to the first target image feature;

performing the weighted summation on the first texture coefficient and the second texture coefficient, and obtaining the first target texture coefficient corresponding to the first target image feature.

17 . The non-transitory storage medium according to claim 16 , wherein after determining the first target expression coefficient, the method further comprises:

using the first target identity coefficient to replace the second identity coefficient corresponding to the first target image feature in the face image of a current frame in the target video;

using the first target texture coefficient to replace the second texture coefficient corresponding to the first target image feature in the face image of the current frame in the target video.

18 . The non-transitory storage medium according to claim 17 , wherein obtaining the target video comprises:

acquiring an initial video;

extracting the face image of each frame in the initial video;

determining the same object by analyzing the face image of each frame, and determining one or more video segments from the initial video, each of the one or more video segments comprising at least two frames and the same object comprised in each of the at least two frames;

determining one of the one or more video segments with a number of frames greater than a preset threshold as the target video.

19 . The non-transitory storage medium according to claim 18 , wherein determining one of the one or more video segments with the number of frames greater than the preset threshold as the target video comprises:

determining the one of the one or more video segments with the number of frames greater than the preset threshold as a first target video segment;

obtaining a second target video segment by performing a style transformation on the first target video segment; and

determining each of the first target video segment and the second target video segment as the target video.

20 . The non-transitory storage medium according to claim 19 , wherein after determining the first target identity coefficient and the first target texture coefficient, the method further comprises:

generating a first target loss function, comprising:

inputting the first target identity coefficient into a second preset backbone model, and outputting a first identity feature;

inputting the first target texture coefficient into a third preset backbone model, and outputting a first texture feature;

splicing the first target image feature, the first identity feature, and the first texture feature, and obtaining a first feature;

inputting the first feature into a preset head network model, and outputting a first predicted expression coefficient;

generating a first predicted face three-dimensional model according to a label identity coefficient, a label texture coefficient, the first predicted expression coefficient, a label posture coefficient, and a label lighting coefficient;

obtaining a first difference between a first face estimated value corresponding to the first predicted face three-dimensional model and an un-occluded area in the face image;

obtaining a second difference between first predicted face three-dimensional key points corresponding to the first predicted face three-dimensional model and face three-dimensional key points in the face image; and

establishing the first target loss function based on the first difference and the second difference.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2023
From: WANG, CHUN; ZENG, DINGHENG; ZHOU, XUNYI; JIANG, NING
To: MASHANG CONSUMER FINANCE CO., LTD.
Reel/Frame 065960/0563 →
Priority Claims (1)
CN 202210369409.5 · Apr 8, 2022 · national
Continuity (2)
Continuation In Part PCTCN2022144220 · Dec 30, 2022
Related Publication 20240135747A1 · Apr 25, 2024
References Cited (45)
US 8249310B2 · Okubo · 2012 [cited by examiner]
US 10783352B2 · Huang · 2020 [cited by examiner]
US 11222466B1 · Naruniec · 2022 [cited by examiner]
US 11276231B2 · Chandran · 2022 [cited by examiner]
US 11321960B2 · Zhou · 2022 [cited by examiner]
US 11380050B2 · Zhe · 2022 [cited by examiner]
US 11645798B1 · Demyanov · 2023 [cited by examiner]
US 11686105B2 · Beyreuther · 2023 [cited by examiner]
US 11869150B1 · Mason · 2024 [cited by examiner]
US 11941753B2 · Zhou · 2024 [cited by examiner]
US 12243349B2 · Bradley · 2025 [cited by examiner]
US 12266042B2 · Kimura · 2025 [cited by examiner]
US 20070189627A1 · Cohen · 2007 [cited by examiner]
US 20180068178A1 · Theobalt · 2018 [cited by examiner]
US 20190138794A1 · Huang · 2019 [cited by examiner]
US 20190147642A1 · Cole · 2019 [cited by examiner]
US 20210192192A1 · Li · 2021 [cited by examiner]
US 20210350508A1 · Li · 2021 [cited by examiner]
US 20210357625A1 · Song · 2021 [cited by examiner]
US 20210386383A1 · McDuff · 2021 [cited by examiner]
US 20210406568A1 · Liberman · 2021 [cited by examiner]
US 20220129689A1 · Kim · 2022 [cited by examiner]
US 20220301348A1 · Bradley · 2022 [cited by examiner]
US 20230081982A1 · He · 2023 [cited by examiner]
US 20230326248A1 · He · 2023 [cited by examiner]
US 20240037852A1 · Zhang · 2024 [cited by examiner]
CN 111523413A · 2020 [cited by examiner]
CN 112036356A · 2020 [cited by examiner]
CN 112884881 · 2021 [cited by applicant]
CN 112884881A · 2021 [cited by examiner]
CN 113838173 · 2021 [cited by applicant]
CN 113838173A · 2021 [cited by examiner]
CN 113887529 · 2022 [cited by applicant]
CN 114241558 · 2022 [cited by applicant]
CN 114241558A · 2022 [cited by examiner]
CN 114782864 · 2022 [cited by applicant]
CN 114783022 · 2022 [cited by applicant]
CN 114821404 · 2022 [cited by applicant]
CN 114898244 · 2022 [cited by applicant]
Monocular 3D Facial Expression Features for Continuous Affect Recognition, Ercheng Pei et al., IEEE, 2021, pp. 3540-3550 ( Year: 2021). [cited by examiner]
High Quality Facial Surface and Texture Synthesis via Generative Adversarial Networks, Ron Slossberg et al., ECCV, 2018, pp. 1-18 (Year: 2018). [cited by examiner]
Face Texture Generation and Identity-Preserving Rectification, Stefan Hormann et al., IEEE, 2021, pp. 2448-2452 (Year: 2021). [cited by examiner]
Full Face-and-Head 3D Model With Photorealistic Texture, Yangyu Fan et al., IEEE, 2020, pp. 210709-210721 (Year: 2020). [cited by examiner]
Simultaneous Facial Feature Tracking and Facial Expression Recognition, Yongqiang Li et al., IEEE, 2013, pp. 2559-2573 (Year : 2013). [cited by examiner]
Ercheng Pei et al: “Monocular 3D Facial Expression Features for Continuous Affect Recognition”, IEEE Transactions on Multimedia, IEEE, USA, vol. 23, Sep. 25, 2020, pp. 3540-3550, XP011883722, ISSN: 1520-9210,DOI: 10.110… [cited by applicant]