IP Library Patent Application 18408100
Patent Application
App. No. 18/408,100

METHOD AND APPARATUS FOR FACE VIDEO COMPRESSION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/408,100
Abstract

A method of encoding a video sequence into a bitstream includes receiving a video sequence; encoding one or more pictures of the video sequence; and generating a bitstream. The encoding includes compressing a reference picture; transforming, based on the reference picture, a plurality of inter pictures associated with the reference picture into facial semantics; and encoding the facial semantics.

Claims (57)

1 . A method of encoding a video sequence into a bitstream, the method comprising:

receiving a video sequence;

encoding one or more pictures of the video sequence; and

generating a bitstream,

wherein the encoding comprises:

compressing a reference picture;

transforming, based on the reference picture, a plurality of inter pictures associated with the reference picture into facial semantics; and

encoding the facial semantics.

2 . The method according to claim 1 , wherein the video sequence comprises a talking face video.

3 . The method according to claim 1 , wherein transforming, based on the reference frame, the plurality of inter pictures associated with the reference picture into facial semantics further comprises:

characterizing the facial semantics using a plurality of three-dimensional (3D) regression parameters.

4 . The method according to claim 3 , wherein the plurality of 3D regression parameters comprises: identity coefficients, albedo coefficients, scene illumination coefficients, expression coefficients, rotation coefficients, translation coefficients, and location coefficient.

5 . The method according to claim 4 , wherein a dimension of the expression coefficients is less than 64.

6 . The method according to claim 5 , wherein the dimension of the expression coefficients is 6, and the expression coefficients represent motion of mouth area.

7 . The method according to claim 1 , wherein the facial semantics are descriptive of one or more of:

a head posture, a face location, a head translation, a mouth motion, or eye blinking.

8 . The method according to claim 7 , wherein an eye-blinking intensity is predicted using facial behaviour analysis.

9 . The method according to claim 1 , wherein encoding the facial semantics further comprises:

obtaining a residual by inter-predicting the facial semantics; and

encoding the residual into the bitstream.

10 . A method of decoding a bitstream to output one or more pictures for a video stream, the method comprising:

receiving a bitstream; and

decoding, using facial semantics in the bitstream, one or more pictures,

wherein the decoding comprises:

reconstructing a reference picture;

decoding the facial semantics of inter pictures;

reconstructing a three-dimensional (3D) mesh based on the facial semantics; and

generating the one or more pictures based on the reconstructed reference frame and facial semantics.

11 . The method according to claim 10 , wherein generating the one or more pictures according to the 3D mesh further comprises:

obtaining a dense motion field and a facial attention map of the 3D mesh; and

reconstructing and compensating the one or more pictures based on the dense motion field and facial attention map.

12 . The method according to claim 10 , wherein reconstructing the 3D mesh based on the reconstructed reference frame and facial semantics further comprises:

reconstructing a 3D face mesh of the reconstructed reference picture or the inter pictures;

obtaining a corresponding 2D face mesh of the 3D face mesh; and

recalibrating a motion of eye regions based on the 2D face mesh.

13 . The method according to claim 11 , wherein obtaining the dense motion field and the facial attention map further comprises:

obtaining a coarse motion field based on motions of each vertex in a 2D face mesh from the reconstructed reference picture and a current inter picture;

obtaining a coarse deformed picture based on the coarse motion field and the reconstructed reference picture; and

obtaining the dense motion field and the facial attention map based on the coarse motion field, the coarse deformed picture, and an eye-blinking motion map.

14 . The method according to claim 11 , further comprising:

obtaining multi-scale spatial features;

obtaining warped facial spatial features by an attention-based feature warping operation on the multi-scale spatial features;

obtaining transformed facial features based on the warped facial spatial features; and

generating the one or more pictures by concatenating the warped facial spatial features and the transformed facial features.

15 . The method according to claim 10 , wherein the video stream comprises a talking face video.

16 . The method according to claim 10 , wherein the facial semantics are descriptive of one or more of:

a head posture, a face location, a head translation, a mouth motion, or eye blinking.

17 . The method according to claim 10 , wherein reconstructing the 3D mesh based on the facial semantics further comprising:

modifying the facial semantics; and

constructing the 3D mesh based on the modified facial semantics.

18 . The method according to claim 10 , wherein before generating the one or more pictures, the method comprises:

applying a virtual character to the one or more frames.

19 . A non-transitory computer readable storage medium storing a bitstream of a video, the bitstream comprising:

an encoded reference picture; and

encoded facial semantics of a plurality of inter frames, wherein the facial semantics are determined based on the reference frame and the plurality of inter frames.

20 . The non-transitory computer readable storage medium according to claim 19 , wherein the facial semantics are descriptive of one or more of:

a head posture, a face location, a head translation, a mouth motion, or eye blinking.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA (CHINA) CO., LTD.
To: ALIBABA INNOVATION PRIVATE LIMITED
Reel/Frame 075529/0567 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA INNOVATION PRIVATE LIMITED
To: SIM IP 5 LLC
Reel/Frame 075529/0713 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2024
From: CHEN, BOLIN; WANG, ZHAO; YE, YAN; WANG, SHIQI
To: ALIBABA (CHINA) CO., LTD.
Reel/Frame 066086/0550 →