IP Library › Granted Patent US 12,744,958
Granted Patent B2
US 12,744,958 · App. 19/099,533 · Granted Sep 22, 2026

Video determination method and apparatus, electronic device and storage medium

Inventors: Jianwei Li (Los Angeles, CA); Xiao Yang (Los Angeles, CA)
Assignee: Lemon Inc.
H04N21/43072G06V40/168H04N21/439H04N21/44008H04N21/8547
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,958
App. No.
19/099,533
Granted
Sep 22, 2026
Kind
B2
Abstract

Embodiments of the present disclosure provide a video determination method and apparatus, an electronic device, and a storage medium. The method includes: acquiring, in response to an effect trigger operation, a target facial image including a target object; determining a target audio, and determining a key video frame sequence corresponding to the target audio; determining, based on the key video frame sequence and the target facial image, a target facial feature in the target facial image that is presented when the target audio is played; and determining a target effect audio and video based on the target facial feature and the target audio.

Claims (82)

1 . A video determination method applied to a client, the method comprising:

acquiring, in response to an effect trigger operation, a first facial image comprising a first object;

determining a first audio, and determining a key video frame sequence corresponding to the first audio;

determining, based on the key video frame sequence and the first facial image, a first facial feature in the first facial image that is presented when the first audio is played; and

determining a first effect audio and video based on the first facial feature and the first audio,

wherein the key video frame sequence comprises at least one video frame, and a facial feature of a user in the video frame is inconsistent with a second facial feature.

2 . The method according to claim 1 , wherein the effect trigger operation comprises at least one of the following:

triggering an effect prop;

a frame-in picture comprising the first object;

triggering an effect wake-up word using audio information; and

a current body movement being consistent with a first effect movement.

3 . The method according to claim 1 , wherein the determining a first audio comprises:

displaying at least one second audio, and determining the first audio based on a trigger operation on the at least one second audio within first duration; or

receiving an uploaded third audio as the first audio.

4 . The method according to claim 1 , further comprising:

determining the key video frame sequence corresponding to the first audio based on a pre-selected first language type corresponding to the first audio.

5 . The method according to claim 1 , wherein the determining a key video frame sequence corresponding to the first audio comprises:

retrieving the key video frame sequence corresponding to the first audio from a pre-determined key video frame sequence library, wherein the first audio is determined from at least one second audio that is displayed in a display interface, and the key video frame sequence library comprises a corresponding key video frame sequence obtained after the at least one second audio is processed; or

processing the first audio to obtain the key video frame sequence corresponding to the first audio.

6 . The method according to claim 5 , wherein the facial feature comprises a mouth shape feature, and the determining the key video frame sequence comprises:

obtaining a second facial image;

obtaining, based on the second audio or the first audio, and the second facial image, a first audio and video in which a mouth shape feature in the second facial image is consistent with a mouth shape feature presented when a third audio or the first audio is played; and

using, as the key video frame sequence, a plurality of audio and video frames of the first audio and video in which mouth shape features are inconsistent with a first mouth shape feature.

7 . The method according to claim 6 , wherein the facial feature further comprises a facial organ feature, and the determining the key video frame sequence comprises:

processing, based on a pre-trained facial driving model, the first audio and video and a third facial image, so as to obtain a second audio and video in which a facial organ feature in the third facial image changes;

sequentially determining a facial organ feature in each audio and video frame of the second audio and video;

using, as a key video frame, an audio and video frame in which the facial organ feature is inconsistent with a first facial organ feature; and

determining, based on timestamps of a plurality of key video frames, the key video frame sequence corresponding to the first audio and video.

8 . The method according to claim 1 , wherein the determining, based on the key video frame sequence and the first facial image, a first facial feature in the first facial image that is presented when the first audio is played comprises:

determining, based on reference feature point data of each key video frame in the key video frame sequence and basic feature point data of the first facial image, first feature point data corresponding to the key video frame; and

determining, based on the first feature point data, the first facial image, and the corresponding basic feature point data, the first facial feature in the first facial image that is presented when the first audio is played,

wherein the reference feature point data corresponds to facial organ feature point data or mouth shape feature point data.

9 . The method according to claim 8 , wherein the determining, based on reference feature point data of each key video frame in the key video frame sequence and basic feature point data of the first facial image, first feature point data corresponding to the key video frame comprises:

determining, for each key video frame, the reference feature point data of the current key video frame and the basic feature point data of the first facial image, and determining difference feature data from the current key video frame; and

determining, based on the difference feature data from each key video frame and the basic feature point data, the first feature point data corresponding to the first facial image in each key video frame.

10 . The method according to claim 8 , wherein the determining, based on the first feature point data, the first facial image, and the corresponding basic feature point data, the first facial feature in the first facial image that is presented when the first audio is played comprises:

inputting the first feature point data, the first facial image, and the corresponding basic feature point data into a pre-trained effect audio and video generation model to obtain the first facial feature of the first object.

11 . The method according to claim 10 , further comprising:

determining sound-picture synchronized videos that correspond to at least one fourth audio in different language types;

determining, based on the sound-picture synchronized video and fourth facial images of different second objects, a first key video frame sequence of the different second objects in the corresponding sound-picture synchronized video; and

obtaining a fifth facial image of a third object, the first key video frame sequence, and facial feature data that corresponds to the fifth facial image, to construct a training sample for training the effect audio and video generation model.

12 . An electronic device, comprising:

one or more processors; and

a storage apparatus configured to store one or more programs,

wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to:

acquire, in response to an effect trigger operation, a first facial image comprising a first object;

determine a first audio, and determining a key video frame sequence corresponding to the first audio;

determine, based on the key video frame sequence and the first facial image, a first facial feature in the first facial image that is presented when the first audio is played; and

determine a first effect audio and video based on the first facial feature and the first audio;

wherein the key video frame sequence comprises at least one video frame, and a facial feature of a user in the video frame is inconsistent with a second facial feature.

13 . A non-transitory storage medium comprising computer-executable instructions that, when executed by a computer processor, cause the computer processor to:

acquire, in response to an effect trigger operation, a first facial image comprising a first object;

determine a first audio, and determining a key video frame sequence corresponding to the first audio;

determine, based on the key video frame sequence and the first facial image, a first facial feature in the first facial image that is presented when the first audio is played; and

determine a first effect audio and video based on the first facial feature and the first audio;

wherein the key video frame sequence comprises at least one video frame, and a facial feature of a user in the video frame is inconsistent with a second facial feature.

14 . The device according to claim 12 , wherein the effect trigger operation comprises at least one of the following:

triggering an effect prop;

a frame-in picture comprising the first object;

triggering an effect wake-up word using audio information; and

a current body movement being consistent with a first effect movement.

15 . The device according to claim 12 , wherein the one or more programs causing the one or more processors to determine the first audio further cause the one more processors to:

display at least one second audio, and determining the first audio based on a trigger operation on the at least one second audio within first duration; or

receive an uploaded third audio as the first audio.

16 . The device according to claim 12 , the one or more programs further cause the one more processors to:

determine the key video frame sequence corresponding to the first audio based on a pre-selected first language type corresponding to the first audio.

17 . The device according to claim 12 , wherein the one or more programs causing the one more processors to determine the key video frame sequence corresponding to the first audio further cause the one more processors to:

retrieve the key video frame sequence corresponding to the first audio from a pre-determined key video frame sequence library, wherein the first audio is determined from at least one second audio that is displayed in a display interface, and the key video frame sequence library comprises a corresponding key video frame sequence obtained after the at least one second audio is processed; or

process the first audio to obtain the key video frame sequence corresponding to the first audio.

18 . The device according to claim 17 , wherein the facial feature comprises a mouth shape feature, and the one or more programs causing the one more processors to determine the key video frame sequence further cause the one more processors to:

obtain a second facial image;

obtain, based on the second audio or the first audio, and the second facial image, a first audio and video in which a mouth shape feature in the second facial image is consistent with a mouth shape feature presented when a third audio or the first audio is played; and

use, as the key video frame sequence, a plurality of audio and video frames of the first audio and video in which mouth shape features are inconsistent with a first mouth shape feature.

19 . The device according to claim 18 , wherein the facial feature further comprises a facial organ feature, and the one or more programs causing the one more processors to determine the key video frame sequence further cause the one more processors to:

process, based on a pre-trained facial driving model, the first audio and video and a third facial image, so as to obtain a second audio and video in which a facial organ feature in the third facial image changes;

sequentially determine a facial organ feature in each audio and video frame of the second audio and video;

use, as a key video frame, an audio and video frame in which the facial organ feature is inconsistent with a first facial organ feature; and

determine, based on timestamps of a plurality of key video frames, the key video frame sequence corresponding to the first audio and video.

20 . The device according to claim 12 , wherein the one or more programs causing the one more processors to determine, based on the key video frame sequence and the first facial image, the first facial feature in the first facial image that is presented when the first audio is played further cause the one more processors to:

determine, based on reference feature point data of each key video frame in the key video frame sequence and basic feature point data of the first facial image, first feature point data corresponding to the key video frame; and

determine, based on the first feature point data, the first facial image, and the corresponding basic feature point data, the first facial feature in the first facial image that is presented when the first audio is played,

wherein the reference feature point data corresponds to facial organ feature point data or mouth shape feature point data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2026
From: LI, JIANWEI
To: BYTEDANCE INC.
Reel/Frame 075675/0894 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2026
From: BYTEDANCE INC.; BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
To: LEMON INC.
Reel/Frame 075675/0899 →
Priority Claims (1)
CN 202210911515.1 · Jul 30, 2022 · national
Continuity (1)
Related Publication 20260046469A1 · Feb 12, 2026
References Cited (80)
US 6889325B1 · Sipman · 2005 [cited by examiner]
US 8190435B2 · Li-Chun Wang · 2012 [cited by examiner]
US 8290423B2 · Wang · 2012 [cited by examiner]
US 8537166B1 · Diard · 2013 [cited by examiner]
US 8650603B2 · Doets · 2014 [cited by examiner]
US 8688600B2 · Barton · 2014 [cited by examiner]
US 8725829B2 · Wang · 2014 [cited by examiner]
US 8811885B2 · Wang · 2014 [cited by examiner]
US 8990842B2 · Rowley · 2015 [cited by examiner]
US 10885693B1 · Saragih · 2021 [cited by examiner]
US 11089238B2 · Shaburov · 2021 [cited by examiner]
US 11284144B2 · Kotsopoulos · 2022 [cited by examiner]
US 11831937B2 · Kotsopoulos · 2023 [cited by examiner]
US 12231709B2 · Kotsopoulos · 2025 [cited by examiner]
US 20020072982A1 · Barton · 2002 [cited by examiner]
US 20020083060A1 · Wang · 2002 [cited by examiner]
US 20040199387A1 · Wang · 2004 [cited by examiner]
US 20050028195A1 · Feinleib · 2005 [cited by examiner]
US 20050091274A1 · Stanford · 2005 [cited by examiner]
US 20050192863A1 · Mohan · 2005 [cited by examiner]
US 20050209917A1 · Anderson · 2005 [cited by examiner]
US 20060094500A1 · Dyke · 2006 [cited by examiner]
US 20060195359A1 · Robinson · 2006 [cited by examiner]
US 20060224452A1 · Ng · 2006 [cited by examiner]
US 20060256133A1 · Rosenberg · 2006 [cited by examiner]
US 20070124756A1 · Covell · 2007 [cited by examiner]
US 20070130580A1 · Covell · 2007 [cited by examiner]
US 20070143778A1 · Covell · 2007 [cited by examiner]
US 20070179850A1 · Ganjon · 2007 [cited by examiner]
US 20070192784A1 · Postrel · 2007 [cited by examiner]
US 20070214049A1 · Postrel · 2007 [cited by examiner]
US 20080052062A1 · Stanford · 2008 [cited by examiner]
US 20090044113A1 · Jones · 2009 [cited by examiner]
US 20090198701A1 · Haileselassie · 2009 [cited by examiner]
US 20090313670A1 · Takao · 2009 [cited by examiner]
US 20100034466A1 · Jing · 2010 [cited by examiner]
US 20100114713A1 · Anderson · 2010 [cited by examiner]
US 20110273455A1 · Powar · 2011 [cited by examiner]
US 20120011545A1 · Doets · 2012 [cited by examiner]
US 20120076310A1 · DeBusk · 2012 [cited by examiner]
US 20120117596A1 · Mountain · 2012 [cited by examiner]
US 20120124608A1 · Postrel · 2012 [cited by examiner]
US 20120191231A1 · Wang · 2012 [cited by examiner]
US 20120221131A1 · Wang · 2012 [cited by examiner]
US 20120295560A1 · Mufti · 2012 [cited by examiner]
US 20120297400A1 · Hill · 2012 [cited by examiner]
US 20120316969A1 · Metcalf, III · 2012 [cited by examiner]
US 20120317240A1 · Wang · 2012 [cited by examiner]
US 20130010204A1 · Wang · 2013 [cited by examiner]
US 20130029762A1 · Klappert · 2013 [cited by examiner]
US 20130031579A1 · Klappert · 2013 [cited by examiner]
US 20130042262A1 · Riethmueller · 2013 [cited by examiner]
US 20130044051A1 · Jeong · 2013 [cited by examiner]
US 20130067512A1 · Dion · 2013 [cited by examiner]
US 20130073366A1 · Heath · 2013 [cited by examiner]
US 20130073377A1 · Heath · 2013 [cited by examiner]
US 20130080242A1 · Alhadeff · 2013 [cited by examiner]
US 20130080262A1 · Scott · 2013 [cited by examiner]
US 20130085828A1 · Schuster · 2013 [cited by examiner]
US 20130111519A1 · Rice · 2013 [cited by examiner]
US 20130124073A1 · Ren · 2013 [cited by examiner]
US 20140078148A1 · Diard · 2014 [cited by examiner]
US 20140137139A1 · Jones · 2014 [cited by examiner]
US 20140214532A1 · Barton · 2014 [cited by examiner]
US 20140278845A1 · Teiser · 2014 [cited by examiner]
US 20150128180A1 · Mountain · 2015 [cited by examiner]
US 20150229979A1 · Wood · 2015 [cited by examiner]
US 20150237389A1 · Grouf · 2015 [cited by examiner]
US 20160127793A1 · Grouf · 2016 [cited by examiner]
US 20160165287A1 · Wood · 2016 [cited by examiner]
US 20160323650A1 · Grouf · 2016 [cited by examiner]
US 20170078727A1 · Wood · 2017 [cited by examiner]
US 20170324995A1 · Grouf · 2017 [cited by examiner]
US 20180262805A1 · Grouf · 2018 [cited by examiner]
US 20200195882A1 · Yi · 2020 [cited by examiner]
CN 104756188A · 2015 [cited by applicant]
CN 112188304A · 2021 [cited by applicant]
CN 112911192A · 2021 [cited by applicant]
CN 113282791A · 2021 [cited by applicant]
ISA Intellectual Property Office of Singapore, International Search Report Issued in Application No. PCT/SG2023/050489, Jan. 30, 2024, 7 pages. [cited by applicant]