IP Library Granted Patent US 12,670,642
Granted Patent B2
US 12,670,642 · App. 18/697,431 · Granted Jun 30, 2026

Method, apparatus, device and storage medium for video recording

Inventors: Ziyang Wu (Beijing, CN); Lu Tao (Beijing, CN); Yixing Zhu (Beijing, CN); Yi Wang (Singapore, SG); Songda Li (Beijing, CN); Suiyu Feng (Beijing, CN)
Assignee: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
G10L25/51G06T11/60G10L15/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,642
App. No.
18/697,431
Granted
Jun 30, 2026
Kind
B2
Abstract

Example embodiments of the present disclosure relates to a method, apparatus, device and storage medium for video recording. The method comprises: collecting voice data and an image of a target user; determining a match degree between the voice data and a reference audio; determining a target effect based on the match degree; adding the target effect to the collected image to obtain a target image; encoding the voice data and the target image to obtain a target video.

Claims (69)

1 . A method for video recording, comprising:

collecting voice data and an image of a user, wherein the image is a portrait of the user comprising a face of the user;

determining a match degree between the voice data and a reference audio;

determining an effect based on the match degree, wherein the effect comprises effects applied to the portrait of the user;

adding the effect to the collected image to obtain an edited image; and

encoding the voice data and the edited image to obtain a video.

2 . The method according to claim 1 , before the collecting voice data and the image of the user, further comprising:

receiving a reference audio selected by the user;

segmenting the reference audio to obtain a plurality of segments of sub-audios;

playing the plurality of segments of sub-audios sequentially according to timestamps, to cause the user to input voice by imitating the played sub-audios to input voice.

3 . The method according to claim 2 , wherein the segmenting the reference audio to obtain the plurality of segments of sub-audios, in response to determining the reference audio as a song, comprises:

obtaining a MIDI file and lyrics of the song;

decomposing the lyrics to obtain a plurality of sub-lyrics;

the playing the plurality of segments of sub-audios sequentially according to timestamps to cause the user to input voice by imitating the played sub-audios comprises:

playing the plurality of sub-lyrics and the MIDI file sequentially according to timestamps, to cause the user to sing the song according to melodies corresponding to the played sub-lyrics and the MIDI file.

4 . The method of claim 1 , wherein the determining the match degree between the voice data and the reference audio comprises:

extracting a voice feature of the voice data and an audio feature of the reference audio;

determining a similarity between the voice feature and the audio feature;

determining the similarity as the match degree between the voice data and the reference audio.

5 . The method of claim 1 , further comprising:

pre-establishing an association between a match degree and an effect;

the determining an effect based on the match degree comprising:

determining the effect corresponding to the match degree based on the association.

6 . The method according to claim 1 , wherein the determining the effect based on the match degree comprises:

extracting features of the user in the collected image to obtain feature information of the user;

determining the effect based on the feature information and the match degree.

7 . The method according to claim 1 , wherein the adding the effect to the collected image comprises:

adding the effect to an image collected between a current match degree and a next match degree; or adding the effect to a predetermined number of images collected from the current match degree.

8 . The method according to claim 1 , wherein the adding the effect to the collected image to obtain an edited image comprises:

calling an effect package corresponding to the effect to perform effect processing on the collected image to obtain the edited image.

9 . An electronic device, comprising:

one or more processing device;

storage means, configured for storing one or more programs;

the one or more programs, when executed by the one or more processing device, causing the one or more processing device to perform a method comprising:

collecting voice data and an image of a user, wherein the image is a portrait of the user comprising a face of the user;

determining a match degree between the voice data and a reference audio;

determining an effect based on the match degree, wherein the effect comprises effects applied to the portrait of the user;

adding the effect to the collected image to obtain an edited image; and

encoding the voice data and the edited image to obtain a video.

10 . The device of claim 9 , before the collecting voice data and the image of the user, the one or more processing device is further caused to perform:

receiving a reference audio selected by the user;

segmenting the reference audio to obtain a plurality of segments of sub-audios;

playing the plurality of segments of sub-audios sequentially according to timestamps, to cause the user to input voice by imitating the played sub-audios.

11 . The device according to claim 10 , wherein the one or more processing device is further caused to segment the reference audio to obtain the plurality of segments of sub-audios, in response to determining the reference audio as a song, by:

obtaining a MIDI file and lyrics of the song;

decomposing the lyrics to obtain a plurality of sub-lyrics;

the playing the plurality of segments of sub-audios sequentially according to timestamps to cause the user to input voice by imitating the played sub-audios comprises:

playing the plurality of sub-lyrics and the MIDI file sequentially according to timestamps, to cause the user to sing the song according to melodies corresponding to the played sub-lyrics and the MIDI file.

12 . The device according to claim 9 , wherein the one or more processing device is further caused to determine the match degree between the voice data and the reference audio by:

extracting a voice feature of the voice data and an audio feature of the reference audio;

determining a similarity between the voice feature and the audio feature;

determining the similarity as the match degree between the voice data and the reference audio.

13 . The device according to claim 9 , wherein the one or more processing device is further caused to perform:

pre-establishing an association between a match degree and an effect;

the determining an effect based on the match degree comprising:

determining the effect corresponding to the match degree based on the association.

14 . The device according to claim 9 , wherein the one or more processing device is further caused to determine the effect based on the match degree by:

extracting features of the user in the collected image to obtain feature information of the user;

determining the effect based on the feature information and the match degree.

15 . The device according to claim 9 , wherein the one or more processing device is further caused to add the effect to the collected image by:

adding the effect to an image collected between a current match degree and a next match degree; or adding the effect to a predetermined number of images collected from the current match degree.

16 . The device according to claim 9 , wherein the one or more processing device is further caused to add the effect to the collected image to obtain an edited image by:

calling an effect package corresponding to the effect to perform effect processing on the collected image to obtain the edited image.

17 . A non-transitory computer readable medium, having a computer program stored thereon, the computer program, when executed by processing device, performing a method comprising:

collecting voice data and an image of a user, wherein the image is a portrait of the user comprising a face of the user;

determining a match degree between the voice data and a reference audio;

determining an effect based on the match degree, wherein the effect comprises effects applied to the portrait of the user;

adding the effect to the collected image to obtain an edited image; and

encoding the voice data and the edited image to obtain a video.