IP Library Granted Patent US 12,347,459
Granted Patent B2
US 12,347,459 · App. 17/768,462 · Granted Jul 1, 2025

Video file generating method and device, terminal, and storage medium

Inventors: Wei Zheng (Beijing, CN); Weiwei Lyu (Beijing, CN)
Assignee: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
G11B27/031G06F3/0487
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,459
App. No.
17/768,462
Granted
Jul 1, 2025
Kind
B2
Abstract

A method and an apparatus for generating a video file, a terminal, and a storage medium are provided. The method includes: presenting, in response to a received video editing instruction, a video editing interface; determining, in response to a clicking operation on a button in the video editing interface, a target audio and a target image for video synthesis; obtaining an audio parameter corresponding to each audio frame; generating a spectrogram corresponding to each audio frame based on the audio parameter; generating, based on the spectrogram and the target image, multiple video frame images that each corresponds to one of the audio frames and includes the spectrum corresponding to the one of the audio frames; and performing, based on the multiple video frame images and the target audio, video encoding to obtain a target video file.

Claims (78)

1. A method for generating a video file, comprising:

presenting, in response to a received video editing instruction, a video editing interface, wherein the video editing interface is configured to enable a selection of an image and a selection of an audio;

determining a target audio and a target image in response to receiving input via the video editing interface;

obtaining, for each of audio frames of the target audio, an audio parameter corresponding to the audio frame;

generating, for each of the audio frames, a spectrogram corresponding to the audio frame based on the obtained audio parameter;

generating, based on the generated spectrogram and the target image, a plurality of video frame images, wherein each of the plurality of video frame images corresponds to one of the audio frames, and each of the plurality of video frame images comprises an image of the spectrum corresponding to the one of the audio frames; and

performing, based on the plurality of video frame images and the target audio, video encoding to generate a target video file.

2. The method according to claim 1 , wherein the generating, for each of the audio frames, a spectrogram corresponding to the audio frame based on the obtained audio parameter comprises:

sampling the target audio based on a preset sampling frequency to obtain an audio parameter corresponding to each of audio frames after sampling; and

generating, for each of the audio frames after sampling, a spectrogram corresponding to the audio frame after sampling by performing Fourier transform on the audio parameter corresponding to the audio frame after sampling.

3. The method according to claim 2 , wherein the generating, for each of the audio frames after sampling, a spectrogram corresponding to the audio frame after sampling comprises:

determining, for each of the audio frames, an amplitude of the audio frame;

determining, for each of the audio frames, a spectrum envelop corresponding to the spectrogram based on the amplitude of the audio frame, to obtain a plurality of spectrum envelopes; and

for each of the plurality of spectrum envelopes, combining the spectrum envelope with the spectrum corresponding to the spectrum envelop, to obtain a plurality of combined spectrograms.

4. The method according to claim 1 , wherein the generating, based on the generated spectrogram and the target image, a plurality of video frame images, wherein each of the plurality of video frame images corresponds to one of the audio frames and comprises the spectrum corresponding to the one of the audio frames comprises:

blurring the target image to obtain a blurred target image;

obtaining a target region of the target image to obtain a target region image;

combining the target region image with the spectrogram of each of the audio frames to obtain a plurality of combined images; and

using each of the plurality of combined images as a foreground and the blurred target image as a background to generate a plurality of video frame images each comprising the spectrogram.

5. The method according to claim 4 , wherein the obtaining a target region of the target image to obtain a target region image comprises:

determining a region that is in the target image and that corresponds to a target object in the target image; and

obtaining, based on the determined region, a region that comprises the target object and that has a target shape, as the target region image.

6. The method according to claim 4 , wherein before the combining the target region image with the spectrogram of each of the audio frames, the method further comprises:

performing color feature extraction on the blurred target image to obtain color features of respective pixels of the blurred target image;

performing a weighted average on the color features of the respective pixels to determine a color of the blurred target image; and

setting the determined color of the blurred image as a color of the spectrogram.

7. The method according to claim 4 , wherein the spectrogram is a spectral histogram, and the combining the target region image with the spectrogram of each of the audio frames to obtain a plurality of combined images comprises:

arranging the spectral histogram around the target region image to form the plurality of combined images, wherein a height of a spectral column in the spectral histogram represents an amplitude of a corresponding one of the audio frames, and an angle of the spectral column in the spectral histogram relative to an edge of the target region image represents a frequency of the corresponding audio frame.

8. The method according to claim 4 , wherein the using each of the plurality of combined images as a foreground and the blurred target image as a background to generate a plurality of video frame images each comprising the spectrogram comprises:

obtaining a relative positional relationship between the foreground and the background in one of the plurality of video frame images corresponding to an adjacent audio frame of the target audio frame; and

generating, based on the obtained relative positional relationship, a video frame image corresponding to the target audio frame, where a presentation position of the foreground in the video frame image corresponding to the target audio frame is rotated by a present angle relative to a representation position of the foreground in the video frame image corresponding to the adjacent audio frame.

9. An apparatus for generating a video file, comprising:

at least one processor; and

at least one memory communicatively coupled to the at least one processor and storing instructions that upon execution by the at least one processor cause the apparatus to perform operations comprising:

presenting, in response to a received video editing instruction, a video editing interface, wherein the video editing interface is configured to enable a selection of an image and a selection of an audio;

determining a target audio and a target image in response to receiving input via the video editing interface;

obtaining, for each of audio frames of the target audio, an audio parameter corresponding to the audio frame;

generating, for each of the audio frames, a spectrogram corresponding to the audio frame based on the obtained audio parameter;

generating, based on the generated spectrogram and the target image, a plurality of video frame images, wherein each of the plurality of video frame images corresponds to one of the audio frames, and each of the plurality of video frame images comprises an image of the spectrum corresponding to the one of the audio frames; and

performing, based on the plurality of video frame images and the target audio, video encoding to generate a target video file.

10. The apparatus according to claim 9 ,

wherein the generating, for each of the audio frames, a spectrogram corresponding to the audio frame based on the obtained audio parameter comprises:

sampling the target audio based on a preset sampling frequency to obtain an audio parameter corresponding to each of audio frames after sampling; and

generating, for each of the audio frames after sampling, a spectrogram corresponding to the audio frame after sampling by performing Fourier transform on the audio parameter corresponding to the audio frame after sampling.

11. The apparatus according to claim 9 ,

wherein the generating, for each of the audio frames after sampling, a spectrogram corresponding to the audio frame after sampling comprises:

determining, for each of the audio frames, an amplitude of the audio frame;

determining, for each of the audio frames, a spectrum envelop corresponding to the spectrogram based on the amplitude of the audio frame, to obtain a plurality of spectrum envelopes; and

for each of the plurality of spectrum envelopes, combining the spectrum envelope with the spectrum corresponding to the spectrum envelop, to obtain a plurality of combined spectrograms.

12. The apparatus according to claim 9 ,

wherein the generating, based on the generated spectrogram and the target image, a plurality of video frame images, wherein each of the plurality of video frame images corresponds to one of the audio frames and comprises the spectrum corresponding to the one of the audio frames comprises:

blurring the target image to obtain a blurred target image;

obtaining a target region of the target image to obtain a target region image;

combining the target region image with the spectrogram of each of the audio frames to obtain a plurality of combined images; and

using each of the plurality of combined images as a foreground and the blurred target image as a background to generate a plurality of video frame images each comprising the spectrogram.

13. The apparatus according to claim 12 ,

wherein the obtaining a target region of the target image to obtain a target region image comprises:

determining a region that is in the target image and that corresponds to a target object in the target image; and

obtaining, based on the determined region, a region that comprises the target object and that has a target shape, as the target region image.

14. The apparatus according to claim 12 ,

wherein before the combining the target region image with the spectrogram of each of the audio frames, the at least one memory communicatively coupled to the at least one processor and storing instructions that upon execution by the at least one processor also cause the apparatus to:

performing color feature extraction on the blurred target image to obtain color features of respective pixels of the blurred target image;

performing a weighted average on the color features of the respective pixels to determine a color of the blurred target image; and

setting the determined color of the blurred image as a color of the spectrogram.

15. The apparatus according to claim 12 ,

wherein the spectrogram is a spectral histogram, and the combining the target region image with the spectrogram of each of the audio frames to obtain a plurality of combined images comprises:

arranging the spectral histogram around the target region image to form the plurality of combined images, wherein a height of a spectral column in the spectral histogram represents an amplitude of a corresponding one of the audio frames, and an angle of the spectral column in the spectral histogram relative to an edge of the target region image represents a frequency of the corresponding audio frame.

16. The apparatus according to claim 12 ,

wherein the using each of the plurality of combined images as a foreground and the blurred target image as a background to generate a plurality of video frame images each comprising the spectrogram comprises:

obtaining a relative positional relationship between the foreground and the background in one of the plurality of video frame images corresponding to an adjacent audio frame of the target audio frame; and

generating, based on the obtained relative positional relationship, a video frame image corresponding to the target audio frame, where a presentation position of the foreground in the video frame image corresponding to the target audio frame is rotated by a present angle relative to a representation position of the foreground in the video frame image corresponding to the adjacent audio frame.

17. A non-transitory storage medium bearing computer-readable instructions that upon execution on a computing device cause the computing device at least to perform operations comprising:

presenting, in response to a received video editing instruction, a video editing interface, wherein the video editing interface is configured to enable a selection of an image and a selection of an audio;

determining a target audio and a target image in response to receiving input via the video editing interface;

obtaining, for each of audio frames of the target audio, an audio parameter corresponding to the audio frame;

generating, for each of the audio frames, a spectrogram corresponding to the audio frame based on the obtained audio parameter;

generating, based on the generated spectrogram and the target image, a plurality of video frame images, wherein each of the plurality of video frame images corresponds to one of the audio frames, and each of the plurality of video frame images comprises an image of the spectrum corresponding to the one of the audio frames; and

performing, based on the plurality of video frame images and the target audio, video encoding to generate a target video file.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: ZHENG, WEI
To: HANGZHOU OCEAN ENGINE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065146/0264 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: HANGZHOU OCEAN ENGINE NETWORK TECHNOLOGY CO., LTD.
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065146/0364 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: LYU, WEIWEI
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065146/0488 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065179/0836 →
Priority Claims (1)
CN 201910974857.6 · Oct 14, 2019 · national
Continuity (1)
Related Publication 20240296870A1 · Sep 5, 2024
References Cited (81)
US 5191319A · Kiltz · 1993 [cited by applicant]
US 6697564B1 · Toklu · 2004 [cited by examiner]
US 8027743B1 · Johnston · 2011 [cited by examiner]
US 9147166B1 · Drame · 2015 [cited by examiner]
US 9158842B1 · Yagnik · 2015 [cited by examiner]
US 9373320B1 · Lyon · 2016 [cited by examiner]
US 9514722B1 · Kim · 2016 [cited by examiner]
US 9679042B2 · Vlack · 2017 [cited by examiner]
US 9953224B1 · Wills · 2018 [cited by examiner]
US 10129658B2 · Rubinstein · 2018 [cited by examiner]
US 10149958B1 · Tran · 2018 [cited by examiner]
US 10236006B1 · Gurijala · 2019 [cited by examiner]
US 10380164B2 · Raichelgauz · 2019 [cited by examiner]
US 10437791B1 · Bebchuk · 2019 [cited by examiner]
US 10685488B1 · Kumar · 2020 [cited by examiner]
US 10816939B1 · Coleman · 2020 [cited by examiner]
US 10841724B1 · Tran · 2020 [cited by examiner]
US 10848590B2 · Raichelgauz · 2020 [cited by examiner]
US 10891723B1 · Chung · 2021 [cited by examiner]
US 10931976B1 · Joze · 2021 [cited by examiner]
US 11604847B2 · Raichelgauz · 2023 [cited by examiner]
US 20070291958A1 · Jehan · 2007 [cited by examiner]
US 20080111887A1 · Cooper · 2008 [cited by examiner]
US 20090070674A1 · Johnston · 2009 [cited by examiner]
US 20090087161A1 · Roberts · 2009 [cited by examiner]
US 20090245603A1 · Koruga · 2009 [cited by examiner]
US 20100077219A1 · Moskowitz · 2010 [cited by examiner]
US 20110261257A1 · Terry · 2011 [cited by examiner]
US 20120321759A1 · Marinkovich · 2012 [cited by examiner]
US 20150066820A1 · Kapur · 2015 [cited by examiner]
US 20150124999A1 · Ren · 2015 [cited by examiner]
US 20150205570A1 · Johnston · 2015 [cited by examiner]
US 20150220806A1 · Heller · 2015 [cited by examiner]
US 20150221321A1 · Christian · 2015 [cited by applicant]
US 20150279383A1 · Crockett · 2015 [cited by examiner]
US 20150310870A1 · Vouin · 2015 [cited by examiner]
US 20150310891A1 · Pello · 2015 [cited by examiner]
US 20150317945A1 · Andress · 2015 [cited by examiner]
US 20150341572A1 · Kelder · 2015 [cited by examiner]
US 20150356992A1 · Wu · 2015 [cited by examiner]
US 20160065864A1 · Guissin · 2016 [cited by examiner]
US 20160070962A1 · Shetty · 2016 [cited by examiner]
US 20160155066A1 · Drame · 2016 [cited by examiner]
US 20160267179A1 · Mei · 2016 [cited by examiner]
US 20170024615A1 · Allen · 2017 [cited by examiner]
US 20170061625A1 · Estrada · 2017 [cited by examiner]
US 20170076753A1 · Vouin · 2017 [cited by examiner]
US 20170142809A1 · Paolini · 2017 [cited by examiner]
US 20170148433A1 · Catanzaro · 2017 [cited by examiner]
US 20170148468A1 · Kim · 2017 [cited by examiner]
US 20170178661A1 · Cahill · 2017 [cited by examiner]
US 20170200315A1 · Lockhart · 2017 [cited by examiner]
US 20170208245A1 · Castillo · 2017 [cited by examiner]
US 20170244938A1 · Al Mohizea · 2017 [cited by examiner]
US 20170278289A1 · Marino · 2017 [cited by examiner]
US 20170303043A1 · Young · 2017 [cited by examiner]
US 20170337693A1 · Baruch · 2017 [cited by examiner]
US 20180011688A1 · Wei · 2018 [cited by examiner]
US 20180205922A1 · Wu · 2018 [cited by examiner]
US 20180238943A1 · Bernsee · 2018 [cited by examiner]
US 20180343477A1 · Loheide · 2018 [cited by examiner]
US 20190020963A1 · Mor · 2019 [cited by examiner]
US 20190034976A1 · Hamedi · 2019 [cited by examiner]
US 20190035431A1 · Attorre · 2019 [cited by examiner]
US 20190109804A1 · Fu · 2019 [cited by examiner]
US 20190180446A1 · Medoff · 2019 [cited by examiner]
US 20190197362A1 · Campanella · 2019 [cited by examiner]
US 20190285673A1 · Skovenborg · 2019 [cited by examiner]
US 20190304076A1 · Nina Paravecino · 2019 [cited by examiner]
US 20200051544A1 · Laput · 2020 [cited by examiner]
US 20200293783A1 · Ramaswamy · 2020 [cited by examiner]
US 20210012769A1 · Vasconcelos · 2021 [cited by examiner]
CN 107135419A · 2017 [cited by applicant]
CN 107749302A · 2018 [cited by applicant]
CN 108769535A · 2018 [cited by applicant]
CN 109120983A · 2019 [cited by applicant]
CN 109309845A · 2019 [cited by applicant]
JP 2013102333A · 2013 [cited by applicant]
International Patent Application No. PCT/CN2020/116576; Int'l Search Report; dated Dec. 16, 2020; 2 pages. [cited by applicant]
“Douyin Kuaishou music video, how to do the effect of cover CD rotation?”; https://jingyan.baidu.com/article/3065b3b6501262becef8a461.html; Baidu; Apr. 2019; accessed Apr. 2022; 3 pages. [cited by applicant]
“Introductory AE: Visual music is not very familiar with this picture, it turns out that beginners can also do it, with tutorials”; https://www.bilibili.com/read/cv2891235?share_medium=iphone&share_plat=ios&share _sourc… [cited by applicant]