IP Library Granted Patent US 12,555,607
Granted Patent B2
US 12,555,607 · App. 18/023,286 · Granted Feb 17, 2026

Audio data processing method and apparatus, and device and storage medium

Inventors: Cheng Li (Beijing, CN); Hao Huang (Beijing, CN)
Assignee: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
G11B27/031G06F3/0482G06F3/04847G06F3/165G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,607
App. No.
18/023,286
Granted
Feb 17, 2026
Kind
B2
Abstract

The embodiments of the present disclosure relate to an audio data processing method and apparatus, and a device and a storage medium. The method comprises: acquiring a first play position of first audio data, and an audition instruction of a user for a first sound effect; adding the first sound effect to a first audio clip in the first audio data, generating sound effect audition data and playing same; and if a first addition instruction of the user for a second sound effect is received, according to information of a first addition length carried in the first addition instruction, adding the second sound effect to a second audio clip, which takes the first play position as a start position, in the first audio data, so as to obtain second audio data. By means of the solution provided in the embodiments of the present disclosure, a sound effect addition operation can be simplified, sound effect addition results are enriched, and the user experience is also enhanced.

Claims (59)

1 . A method for processing audio data, comprising:

acquiring a first playback position on first audio data and an audition instruction of a user for at least one first sound effect;

adding the at least one first sound effect to a first audio segment in the first audio data to generate sound effect audition data, and auditioning the generated sound effect audition data, wherein the first audio segment starts from the first playback position;

selecting a sound effect from the auditioned at least one first sound effect as a second sound effect, based on the auditioned sound effect audition data generated by adding the at least one first sound effect to the first audio segment;

receiving a first adding instruction of the user for the second sound effect, wherein the first adding instruction comprises information on a first adding length of the second sound effect to be added in the first audio data; and adding the second sound effect to a second audio segment in the first audio data to obtain second audio data, wherein the second audio segment starts from the first playback position, and has a length of the first adding length,

wherein adding the at least one first sound effect to the first audio segment in the first audio data to generate the sound effect audition data comprises:

displaying one or more sound options, and selecting at least one from the one or more sound options as a target sound;

recognizing the target sound from the first audio segment through a preset sound recognition model; and

adding the at least one first sound effect to the target sound in the first audio segment, to generate audition data of the first sound effect on the first audio segment.

2 . The method according to claim 1 , wherein the first audio data is audio data of a to-be-edited video in a video editing interface.

3 . The method according to claim 2 , wherein acquiring the first playback position on the first audio data and the audition instruction of the user for the at least one first sound effect comprising:

displaying the video editing interface, wherein the video editing interface comprises a playback progress control for a video and a sound effect control for the first sound effect; and

acquiring the first playback position selected by the user through the playback progress control, and the audition instruction triggered by the user through the sound effect control.

4 . The method according to claim 1 , further comprising:

returning a playback position of the first audio data to the first playback position, after playing of the sound effect audition data is finished.

5 . The method according to claim 1 , wherein adding the second sound effect to the second audio segment in the first audio data comprises:

adding the second sound effect to a target sound in the second audio segment.

6 . The method according to claim 1 , wherein after adding the second sound effect to the second audio segment in the first audio data to obtain the second audio data, the method further comprises:

obtaining a second playback position on the second audio data, and a second adding instruction of the user for a third sound effect, wherein the second adding instruction comprises information on a second adding length of the third sound effect to be added to the second audio data; and

adding the third sound effect to a third audio segment in the second audio data to obtain third audio data, wherein the third audio segment starts from the second playback position and has a length of the second adding length.

7 . The method according to claim 6 , further comprising:

applying a fade-out effect on the second sound effect and a fade-in effect on the third sound effect, in a case that an end position of the second sound effect and the second playback position are two consecutive playback positions on the third audio data.

8 . A terminal device, comprising:

a memory; and

a processor, wherein:

the memory stores a computer program; and the computer program, when executed by the processor, causes the processor to:

acquire a first playback position on first audio data and an audition instruction of a user for at least one first sound effect;

add the at least one first sound effect to a first audio segment in the first audio data to generate sound effect audition data, and audition the sound effect audition data, wherein the first audio segment starts from the first playback position;

select a sound effect from the auditioned first sound effect as a second sound effect, based on the auditioned sound effect audition data generated by adding the at least one first sound effect to the first audio segment;

receive a first adding instruction of the user for the second sound effect, wherein the first adding instruction comprises information on a first adding length of the second sound effect to be added in the first audio data; and

add the second sound effect to a second audio segment in the first audio data to obtain second audio data, wherein the second audio segment starts from the first playback position, and has a length of the first adding length,

wherein the computer program, when executed by the processor, causes the processor to:

present an interface comprising one or more sound options, and select at least one from the one or more sound options as a target sound;

recognize the target sound from the first audio segment through a preset sound recognition model; and

add the at least one first sound effect to the target sound in the first audio segment, to generate audition data of the first sound effect on the first audio segment.

9 . The terminal device according to claim 8 , wherein the first audio data is audio data of a to-be-edited video in a video editing interface.

10 . The terminal device according to claim 9 , wherein the computer program, when executed by the processor, causes the processor to:

display the video editing interface, wherein the video editing interface comprises a playback progress control for a video and a sound effect control for the first sound effect; and

acquire the first playback position selected by the user through the playback progress control, and the audition instruction triggered by the user through the sound effect control.

11 . The terminal device according to claim 8 , wherein the computer program, when executed by the processor, causes the processor to:

return a playback position of the first audio data to the first playback position, after playing of the sound effect audition data is finished.

12 . The terminal device according to claim 8 , wherein the computer program, when executed by the processor, causes the processor to:

add the second sound effect to a target sound in the second audio segment.

13 . The terminal device according to claim 8 , wherein the computer program, when executed by the processor, causes the processor to:

obtain a second playback position on the second audio data, and a second adding instruction of the user for a third sound effect, wherein the second adding instruction comprises information on a second adding length of the third sound effect to be added to the second audio data; and

add the third sound effect to a third audio segment in the second audio data to obtain third audio data, wherein the third audio segment starts from the second playback position and has a length of the second adding length.

14 . The terminal device according to claim 13 , wherein the computer program, when executed by the processor, causes the processor to:

apply a fade-out effect on the second sound effect and a fade-in effect on the third sound effect, in a case that an end position of the second sound effect and the second playback position are two consecutive playback positions on the third audio data.

15 . A non-transitory computer-readable storage medium storing a computer program, wherein;

the computer program, when executed by a processor, causes the processor to:

acquire a first playback position on first audio data and an audition instruction of a user for at least one first sound effect;

add the at least one first sound effect to a first audio segment in the first audio data to generate sound effect audition data, and audition the generated sound effect audition data, wherein the first audio segment starts from the first playback position;

select a sound effect from the auditioned first sound effect as a second sound effect, based on the audition sound effect audition data generated by adding the at least one first sound effect to the first audio segment;

receive a first adding instruction of the user for the second sound effect, wherein the first adding instruction comprises information on a first adding length of the second sound effect to be added in the first audio data; and

add the second sound effect to a second audio segment in the first audio data to obtain second audio data, wherein the second audio segment starts from the first playback position, and has a length of the first adding length,

wherein the computer program, when executed by a processor, causes the processor to:

present an interface comprising one or more sound options, and select at least one from the one or more sound options as a target sound;

recognize the target sound from the first audio segment through a preset sound recognition model; and

add the at least one first sound effect to the target sound in the first audio segment, to generate audition data of the first sound effect on the first audio segment.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: LI, CHENG
To: MIYOU INTERNET TECHNOLOGY (SHANGHAI) CO., LTD.
Reel/Frame 064033/0423 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: HUANG, HAO
To: SHENZHEN JINRITOUTIAO TECHNOLOGY CO., LTD.
Reel/Frame 064033/0505 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: MIYOU INTERNET TECHNOLOGY (SHANGHAI) CO., LTD.
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 064033/0525 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: SHENZHEN JINRITOUTIAO TECHNOLOGY CO., LTD.
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 064033/0580 →
Priority Claims (1)
CN 202010873112.3 · Aug 26, 2020 · national
Continuity (1)
Related Publication 20230307004A1 · Sep 28, 2023
References Cited (26)
US 9728225B2 · Hsu · 2017 [cited by applicant]
US 10062367B1 · Evans et al. · 2018 [cited by applicant]
US 20050066279A1 · LeBarton · 2005 [cited by examiner]
US 20190026068A1 · Gan et al. · 2019 [cited by applicant]
US 20200302933A1 · Arciero · 2020 [cited by examiner]
US 20200410967A1 · Jiang · 2020 [cited by examiner]
US 20210397411A1 · Plom · 2021 [cited by examiner]
US 20220028427A1 · Matsuda · 2022 [cited by examiner]
CN 102724423A · 2012 [cited by applicant]
CN 106559572A · 2017 [cited by applicant]
CN 108965757A · 2018 [cited by applicant]
CN 109346111A · 2019 [cited by applicant]
CN 109754825A · 2019 [cited by applicant]
CN 110377212A · 2019 [cited by applicant]
CN 111142838A · 2020 [cited by applicant]
CN 112165647A · 2021 [cited by applicant]
JP 2009204907A · 2009 [cited by applicant]
JP 2014095806A · 2014 [cited by applicant]
JP 2019205151A · 2019 [cited by applicant]
WO 2017013762A1 · 2017 [cited by applicant]
WO 2020151008A1 · 2020 [cited by applicant]
Extended European Search Report in EP21860468.4, mailed Dec. 4, 2023, 10 pages. [cited by applicant]
International Search Report issued in International Patent Application No. PCT/CN2021/114706 on Nov. 25, 2021. [cited by applicant]
Communication pursuant to Article 94(3) EPC for European Patent Application No. 21860468.4, mailed on Oct. 8, 2024, 6 pages. [cited by applicant]
Decision to Grant a Patent for Japanese Application No. 2023-513340, mailed on Oct. 1, 2024, 5 pages. [cited by applicant]
Office Action for Japanese Patent Application No. 2023-513340, mailed Mar. 12, 2024, 10 pages. [cited by applicant]