IP Library Granted Patent US 12,437,739
Granted Patent B2
US 12,437,739 · App. 17/766,911 · Granted Oct 7, 2025

Method and apparatus for determining volume adjustment ratio information, device, and storage medium

Inventors: Xiaobin Zhuang (Shenzhen, CN); Sen Lin (Shenzhen, CN)
Assignee: TENCENT MUSIC ENTERTAINMENT TECHNOLOGY (SHENZHEN) CO., LTD.
G10H1/46G10H1/0008G10H1/366G10H2210/005G10H2210/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,739
App. No.
17/766,911
Granted
Oct 7, 2025
Kind
B2
Abstract

A method for determining volume adjustment ratio information comprises acquiring a first singing audio and an original accompaniment audio corresponding to the first singing audio, wherein the first singing audio is a user singing audio; acquiring a first audio of a non-singing part in the first singing audio, and acquiring a loudness characteristic of the first audio; acquiring, in the original accompaniment audio, a second audio whose playback duration corresponds to a playback duration of the first audio, and acquiring a loudness characteristic of the second audio; and determining a ratio of the loudness characteristic of the first audio to the loudness characteristic of the second audio as adjustment ratio information for adjusting an accompaniment volume of the first singing audio.

Claims (65)

1. A method for determining volume adjustment ratio information, applied to a terminal, comprising:

acquiring a first singing audio and an original accompaniment audio corresponding to the first singing audio, wherein the first singing audio is obtained by synthesizing human voice audio recorded by a user through a singing application and an accompaniment audio of a corresponding singing song;

acquiring a first audio of a non-singing part in the first singing audio, and acquiring a loudness characteristic of the first audio;

acquiring, in the original accompaniment audio, a second audio whose playback duration corresponds to a playback duration of the first audio, and acquiring a loudness characteristic of the second audio;

determining a ratio of the loudness characteristic of the first audio to the loudness characteristic of the second audio as adjustment ratio information for adjusting an accompaniment volume of the first singing audio; and

generating an accompaniment audio with volume being adjusted based on the adjustment ratio information.

2. The method according to claim 1 , wherein said acquiring the first audio of the non-singing part in the first singing audio comprises:

acquiring a playback start time point and a playback end time point of each sentence of lyrics in lyric data corresponding to the first singing audio; and

determining a plurality of first audio segments of the non-singing part in the first singing audio based on the playback start time point and the playback end time point of each sentence of lyrics, and acquiring the first audio by combining the plurality of first audio segments according to a playback time sequence.

3. The method according to claim 2 , wherein said acquiring, in the original accompaniment audio, the second audio whose playback duration corresponds to the playback duration of the first audio comprises:

determining, in the original accompaniment audio, a plurality of second audio segments corresponding to the non-singing part in the first singing audio based on the playback start time point and the playback end time point of each sentence of lyrics, and acquiring the second audio by combining the plurality of second audio segments according to the playback time sequence.

4. The method according to claim 1 , wherein said acquiring the loudness characteristic of the first audio comprises: dividing the first audio into a plurality of third audio segments with a predetermined duration, and determining a loudness characteristic of each of the third audio segments; and

said acquiring the loudness characteristic of the second audio comprises: dividing the second audio into a plurality of fourth audio segments with a predetermined duration, and determining a loudness characteristic of each of the fourth audio segments.

5. The method according to claim 4 , wherein said determining the ratio of the loudness characteristic of the first audio to the loudness characteristic of the second audio as the adjustment ratio information for adjusting the accompaniment volume of the first singing audio comprises:

selecting a first predetermined number of first loudness characteristics which are a previous part of loudness characteristics of all of the third audio segments arranged in ascending order, and selecting a first predetermined number of second loudness characteristics which are a previous part of loudness characteristics of all of the fourth audio segments arranged in ascending order; and

determining a ratio of a sum of the first predetermined number of first loudness characteristics to a sum of the first predetermined number of second loudness characteristics as the adjustment ratio information for adjusting the accompaniment volume of the first singing audio.

6. The method according to claim 4 , wherein said determining the loudness characteristic of each of the third audio segments comprises: uniformly selecting a predetermined number of playback time points in each of the third audio segments, and determining a root mean square of audio amplitudes corresponding to the predetermined number of playback time points as the loudness characteristic of a respective third audio segment; and

said determining the loudness characteristic of each of the fourth audio segments comprises: uniformly selecting the predetermined number of playback time points in each of the fourth audio segments, and determining a root mean square of audio amplitudes corresponding to the predetermined number of playback time points as the loudness characteristic of the fourth audio segment.

7. The method according to claim 1 , wherein after said determining the ratio of the loudness characteristic of the first audio to the loudness characteristic of the second audio as the adjustment ratio information for adjusting the accompaniment volume of the first singing audio, further comprising:

acquiring an adjusted accompaniment audio by adjusting a volume of the original accompaniment audio based on the adjustment ratio information; and

recording a second singing audio based on the adjusted accompaniment audio.

8. The method according to claim 7 , wherein said recording the second singing audio based on the adjusted accompaniment audio comprises:

acquiring segment time information for performing segment re-recording on the first singing audio;

extracting a part of the adjusted accompaniment audio based on the segment time information, and recording a singing audio segment based on the part of the adjusted accompaniment audio; and

acquiring the second singing audio by replacing a singing audio segment corresponding to the segment time information in the first singing audio with the singing audio segment.

9. A non-transitory computer-readable storage medium storing at least one instruction therein, wherein the at least one instruction, when loaded and executed by a processor, causes the processor to perform the method for determining volume adjustment ratio information as defined in claim 1 .

10. An apparatus for determining volume adjustment ratio information, comprising:

a processor; and

a memory configured to store at least one instruction executable by the processor; wherein

the processor, when executing the at least one instruction, is caused to perform a method for determining volume adjustment ratio information comprising:

acquiring a first singing audio and an original accompaniment audio corresponding to the first singing audio, wherein the first singing audio is obtained by synthesizing human voice audio recorded by a user through a singing application and an accompaniment audio of a corresponding singing song;

acquiring a first audio of a non-singing part in the first singing audio, and acquiring a loudness characteristic of the first audio;

acquiring, in the original accompaniment audio, a second audio whose playback duration corresponds to a playback duration of the first audio, and acquiring a loudness characteristic of the second audio;

determining a ratio of the loudness characteristic of the first audio to the loudness characteristic of the second audio as adjustment ratio information for adjusting an accompaniment volume of the first singing audio; and

obtaining an accompaniment audio with volume being adjusted based on the adjustment ratio information.

11. The apparatus according to claim 10 , wherein said acquiring the first audio of the non-singing part in the first singing audio comprises:

acquiring a playback start time point and a playback end time point of each sentence of lyrics in lyric data corresponding to the first singing audio; and

determining a plurality of first audio segments of the non-singing part in the first singing audio based on the playback start time point and the playback end time point of each sentence of lyrics, and acquiring the first audio by combining the plurality of first audio segments according to a playback time sequence.

12. The apparatus according to claim 11 , wherein said acquiring, in the original accompaniment audio, the second audio whose playback duration corresponds to the playback duration of the first audio comprises:

determining, in the original accompaniment audio, a plurality of second audio segments corresponding to the non-singing part in the first singing audio based on the playback start time point and the playback end time point of each sentence of lyrics, and acquiring the second audio by combining the plurality of second audio segments according to the playback time sequence.

13. The apparatus according to claim 10 , wherein said acquiring the loudness characteristic of the first audio comprises: dividing the first audio into a plurality of third audio segments with a predetermined duration, and determining a loudness characteristic of each of the third audio segments; and

said acquiring the loudness characteristic of the second audio comprises: dividing the second audio into a plurality of fourth audio segments with a predetermined duration, and determining a loudness characteristic of each of the fourth audio segments.

14. The apparatus according to claim 13 , wherein said determining the ratio of the loudness characteristic of the first audio to the loudness characteristic of the second audio as the adjustment ratio information for adjusting the accompaniment volume of the first singing audio comprises:

selecting a first predetermined number of first loudness characteristics which are a previous part of loudness characteristics of all of the third audio segments arranged in ascending order, and selecting a first predetermined number of second loudness characteristics which are a previous part of loudness characteristics of all of the fourth audio segments arranged in ascending order; and

determining a ratio of a sum of the first predetermined number of first loudness characteristics to a sum of the first predetermined number of second loudness characteristics as the adjustment ratio information for adjusting the accompaniment volume of the first singing audio.

15. The apparatus according to claim 13 , wherein said determining the loudness characteristic of each of the third audio segments comprises: uniformly selecting a predetermined number of playback time points in each of the third audio segments, and determining a root mean square of audio amplitudes corresponding to the predetermined number of playback time points as the loudness characteristic of a respective third audio segment; and

said determining the loudness characteristic of each of the fourth audio segments comprises: uniformly selecting a predetermined number of playback time points in each of the fourth audio segments, and determining a root mean square of audio amplitudes corresponding to the predetermined number of playback time points as the loudness characteristic of the fourth audio segment.

16. The apparatus according to claim 10 , after said determining the ratio of the loudness characteristic of the first audio to the loudness characteristic of the second audio as the adjustment ratio information for adjusting the accompaniment volume of the first singing audio, the method performed by the processor further comprises:

acquiring an adjusted accompaniment audio by adjusting a volume of the original accompaniment audio based on the adjustment ratio information; and

recording a second singing audio based on the adjusted accompaniment audio.

17. The apparatus according to claim 16 , wherein said recording the second singing audio based on the adjusted accompaniment audio comprises:

acquiring segment time information for performing segment re-recording on the first singing audio;

extracting a part of the adjusted accompaniment audio based on the segment time information, and recording a singing audio segment based on the part of the adjusted accompaniment audio; and

acquiring the second singing audio by replacing a singing audio segment in the first singing audio corresponding to the segment time information with the singing audio segment.

18. A computer device, comprising a processor and a memory storing at least one instruction, wherein the processor, when loading and executing the at least one instruction, is caused to perform a method for determining volume adjustment ratio information comprising:

acquiring a first singing audio and an original accompaniment audio corresponding to the first singing audio, wherein the first singing audio is obtained by synthesizing human voice audio recorded by a user through a singing application and an accompaniment audio of a corresponding singing song;

acquiring a first audio of a non-singing part in the first singing audio, and acquiring a loudness characteristic of the first audio;

acquiring, in the original accompaniment audio, a second audio whose playback duration corresponds to a playback duration of the first audio, and acquiring a loudness characteristic of the second audio;

determining a ratio of the loudness characteristic of the first audio to the loudness characteristic of the second audio as adjustment ratio information for adjusting an accompaniment volume of the first singing audio; and

generating an accompaniment audio with volume being adjusted based on the adjustment ratio information.

19. The computer device according to claim 18 , wherein said acquiring the first audio of the non-singing part in the first singing audio comprises:

acquiring a playback start time point and a playback end time point of each sentence of lyrics in lyric data corresponding to the first singing audio; and

determining a plurality of first audio segments of the non-singing part in the first singing audio based on the playback start time point and the playback end time point of each sentence of lyrics, and acquiring the first audio by combining the plurality of first audio segments according to a playback time sequence.

20. The computer device according to claim 19 , wherein said acquiring, in the original accompaniment audio, the second audio whose playback duration corresponds to the playback duration of the first audio comprises:

determining, in the original accompaniment audio, a plurality of second audio segments corresponding to the non-singing part in the first singing audio based on the playback start time point and the playback end time point of each sentence of lyrics, and acquiring the second audio by combining the plurality of second audio segments according to the playback time sequence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2022
From: ZHUANG, XIAOBIN; LIN, SEN
To: TENCENT MUSIC ENTERTAINMENT TECHNOLOGY (SHENZHEN) CO., LTD.
Reel/Frame 059520/0510 →
Priority Claims (1)
CN 201910958720.1 · Oct 10, 2019 · national
Continuity (1)
Related Publication 20230252964A1 · Aug 10, 2023
References Cited (24)
US 6025553A · Lee · 2000 [cited by examiner]
US 8367923B2 · Humphrey · 2013 [cited by examiner]
US 10001968B1 · Slick · 2018 [cited by examiner]
US 20170177295A1 · Bowen · 2017 [cited by examiner]
US 20180219521A1 · Arimoto · 2018 [cited by examiner]
US 20180247675A1 · Feng · 2018 [cited by examiner]
US 20190005929A1 · Dabon · 2019 [cited by examiner]
US 20210279030A1 · Morsy · 2021 [cited by examiner]
US 20220312096A1 · Zhu · 2022 [cited by examiner]
US 20230238016A1 · Schindler · 2023 [cited by examiner]
US 20230252964A1 · Zhuang · 2023 [cited by examiner]
US 20230335091A1 · Morsy · 2023 [cited by examiner]
CN 1924992A · 2007 [cited by applicant]
CN 107680571A · 2018 [cited by examiner]
CN 107705778A · 2018 [cited by applicant]
CN 109003627A · 2018 [cited by applicant]
CN 109300482A · 2019 [cited by applicant]
CN 109828740A · 2019 [cited by applicant]
CN 109859729A · 2019 [cited by applicant]
CN 110688082A · 2020 [cited by examiner]
International Search Report of the International Searching Authority for State Intellectual Property Office of the People's Republic of China in PCT application No. PCT/CN2020/120044 issued on Jan. 12, 2021, which is an… [cited by applicant]
The State Intellectual Property Office of People's Republic of China, First Office Action in Patent Application No. CN201910958720.1 issued on Feb. 3, 2021, which is a foreign counterpart application corresponding to th… [cited by applicant]
Notification of Completion of Formalities for Patent Register and Notification to Grant Patent Right for Invention Application No. 201910958720.1 Issued on Jul. 5, 2021. [cited by applicant]
Jiang, He; “Observation on the music dissemination of mobile KTV “Changba””; Music Communication, No. 1 2014, Mar. 30, 2014, pp. 76-80. [cited by applicant]