IP Library › Granted Patent US 12,046,225
Granted Patent B2
US 12,046,225 · App. 16/990,869 · Granted Jul 23, 2024

Audio synthesizing method, storage medium and computer equipment

Inventors: Lingrui Cui (Shenzhen, CN); Yi Lu (Shenzhen, CN); Yiting Zhou (Shenzhen, CN); Xinwan Wu (Shenzhen, CN); Yidong Liang (Shenzhen, CN); Xiao Mei (Shenzhen, CN); Qihang Feng (Shenzhen, CN); Fangxiao Wang (Shenzhen, CN); Huifu Jiang (Shenzhen, CN); Shangzhen Zheng (Shenzhen, CN); Le Yu (Shenzhen, CN); Shengfei Xia (Shenzhen, CN); Jingxuan Wang (Shenzhen, CN); Ran Zhang (Shenzhen, CN); Yifan Guo (Shenzhen, CN); Zhenyun Zhang (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G10L13/033G06F16/638G06F16/686G10H7/00G11B27/02G10H2210/021G10H2210/101G10H2240/141
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,046,225
App. No.
16/990,869
Granted
Jul 23, 2024
Kind
B2
Abstract

This application relates to an audio synthesis method, a storage medium, and a computer device. The method includes: obtaining a target text; determining a target song according to a selection instruction; synthesizing a self-made song using the target text and tune information of the target song according to a tune control model, the target text being used as the lyrics of the self-made song. The solutions provided in this application improve an audio playback effect.

Claims (71)

1. An audio synthesis method performed at a computer device having a processor and memory storing a plurality of programs to be executed by the processor, the method comprising:

obtaining a target text from a user;

determining a target song according to a selection instruction, the target text being independent from the target song and the target song being associated with an original singer;

selecting a virtual character from a plurality of virtual characters, wherein the plurality of virtual characters are third-parties different from the user and the original singer;

searching for a tune control model corresponding to the virtual character;

synthesizing a self-made song using the target text and tune information of the target song according to the found tune control model, the target text being used as lyrics of the self-made song and timbre of the self-made song conforming to the virtual character different from the user and the original singer; and

playing the self-made song while recording a video using the computer device. wherein the self-made song becomes background audio of the video being recorded.

2. The method according to claim 1 , wherein the synthesizing a self-made song using the target text and tune information of the target song according to a tune control model comprises:

searching for the tune information matching the target song;

inputting the target text and the tune information into the tune control model, and determining a tune feature corresponding to each character in the target text according to the tune information by using a hidden layer of the tune control model; and

outputting, by using an output layer of the tune control model, the self-made song obtained after speech synthesis is performed on each character in the target text according to the corresponding tune feature.

3. The method according to claim 1 , wherein the target song is selected from multiple candidate songs; and the tune control model is trained by:

collecting candidate song audio corresponding to the candidate songs;

determining a candidate song tune corresponding to each candidate song according to the collected candidate song audio;

obtaining a text sample; and

obtaining the tune control model through training according to the text sample and the candidate song tune.

4. The method according to claim 1 , further comprising synthesizing self-made audio using the target text according to a timbre control model matching a target speaking object, wherein the timbre control model is trained by:

collecting an audio material corresponding to each candidate speaking object;

determining a phoneme material sequence corresponding to the corresponding candidate speaking object according to each audio material; and

obtaining the timbre control model matching each candidate speaking object through training by using the phoneme material sequence corresponding to each candidate speaking object.

5. The method according to claim 4 , wherein the synthesizing the self-made audio using the target text according to a timbre control model comprises:

searching for the timbre control model matching the target speaking object;

determining a phoneme sequence corresponding to the target text;

synthesizing a self-made speech according to the phoneme sequence by using the timbre control model; and

synthesizing the self-made audio according to the self-made speech and a background accompaniment.

6. The method according to claim 1 , further comprising:

configuring the self-made audio as background audio;

superimposing a virtual object additional element on an acquired image to obtain a video frame; and

generating a recorded video using the background audio and the video frame obtained through superimposition.

7. The method according to claim 1 , further comprising:

configuring the self-made audio as background audio;

generating a call video frame according to a picture corresponding to a target speaking object and an acquired image; and

generating a recorded video using the background audio and the generated call video frame.

8. A computer device, comprising a memory and a processor, the memory storing a plurality of computer programs, the computer programs, when executed by the processor, causing the computer device to perform a plurality of operations including:

obtaining a target text from a user;

determining a target song according to a selection instruction, the target text being independent from the target song and the target song being associated with an original singer;

selecting a virtual character from a plurality of virtual characters, wherein the plurality of virtual characters are third-parties different from the user and the original singer;

searching for a tune control model corresponding to the virtual character;

synthesizing a self-made song using the target text and tune information of the target song according to the found tune control model, the target text being used as lyrics of the self-made song and timbre of the self-made song conforming to the virtual character different from the user and the original singer; and

playing the self-made song while recording a video using the computer device. wherein the self-made song becomes background audio of the video being recorded.

9. The computer device according to claim 8 , wherein the synthesizing a self-made song using the target text and tune information of the target song according to a tune control model comprises:

searching for the tune information matching the target song;

inputting the target text and the tune information into the tune control model, and determining a tune feature corresponding to each character in the target text according to the tune information by using a hidden layer of the tune control model; and

outputting, by using an output layer of the tune control model, the self-made song obtained after speech synthesis is performed on each character in the target text according to the corresponding tune feature.

10. The computer device according to claim 8 , wherein the plurality of operations further comprise:

configuring the self-made song as background audio; and

recording a video using the background audio.

11. The computer device according to claim 8 , wherein the plurality of operations further comprises synthesizing self-made audio using the target text according to a timbre control model, comprising:

searching for the timbre control model matching a target speaking object;

determining a phoneme sequence corresponding to the target text;

synthesizing a self-made speech according to the phoneme sequence by using the timbre control model; and

synthesizing the self-made audio according to the self-made speech and a background accompaniment.

12. The computer device according to claim 8 , wherein the plurality of operations further comprise:

configuring the self-made audio as background audio;

superimposing a virtual object additional element on an acquired image to obtain a video frame; and

generating a recorded video using the background audio and the video frame obtained through superimposition.

13. The computer device according to claim 8 , wherein the plurality of operations further comprise:

configuring the self-made audio as background audio;

generating a call video frame according to a picture corresponding to a target speaking object and an acquired image; and

generating a recorded video using the background audio and the generated call video frame.

14. A non-transitory computer-readable storage medium storing a plurality of computer programs, the computer programs, when executed by a processor of a computer device, causing the computer device to perform a plurality of operations including:

obtaining a target text from a user;

determining a target song according to a selection instruction, the target text being independent from the target song and the target song being associated with an original singer;

selecting a virtual character from a plurality of virtual characters, wherein the plurality of virtual characters are third-parties different from the user and the original singer;

searching for a tune control model corresponding to the virtual character;

synthesizing a self-made song using the target text and tune information of the target song according to the found tune control model, the target text being used as lyrics of the self-made song and timbre of the self-made song conforming to the virtual character different from the user and the original singer;

playing the self-made song while recording a video using the computer device. wherein the self-made song becomes background audio of the video being recorded.

15. The non-transitory computer-readable storage medium according to claim 14 , wherein the synthesizing a self-made song using the target text and tune information of the target song according to a tune control model comprises:

searching for the tune information matching the target song;

inputting the target text and the tune information into the tune control model, and determining a tune feature corresponding to each character in the target text according to the tune information by using a hidden layer of the tune control model; and

outputting, by using an output layer of the tune control model, the self-made song obtained after speech synthesis is performed on each character in the target text according to the corresponding tune feature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2021
From: CUI, LINGRUI; LU, YI; ZHOU, YITING; WU, XINWAN; LIANG, YIDONG; MEI, XIAO; FENG, QIHANG; WANG, FANGXIAO; JIANG, HUIFU; ZHENG, SHANGZHEN; YU, LE; XIA, SHENGFEI; WANG, JINGXUAN; ZHANG, RAN; GUO, YIFAN; ZHANG, ZHENYUN
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 055218/0963 →
Priority Claims (1)
CN 201810730283.3 · Jul 5, 2018 · national
Continuity (2)
Continuation PCTCN2019089678 · May 31, 2019
Related Publication 20200372896A1 · Nov 26, 2020