IP Library Granted Patent US 12,505,826
Granted Patent B2
US 12,505,826 · App. 18/203,469 · Granted Dec 23, 2025

Audio processing method and apparatus based on artificial intelligence, electronic device, computer program product, and computer-readable storage medium

Inventors: Binghuai Lin (Shenzhen, CN); Liyuan Wang (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G10L15/02G10L15/16G10L15/22G10L2015/025G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,826
App. No.
18/203,469
Granted
Dec 23, 2025
Kind
B2
Abstract

This application provides an audio processing method performed by an electronic device. The method includes: determining a phoneme feature of at least one phoneme of a given text; determining an audio feature of an audio frame in audio data corresponding to the text; for the audio frame: obtaining a weight of the phoneme feature of the at least one phoneme based on a mapping relationship between the phoneme feature of the at least one phoneme and the audio feature of the audio frame, fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame, and determining a start time and a stop time of a phoneme in the audio data based on the fused feature of the audio frame.

Claims (68)

1 . An audio processing method performed by an electronic device, the method comprising:

determining a phoneme feature of at least one phoneme of a given text;

determining an audio feature of an audio frame in audio data corresponding to the given text;

for the audio frame:

obtaining a weight of the phoneme feature of the at least one phoneme based on a mapping relationship between the phoneme feature of the at least one phoneme and the audio feature of the audio frame;

fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame; and

determining a start time and a stop time of a phoneme in the audio data based on the fused feature of the audio frame.

2 . The method according to claim 1 , wherein a phoneme corresponding to an audio frame in the audio data is determined based on the fused feature of the audio frame.

3 . The method according to claim 1 , wherein the determining a phoneme feature of at least one phoneme of a given text comprises:

for each phoneme:

determining a characteristic representation feature of the phoneme;

determining a location representation feature of the phoneme representing a location of the phoneme in a corresponding text unit; and

adding the location representation feature and the characteristic representation feature to obtain the phoneme feature of the phoneme.

4 . The method according to claim 1 , wherein the fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame comprises:

performing value vector transformation on the phoneme feature of the at least one phoneme to obtain a value vector;

multiplying the weight of the phoneme feature of the at least one phoneme with the value vector to obtain an attention result corresponding to the at least one phoneme; and

fusing the attention result corresponding to the at least one phoneme and the audio feature of the audio frame to obtain the fused feature corresponding to the audio frame.

5 . The method according to claim 1 , wherein the determining a start time and a stop time of a phoneme in the audio data based on the fused feature of the audio frame comprises:

determining at least one audio frame corresponding to the phoneme;

determining a start time and a stop time of consecutive audio frames corresponding to the phoneme as the start time and stop time of the phoneme when the phoneme corresponds to a plurality of consecutive audio frames; and

determining a start time and a stop time of the audio frame corresponding to the phoneme as the start time and stop time of the phoneme in the audio data when the phoneme corresponds to one audio frame.

6 . The method according to claim 1 , wherein the fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame is implemented by invoking an attention fusion network, the determination of a phoneme corresponding to each audio frame is implemented by invoking a phoneme classification network.

7 . An electronic device, comprising:

a memory, configured to store a computer-executable instruction; and

a processor, configured to implement a audio processing method by executing the computer-executable instruction stored in the memory, the method including:

determining a phoneme feature of at least one phoneme of a given text;

determining an audio feature of an audio frame in audio data corresponding to the given text;

for the audio frame:

obtaining a weight of the phoneme feature of the at least one phoneme based on a mapping relationship between the phoneme feature of the at least one phoneme and the audio feature of the audio frame;

fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame; and

determining a start time and a stop time of a phoneme in the audio data based on the fused feature of the audio frame.

8 . The electronic device according to claim 7 , wherein a phoneme corresponding to an audio frame in the audio data is determined based on the fused feature of the audio frame.

9 . The electronic device according to claim 7 , wherein the determining a phoneme feature of at least one phoneme of a given text comprises:

for each phoneme:

determining a characteristic representation feature of the phoneme;

determining a location representation feature of the phoneme representing a location of the phoneme in a corresponding text unit; and

adding the location representation feature and the characteristic representation feature to obtain the phoneme feature of the phoneme.

10 . The electronic device according to claim 7 , wherein the fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame comprises:

performing value vector transformation on the phoneme feature of the at least one phoneme to obtain a value vector;

multiplying the weight of the phoneme feature of the at least one phoneme with the value vector to obtain an attention result corresponding to the at least one phoneme; and

fusing the attention result corresponding to the at least one phoneme and the audio feature of the audio frame to obtain the fused feature corresponding to the audio frame.

11 . The electronic device according to claim 7 , wherein the determining a start time and a stop time of each phoneme in the audio data based on the fused feature of the audio frame comprises:

determining at least one audio frame corresponding to the phoneme;

determining a start time and a stop time of consecutive audio frames corresponding to the phoneme as the start time and stop time of the phoneme when the phoneme corresponds to a plurality of consecutive audio frames; and

determining a start time and a stop time of the audio frame corresponding to the phoneme as the start time and stop time of the phoneme in the audio data when the phoneme corresponds to one audio frame.

12 . The electronic device according to claim 7 , wherein the fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame is implemented by invoking an attention fusion network, the determination of a phoneme corresponding to each audio frame is implemented by invoking a phoneme classification network.

13 . A non-transitory computer-readable storage medium, storing a computer-executable instruction that, when executed by a processor of an electronic device, causing the electronic device to implement an audio processing method including:

determining a phoneme feature of at least one phoneme of a given text;

determining an audio feature of an audio frame in audio data corresponding to the given text;

for the audio frame:

obtaining a weight of the phoneme feature of the at least one phoneme based on a mapping relationship between the phoneme feature of the at least one phoneme and the audio feature of the audio frame;

fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame; and

determining a start time and a stop time of a phoneme in the audio data based on the fused feature of the audio frame.

14 . The non-transitory computer-readable storage medium according to claim 13 , wherein a phoneme corresponding to an audio frame in the audio data is determined based on the fused feature of the audio frame.

15 . The non-transitory computer-readable storage medium according to claim 13 , wherein the determining a phoneme feature of at least one phoneme of a given text comprises:

for each phoneme:

determining a characteristic representation feature of the phoneme;

determining a location representation feature of the phoneme representing a location of the phoneme in a corresponding text unit; and

adding the location representation feature and the characteristic representation feature to obtain the phoneme feature of the phoneme.

16 . The non-transitory computer-readable storage medium according to claim 13 , wherein the fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame comprises:

performing value vector transformation on the phoneme feature of the at least one phoneme to obtain a value vector;

multiplying the weight of the phoneme feature of the at least one phoneme with the value vector to obtain an attention result corresponding to the at least one phoneme; and

fusing the attention result corresponding to the at least one phoneme and the audio feature of the audio frame to obtain the fused feature corresponding to the audio frame.

17 . The non-transitory computer-readable storage medium according to claim 13 , wherein the determining a start time and a stop time of each phoneme in the audio data based on the fused feature of the audio frame comprises:

determining at least one audio frame corresponding to the phoneme;

determining a start time and a stop time of consecutive audio frames corresponding to the phoneme as the start time and stop time of the phoneme when the phoneme corresponds to a plurality of consecutive audio frames; and

determining a start time and a stop time of the audio frame corresponding to the phoneme as the start time and stop time of the phoneme in the audio data when the phoneme corresponds to one audio frame.

18 . The non-transitory computer-readable storage medium according to claim 13 , wherein the fusing the audio feature of the audio frame and the phoneme feature of the at least one phoneme based on the weight of the phoneme feature of the at least one phoneme to obtain a fused feature of the audio frame is implemented by invoking an attention fusion network, the determination of a phoneme corresponding to each audio frame is implemented by invoking a phoneme classification network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2023
From: LIN, BINGHUAI; WANG, LIYUAN
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 063816/0449 →
Priority Claims (1)
CN 202111421900.X · Nov 26, 2021 · national
Continuity (2)
Continuation PCTCN2022122553 · Sep 29, 2022
Related Publication 20230306959A1 · Sep 28, 2023
References Cited (17)
US 11072344B2 · Provost · 2021 [cited by examiner]
US 20210233513A1 · Su et al. · 2021 [cited by applicant]
CN 104756182A · 2015 [cited by applicant]
CN 106297828A · 2017 [cited by applicant]
CN 111105785A · 2020 [cited by applicant]
CN 111312231A · 2020 [cited by applicant]
CN 111754978A · 2020 [cited by applicant]
CN 113436608A · 2021 [cited by applicant]
CN 113536029A · 2021 [cited by applicant]
CN 114360504A · 2022 [cited by applicant]
GB 2575423A · 2020 [cited by applicant]
JP 2001175275A · 2001 [cited by applicant]
JP 2004077901A · 2004 [cited by applicant]
Tencent Technology, WO, PCT/CN2022/122553, Jan. 5, 2023, 4 pgs. [cited by applicant]
Tencent Technology, IPRP, PCT/CN2022/122553, May 2, 2024, 5 pgs. [cited by applicant]
Tencent Technology, Extended European Search Report, EP Patent Application No. 22897381.4, Oct. 8, 2024, 12 pgs. [cited by applicant]
Binghuai Lin et al., “Learning Acoustic Frame Labeling for Phoneme Segmentation with Regularized Attention Mechanism”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), DOI: 10.1109/ICAS… [cited by applicant]