IP Library › Granted Patent US 11,948,595
Granted Patent B2
US 11,948,595 · App. 17/282,732 · Granted Apr 2, 2024

Method for detecting audio, device, and storage medium

Inventors: Zhen Li (Guangzhou, CN); Zhenchuan Huang (Guangzhou, CN); Yu Zou (Guangzhou, CN)
Assignee: BIGO TECHNOLOGY PTE. LTD.
G10L25/30G10L15/16G10L17/06G10L17/18G10L17/26G10L25/03G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,948,595
App. No.
17/282,732
Granted
Apr 2, 2024
Kind
B2
Abstract

Provided is a method for detecting audio, which includes acquiring audio file data; determining attribute detection data corresponding to the audio file data; and generating a voice detection result corresponding to the audio file data by voice violation detection on the attribute detection data by a fully connected network model. A device for detecting audio and a non-transitory computer-readable storage medium are also provided.

Claims (97)

1. A method for detecting audio, comprising:

acquiring audio file data;

determining attribute detection data corresponding to the audio file data, wherein the attribute detection data comprises at least two of: user rating data, classification probability data, and voiceprint feature data, the user rating data being indicative of a user rating, the classification probability data being indicative of a classification probability corresponding to a voice violation, and the voiceprint feature data being indicative of a voiceprint feature corresponding to the audio file data; and

generating a voice detection result corresponding to the audio file data by voice violation detection on the attribute detection data by a fully connected network model,

wherein the attribute detection data comprises the classification probability data and at least one of the user rating data and the voiceprint feature data; and

wherein determining the attribute detection data corresponding to the audio file data comprises:

acquiring at least two frames of audio time-domain information by slicing the audio file data;

acquiring magnitude spectrum feature data and voiceprint feature data by feature extraction on the at least two frames of audio time-domain information;

generating feature vector data by splicing the magnitude spectrum feature data and the voiceprint feature data; and

acquiring the classification probability data by voice classification on the feature vector data by a voice classification model.

2. The method according to claim 1 , wherein acquiring the magnitude spectrum feature data and the voiceprint feature data by the feature extraction on the at least two frames of audio time-domain information comprises:

acquiring audio frequency-domain information by frequency-domain transformation on the at least two frames of audio time-domain information;

acquiring the magnitude spectrum feature data by magnitude spectrum feature extraction based on the audio frequency-domain information; and

acquiring the voiceprint feature data by voiceprint feature extraction based on the audio frequency-domain information.

3. The method according to claim 1 , wherein the attribute detection data comprises the voiceprint feature data;

determining the attribute detection data corresponding to the audio file data comprises:

acquiring audio frequency-domain information by frequency-domain transformation on the at least two frames of audio time-domain information;

acquiring first fixed-length data by averaging the audio frequency-domain information; and

acquiring the voiceprint feature data by voiceprint feature extraction by a neural network model based on the first fixed-length data.

4. The method according to claim 3 , further comprising:

acquiring audio file data to be trained from a training set;

acquiring frame time-domain information by slicing the audio file data to be trained by using a moving window;

acquiring frame frequency-domain information by frequency-domain transformation on the frame time-domain information;

acquiring second fixed-length data by averaging the frame frequency-domain information; and

acquiring the neural network model by training according to a neural network algorithm based on the second fixed-length data and label data corresponding to the audio file data.

5. The method according to claim 1 , wherein the attribute detection data comprises the user rating data, the method further comprises:

acquiring history behavior data of a target user, wherein the history behavior data comprises at least one of: history login data, user consumption behavior data, violation history data, and recharge history data; and

acquiring the user rating data based on the history behavior data, and storing the user rating data in a database;

determining the attribute detection data corresponding to the audio file data comprises:

acquiring the user rating data corresponding to the target user from the database based on user information corresponding to the audio file data.

6. The method according to claim 1 , further comprising:

determining that the audio file data comprises violation voice data in response to the voice detection result being a voice violation detection result; and

prohibiting a transmission or playback of the violation voice data, or shielding a voice input from a user corresponding to the violation voice data.

7. The method according to claim 1 , further comprising:

acquiring frame time-domain information by slicing acquired audio file data to be trained by using a moving window;

acquiring magnitude spectrum feature training data and voiceprint feature training data by feature extraction on the frame time-domain information, wherein the feature extraction of the frame time-domain information comprises: magnitude spectrum feature extraction and voiceprint feature extraction;

acquiring third fixed-length data by averaging the magnitude spectrum feature training data;

generating feature vector training data by splicing the magnitude spectrum feature training data and the voiceprint feature training data; and

acquiring the voice classification model by training the third fixed-length data and the feature vector training data.

8. The method according to claim 1 , further comprising:

acquiring attribute detection data to be trained; and

acquiring the fully connected network model by training the attribute detection data to be trained.

9. A device, comprising:

a processor; and

a memory storing at least one instruction therein;

wherein the processor, when loading and executing the at least one instruction, is caused to:

acquire audio file data;

determine attribute detection data corresponding to the audio file data, wherein the attribute detection data comprises at least two of: user rating data, classification probability data, and voiceprint feature data, the user rating data being indicative of a user rating, the classification probability data being indicative of a classification probability corresponding to a voice violation, and the voiceprint feature data being indicative of a voiceprint feature corresponding to the audio file data; and

generate a voice detection result corresponding to the audio file data by voice violation detection on the attribute detection data by a fully connected network model,

wherein the attribute detection data comprises the voiceprint feature data and at least one of the user rating data and the classification probability data; and

wherein in order to determine the attribute detection data corresponding to the audio file data, the processor is caused to:

acquire at least two frames of audio time-domain information by slicing the audio file data,

acquire audio frequency-domain information by frequency-domain transformation on the at least two frames of audio time-domain information;

acquire first fixed-length data by averaging the audio frequency-domain information; and

acquire the voiceprint feature data by voiceprint feature extraction by a neural network model based on the first fixed-length data.

10. The device according to claim 9 , wherein the attribute detection data comprises the classification probability data;

in order to determine the attribute detection data corresponding to the audio file data, the processor is caused to:

acquire magnitude spectrum feature data and voiceprint feature data by feature extraction on the at least two frames of audio time-domain information;

generate feature vector data by splicing the magnitude spectrum feature data and the voiceprint feature data; and

acquire the classification probability data by voice classification on the feature vector data by a voice classification model.

11. The device according to claim 10 , wherein in order to acquire the magnitude spectrum feature data and the voiceprint feature data by the feature extraction on the at least two frames of audio time-domain information, the processor is caused to:

acquire audio frequency-domain information by frequency-domain transformation on the at least two frames of audio time-domain information;

acquire the magnitude spectrum feature data by magnitude spectrum feature extraction based on the audio frequency-domain information; and

acquire the voiceprint feature data by voiceprint feature extraction based on the audio frequency-domain information.

12. The device according to claim 9 , wherein the at least one instruction further causes the processor to:

acquire audio file data to be trained from a training set;

acquire frame time-domain information by slicing the audio file data to be trained by using a moving window;

acquire frame frequency-domain information by frequency-domain transformation on the frame time-domain information;

acquire second fixed-length data by averaging the frame frequency-domain information; and

acquire the neural network model by training according to a neural network algorithm based on the second fixed-length data and label data corresponding to the audio file data.

13. The device according to claim 9 , wherein the attribute detection data comprises the user rating data, the at least one instruction further causes the processor to:

acquire history behavior data of a target user, wherein the history behavior data comprises at least one of: history login data, user consumption behavior data, violation history data, and recharge history data; and

acquire the user rating data based on the history behavior data, and storing the user rating data in a database;

in order to determine the attribute detection data corresponding to the audio file data, the processor is caused to:

acquire the user rating data corresponding to the target user from the database based on user information corresponding to the audio file data.

14. The device according to claim 9 , wherein the at least one instruction further causes the processor to:

determine that the audio file data comprises violation voice data in response to the voice detection result being a voice violation detection result; and

prohibit a transmission or playback of the violation voice data, or shielding a voice input from a user corresponding to the violation voice data.

15. The device according to claim 9 , wherein the at least one instruction further causes the processor to:

acquire frame time-domain information by slicing acquired audio file data to be trained by using a moving window;

acquire magnitude spectrum feature training data and voiceprint feature training data by feature extraction on the frame time-domain information, wherein the feature extraction of the frame time-domain information comprises: magnitude spectrum feature extraction and voiceprint feature extraction;

acquire third fixed-length data by averaging the magnitude spectrum feature training data;

generate feature vector training data by splicing the magnitude spectrum feature training data and the voiceprint feature training data; and

acquire the voice classification model by training the third fixed-length data and the feature vector training data.

16. The device according to claim 9 , wherein the at least one instruction further causes the processor to:

acquire attribute detection data to be trained; and

acquire the fully connected network model by training the attribute detection data to be trained.

17. A non-transitory computer-readable storage medium storing at least one instruction therein, wherein the at least one instruction, when loaded and executed by a processor of a device, causes the device to:

acquire audio file data;

determine attribute detection data corresponding to the audio file data, wherein the attribute detection data comprises at least two of: user rating data, classification probability data, and voiceprint feature data, the user rating data being indicative of a user rating, the classification probability data being indicative of a classification probability corresponding to a voice violation, and the voiceprint feature data being indicative of a voiceprint feature corresponding to the audio file data; and

generate a voice detection result corresponding to the audio file data by voice violation detection on the attribute detection data by a fully connected network model,

wherein the attribute detection data comprises the classification probability data and at least one of the user rating data and the voiceprint feature data; and

wherein in order to determine the attribute detection data corresponding to the audio file data, the processor is caused to:

acquire at least two frames of audio time-domain information by slicing the audio file data;

acquiring magnitude spectrum feature data and voiceprint feature data by feature extraction on the at least two frames of audio time-domain information;

generating feature vector data by splicing the magnitude spectrum feature data and the voiceprint feature data; and

acquiring the classification probability data by voice classification on the feature vector data by a voice classification model.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE ADDRESS PREVIOUSLY RECORDED AT REEL: 055823 FRAME: 0254. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 10, 2021
From: LI, ZHEN; HUANG, ZHENCHUAN; ZOU, YU
To: BIGO TECHNOLOGY PTE. LTD.
Reel/Frame 056185/0449 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2021
From: LI, ZHEN; HUANG, ZHENCHUAN; ZOU, YU
To: BIGO TECHNOLOGY PTE. LTD.
Reel/Frame 055823/0254 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2021
From: LI, ZHEN; HUANG, ZHENCHUAN
To: BIGO TECHNOLOGY PTE. LTD.
Reel/Frame 055813/0403 →
Priority Claims (1)
CN 201811178750.2 · Oct 10, 2018 · national
Continuity (1)
Related Publication 20220005493A1 · Jan 6, 2022