IP Library › Granted Patent US 11,967,340
Granted Patent B2
US 11,967,340 · App. 18/340,767 · Granted Apr 23, 2024

Method for detecting speech in audio data

Inventors: Subong Choi (Seoul, KR); Dongchan Shin (Seoul, KR); Jihwa Lee (Seoul, KR)
Assignee: ActionPower Corp.
G10L25/78G10L21/0272G10L25/18G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,967,340
App. No.
18/340,767
Granted
Apr 23, 2024
Kind
B2
Abstract

Disclosed is a method for detecting a voice from audio data, performed by a computing device according to an exemplary embodiment of the present disclosure. The method includes obtaining audio data; generating image data based on a spectrum of the obtained audio data; analyzing the generated image data by utilizing a pre-trained neural network model; and determining whether an automated response system (ARS) voice is included in the audio data, based on the analysis of the image data.

Claims (50)

1. A method for detecting a voice from audio data, performed by a computing device including at least one processor, the method comprising:

obtaining audio data;

generating image data based on a spectrum of the obtained audio data;

extracting a feature on the image data by utilizing a first neural network model; and

extracting a time-series feature of the image data by utilizing a second neural network model sequentially connected with the first neural network model;

determining whether an automated response system (ARS) voice is included in each of a plurality of sections of the audio data, based on the feature on the image data and the time-series feature of the image data;

eliminating a section including the ARS voice, among the plurality of sections of the audio data; and

generating input data or training data for analysis of spoken voice based on remaining sections excluding the eliminated section, among the plurality of sections of the audio data.

2. The method according to claim 1 , wherein the generating of the image data includes:

generating a mel-spectrogram based on the spectrum of the obtained audio data.

3. The method according to claim 2 , wherein the generating of the image data further includes:

generating a plurality of frames by dividing the generated mel-spectrogram in a predetermined time unit.

4. The method according to claim 3 ,

wherein the extracting the feature on the image data includes:

extracting a feature on the image of the mel-spectrogram by utilizing the first neural network model; and

wherein the extracting the time-series feature of the image data includes:

further extracting a time-series feature of the mel-spectrogram by utilizing the second neural network model.

5. The method according to claim 4 , wherein the first neural network model includes a convolutional neural network (CNN) model and the second neural network model includes a long short-term memory (LSTM) model.

6. The method according to claim 4 , wherein the mel-spectrogram is divided into the plurality of frames,

the first neural network model includes a plurality of first sub modules, and

the extracting of the feature on the image of the mel-spectrogram includes:

extracting the feature on the image for each of the plurality of frames by utilizing the plurality of first sub modules.

7. The method according to claim 2 , wherein the determining of whether an ARS voice is included in the audio data includes:

determining a first probability representing whether a spoken voice is included in the audio data and a second probability representing whether the ARS voice is included in the audio data;

determining whether the spoken voice is included in the audio data based on the first probability; and

determining whether the ARS voice is included in the audio data based on the second probability.

8. The method according to claim 7 , wherein the mel-spectrogram is divided into a plurality of frames,

the determining of whether the ARS voice is included in the audio data further includes:

determining the first probability and the second probability with respect to each of the plurality of frames;

determining whether the spoken voice is included in each of the plurality of frames based on the first probability for each of the plurality of frames; and

determining whether the ARS voice is included in each of the plurality of frames based on the second probability for each of the plurality of frames.

9. A computer program stored in a non-transitory computer readable storage medium, wherein when the computer program is executed by one or more processors to perform the following operations to detect a voice from audio data, the operations including:

obtaining audio data;

generating image data based on a spectrum of the obtained audio data;

extracting a feature on the image data by utilizing a first neural network model; and

extracting a time-series feature of the image data by utilizing a second neural network model sequentially connected with the first neural network model;

determining whether an automated response system (ARS) voice is included in each of a plurality of sections of the audio data, based on the feature on the image data and the time-series feature of the image data;

eliminating a section including the ARS voice, among the plurality of sections of the audio data; and

generating input data or training data for analysis of spoken voice based on remaining sections excluding the eliminated section, among the plurality of sections of the audio data.

10. A computing device, comprising:

a processor including at least one core;

a memory including executable program codes in the processor; and

a network unit for obtaining an audio file,

wherein the processor is configured to:

generate image data based on a spectrum of the obtained audio data;

extract time-series feature of the image data by utilizing a first neural network model; and

extract time-series feature of the image data by utilizing a second neural network model sequentially connected with the first neural network model;

determine whether an automated response system (ARS) voice is included in each of a plurality of sections of the audio data, based on the feature on the image data and the time-series feature of the image data;

eliminate a section including the ARS voice, among the plurality of sections of the audio data; and

generate input data or training data for analysis of spoken voice based on the remaining sections excluding the eliminated section, among the plurality of sections of the audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2023
From: CHOI, SUBONG; SHIN, DONGCHAN; LEE, JIHWA
To: ACTIONPOWER CORP.
Reel/Frame 064582/0342 →
Priority Claims (1)
KR 10-2022-0077482 · Jun 24, 2022 · national
Continuity (1)
Related Publication 20230419988A1 · Dec 28, 2023