SYSTEM AND METHOD FOR DATA AUGMENTATION OF FEATURE-BASED VOICE DATA
A method, computer program product, and computing system for receiving feature-based voice data associated with a first acoustic domain. One or more audio feature-based augmentations may be performed on at least a portion of the feature-based voice data. Performing the one or more audio feature-based augmentations may include adding one or more audio features to the at least a portion of the feature-based voice data and/or removing one or more audio features from the at least a portion of the feature-based voice data.
1 . A computer-implemented method, executed on a computing device, comprising:
receiving feature-based voice data associated with a first acoustic domain; and
performing one or more audio feature-based augmentations on at least a portion of the feature-based voice data including one or more of:
adding one or more audio features to the at least a portion of the feature-based voice data; and
removing one or more audio features from the at least a portion of the feature-based voice data.
2 . The computer-implemented method of claim 1 , wherein receiving the feature-based voice data associated with the first acoustic domain includes converting an audio signal, generated in the first acoustic domain, from the time domain to the feature domain, thus defining the feature-based voice data associated with the first acoustic domain.
3 . The computer-implemented method of claim 1 , further comprising:
receiving a selection of a target acoustic domain.
4 . The computer-implemented method of claim 3 , wherein performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data includes performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain.
5 . The computer-implemented method of claim 4 , wherein adding the one or more audio features to the at least a portion of the feature-based voice data includes adding one or more noise features associated with the target domain to the at least a portion of the feature-based voice data.
6 . The computer-implemented method of claim 4 , wherein removing the one or more audio features to the at least a portion of the feature-based voice data includes removing one or more noise features associated with the first acoustic domain from the at least a portion of the feature-based voice data.
7 . The computer-implemented method of claim 1 , wherein performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data includes performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, a target signal-to-noise ratio (SNR).
8 . The computer-implemented method of claim 1 , further comprising:
training a machine learning model with a plurality of audio features associated with the target acoustic domain.
9 . The computer-implemented method of claim 8 , wherein performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data includes performing the one or more audio feature-based augmentations to the at least a portion of the feature-based voice data using the trained machine learning model configured to model the plurality of audio features associated with the target acoustic domain.
10 . A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:
receiving feature-based voice data associated with a first acoustic domain; and
performing one or more audio feature-based augmentations on at least a portion of the feature-based voice data including one or more of:
adding one or more audio features to the at least a portion of the feature-based voice data; and
removing one or more audio features from the at least a portion of the feature-based voice data.
11 . The computer program product of claim 10 , wherein receiving the feature-based voice data associated with the first acoustic domain includes converting an audio signal, generated in the first acoustic domain, from the time domain to the feature domain, thus defining the feature-based voice data associated with the first acoustic domain.
12 . The computer program product of claim 10 , wherein the operations further comprise:
receiving a selection of a target acoustic domain.
13 . The computer program product of claim 12 , wherein performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data includes performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain.
14 . The computer program product of claim 13 , wherein adding the one or more audio features to the at least a portion of the feature-based voice data includes adding one or more noise features associated with the target domain to the at least a portion of the feature-based voice data.
15 . The computer program product of claim 13 , wherein removing the one or more audio features to the at least a portion of the feature-based voice data includes removing one or more noise features associated with the first acoustic domain from the at least a portion of the feature-based voice data.
16 . The computer program product of claim 10 , wherein performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data includes performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, a target signal-to-noise ratio (SNR).
17 . The computer program product of claim 12 , further comprising:
training a machine learning model with a plurality of audio features associated with the target acoustic domain.
18 . The computer program product of claim 17 , wherein performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data includes performing the one or more audio feature-based augmentations to the at least a portion of the feature-based voice data using the trained machine learning model configured to model the plurality of audio features associated with the target acoustic domain.
19 . A computing system comprising:
a memory; and
a processor configured to receive feature-based voice data associated with a first acoustic domain, and wherein the processor is further configured to perform one or more audio feature-based augmentations on at least a portion of the feature-based voice data including one or more of adding one or more audio features to the at least a portion of the feature-based voice data; and removing one or more audio features from the at least a portion of the feature-based voice data.
20 . The computing system of claim 19 , wherein performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data includes performing the one or more audio feature-based augmentations on at least a portion of the feature-based voice data based upon, at least in part, a target signal-to-noise ratio (SNR).