System and method for data augmentation of feature-based voice data
A method, computer program product, and computing system for receiving feature-based voice data associated with a first acoustic domain. One or more rate-based augmentations may be performed on at least a portion of the feature-based voice data, thus defining rate-based augmented feature-based voice data.
1. A computer-implemented method, executed on a computing device, comprising:
extracting acoustic metadata from each portion of a signal before converting the signal from a first acoustic domain to a feature domain, wherein the signal is divided into a plurality of portions;
receiving feature-based voice data associated with the first acoustic domain, wherein the feature-based voice data is converted from the signal in the first acoustic domain to the feature domain;
determining a phoneme-rate associated with the first acoustic domain based upon, at least in part, the acoustic metadata;
receiving a selection of a target acoustic domain;
determining a phoneme-rate associated with the target acoustic domain; and
performing one or more rate-based augmentations on at least a portion of the feature-based voice data converted from the signal in the first acoustic domain to the feature domain, thus defining rate-based augmented feature-based voice data, wherein the one or more rate-based augmentation is a change to a speaking rate within the at least the portion of the feature-based voice data, wherein performing the one or more rate-based augmentations to the at least a portion of the feature-based voice data includes adjusting the phoneme-rate associated with the first acoustic domain toward the phoneme-rate associated with the target acoustic domain.
2. The computer-implemented method of claim 1 , wherein adjusting the phoneme-rate associated with the first acoustic domain toward the phoneme-rate associated with the target acoustic domain includes decreasing a phoneme-rate of at least a portion of the feature-based voice data.
3. The computer-implemented method of claim 2 , wherein decreasing a phoneme-rate of at least a portion of the feature-based voice data includes adding one or more frames to the feature-based voice data.
4. The computer-implemented method of claim 1 , wherein adjusting the phoneme-rate associated with the first acoustic domain toward the phoneme-rate associated with the target acoustic domain includes increasing a phoneme-rate of at least a portion of the feature-based voice data.
5. The computer-implemented method of claim 4 , wherein increasing a phoneme-rate of at least a portion of the feature-based voice data includes dropping one or more frames from the feature-based voice data.
6. The computer-implemented method of claim 1 , further comprising:
training a machine learning model to one or more of add at least one frame to the feature-based voice data and remove at least one frame from the feature-based voice data based upon, at least in part, the target acoustic domain.
7. The computer-implemented method of claim 6 , wherein performing the one or more rate-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain includes performing the one or more rate-based augmentations to the at least a portion of the feature-based voice data using the trained machine learning model configured to one or more of add at least one frame to the feature-based voice data and remove at least one frame from the feature-based voice data based upon, at least in part, the target acoustic domain.
8. The computer-implemented method of claim 7 , wherein the trained machine learning model is configured to perform smoothing of the feature-based voice data when one or more of adding at least one frame to the feature-based voice data and removing at least one frame from the feature-based voice data.
9. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:
extracting acoustic metadata from each portion of a signal before converting the signal from a first acoustic domain to a feature domain, wherein the signal is divided into a plurality of portions;
receiving feature-based voice data associated with the first acoustic domain, wherein the feature-based voice data is converted from the signal in the first acoustic domain to the feature domain;
determining a phoneme-rate associated with the first acoustic domain based upon, at least in part, the acoustic metadata;
receiving a selection of a target acoustic domain;
determining a phoneme-rate associated with the target acoustic domain; and
performing one or more rate-based augmentations on at least a portion of the feature-based voice data converted from the signal in the first acoustic domain to the feature domain, thus defining rate-based augmented feature-based voice data, wherein the one or more rate-based augmentation is a change to a speaking rate within the at least the portion of the feature-based voice data, wherein performing the one or more rate-based augmentations to the at least a portion of the feature-based voice data includes adjusting the phoneme-rate associated with the first acoustic domain toward the phoneme-rate associated with the target acoustic domain.
10. The computer program product of claim 9 , wherein adjusting the phoneme-rate associated with the first acoustic domain toward the phoneme-rate associated with the target acoustic domain includes decreasing a phoneme-rate of at least a portion of the feature-based voice data.
11. The computer program product of claim 10 , wherein decreasing a phoneme-rate of at least a portion of the feature-based voice data includes adding one or more frames to the feature-based voice data.
12. The computer program product of claim 9 , wherein adjusting the phoneme-rate associated with the first acoustic domain toward the phoneme-rate associated with the target acoustic domain includes increasing a phoneme-rate of at least a portion of the feature-based voice data.
13. The computer program product of claim 12 , wherein increasing a phoneme-rate of at least a portion of the feature-based voice data includes dropping one or more frames from the feature-based voice data.
14. The computer program product of claim 9 , further comprising: training a machine learning model to one or more of add at least one frame to the feature-based voice data and drop at least one frame from the feature-based voice data based upon, at least in part, the target acoustic domain.
15. The computer program product of claim 14 , wherein performing the one or more rate-based augmentations to the at least a portion of the feature-based voice data based upon, at least in part, the target acoustic domain includes performing the one or more rate-based augmentations to the at least a portion of the feature-based voice data using the trained machine learning model configured to one or more of add at least one frame to the feature-based voice data and drop at least one frame from the feature-based voice data based upon, at least in part, the target acoustic domain.
16. The computer program product of claim 15 , wherein the trained machine learning model is configured to perform smoothing of the feature-based voice data when one or more of adding at least one frame to the feature-based voice data and dropping at least one frame from the feature-based voice data.
17. A computing system comprising:
a memory; and
a processor configured to extract acoustic metadata from each portion of a signal before converting the signal from a first acoustic domain to a feature domain, wherein the signal is divided into a plurality of portions, to receive feature-based voice data associated with the first acoustic domain, wherein the feature-based voice data is converted from the signal in the first acoustic domain to the feature domain, to determine a phoneme-rate associated with the first acoustic domain based upon, at least in part, the acoustic metadata, to receive a selection of a target acoustic domain, to determine a phoneme-rate associated with the target acoustic domain, and to perform one or more rate-based augmentations on at least a portion of the feature-based voice data converted from the signal in the first acoustic domain to the feature domain, thus defining rate-based augmented feature-based voice data, wherein the one or more rate-based augmentation is a change to a speaking rate within the at least the portion of the feature-based voice data, wherein performing the one or more rate-based augmentations to the at least a portion of the feature-based voice data includes adjusting the phoneme-rate associated with the first acoustic domain toward the phoneme-rate associated with the target acoustic domain.