Methods and systems for using machine learning models to detect anomalies and predict trends in speech data
A non-transitory processor-readable medium stores instructions that, when executed by a processor, cause the processor to receive audio data and provide the audio data as input to a first machine learning model to generate (1) transcription data, (2) timestamp data associated with the transcription data, and (3) a confidence metric associated with the transcription data. A speaking metric is calculated based on the transcription data and the timestamp data. The speaking metric and the confidence metric are provided as input to a second machine learning model to predict a listener effort metric.
1 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:
receive audio data associated with a person;
provide the audio data as input to a first machine learning model to generate (1) transcription data, (2) timestamp data associated with the transcription data, and (3) a confidence metric that represents an accuracy of the transcription data relative to the audio data;
calculate a speaking metric based on the transcription data and the timestamp data; and
predict, by providing the speaking metric and the confidence metric as input to a second machine learning model, a listener effort metric that is from a perspective other than that of the person and that indicates a level of dysarthria associated with the person.
2 . The non-transitory, processor-readable medium of claim 1 , further storing instructions that cause the processor to provide the listener effort metric as input to a third machine learning model to predict a progression rate associated with the dysarthria of the person.
3 . The non-transitory, processor-readable medium of claim 2 , wherein the progression rate is associated with a change in the listener effort metric during a predefined time period.
4 . The non-transitory, processor-readable medium of claim 2 , wherein the progression rate is associated with a disease quantification for at least one of amyotrophic lateral sclerosis (ALS), Alzheimer's disease, Huntington's disease, or Parkinson's disease.
5 . The non-transitory, processor-readable medium of claim 1 , wherein:
the audio data includes speech data; and
the first machine learning model is a speech recognition model.
6 . The non-transitory, processor-readable medium of claim 1 , wherein the instructions to cause the processor to predict include instructions to cause the processor to predict the listener effort metric by providing as input to the second machine learning model, at least one of a fundamental frequency, a jitter metric, a shimmer metric, a formants variation metric, a Wiener entropy metric, or a Cepstral peak prominence metric, associated with the audio data.
7 . The non-transitory, processor-readable medium of claim 1 , wherein the second machine learning model is at least one of a Lasso regression model, a neural network, a random forest model, or a nearest neighbor model.
8 . The non-transitory, processor-readable medium of claim 1 , wherein the second machine learning model is trained based on a perceived listener effort metric received via a user interface.
9 . The non-transitory, processor-readable medium of claim 1 , wherein the confidence metric is a collective confidence metric, the instructions to cause the processor to generate the collective confidence metric include instructions to cause the processor to generate the collective confidence metric based on a confidence metric associated with each word from a plurality of words included in the transcription data.
10 . The non-transitory, processor-readable medium of claim 1 , wherein the listener effort metric is from a plurality of listener effort metrics associated with the person over a period of time, the non-transitory, processor-readable medium further storing instructions that cause the processor to:
calculate, using the plurality of listener effort metrics, a progression trend; and
send a signal to display to a user a representation of the progression trend.
11 . The non-transitory, processor-readable medium of claim 1 , further storing instructions that cause the processor to send a signal to display to a user the listener effort metric.
12 . A method, comprising:
receiving speech data that represents a set of words spoken by a person;
providing the speech data as input to a first machine learning model to produce (1) transcription data that represents the set of words, (2) a set of confidence metrics for the set of words, and (3) timestamp data representing a speech rate;
providing the speech data, the transcription data, the set of confidence metrics, and the timestamp data as input to a second machine learning model to produce a listener effort metric that quantifies an effort to listen to the person; and
providing the listener effort metric as input to a third machine learning model to forecast a progression rate of a condition for the person.
13 . The method of claim 12 , wherein:
the first machine learning model includes an encoder-decoder model.
14 . The method of claim 12 , wherein:
the second machine learning model is a Least Absolute Shrinkage and Selection Operator (Lasso) regression model.
15 . The method of claim 12 , wherein:
the third machine learning model includes a mixture of Gaussian processes (MoGP) model.
16 . The method of claim 12 , further comprising:
providing the speech data as input to a fourth machine learning model to produce at least one of a dysarthria severity metric, a voice strain metric, a consistency metric, an intelligibility metric, an articulatory precision metric, a dysphonia severity metric, a hypernasality metric, a breath support metric, or a prosody metric.
17 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:
receive speech data for a person who has a condition;
perform a frequency decomposition on the speech data to produce spectrogram data;
provide the spectrogram data as input to a first machine learning model to produce acoustical feature data; and
provide the acoustical feature data as input to a second machine learning model to produce a listener effort metric (1) from a perspective other than that of the person and (2) that quantifies the condition of the person.
18 . The non-transitory, processor-readable medium of claim 17 , wherein:
the first machine learning model includes a convolutional neural network (CNN).
19 . The non-transitory, processor-readable medium of claim 17 , wherein:
the acoustical feature data represents at least one of a sound envelope metric, a fundamental frequency, a jitter metric, a shimmer metric, a pitch metric, a formants metric, a formants variation metric, a harmonic-to-noise ratio (HNR), a Wiener entropy metric, or a Cepstral peak prominence (CPP) metric.
20 . The non-transitory, processor-readable medium of claim 17 , further storing instructions to cause the processor to:
provide the listener effort metric as in input to a mixture of Gaussian processes (MoGP) model to predict a progression rate associated with the condition.