IP Library Granted Patent US 12682916
Granted Patent B1
US 12682916 · App. 19/170,611 · Granted Jul 14, 2026

Methods and systems for using machine learning models to detect anomalies and predict trends in speech data

Inventors: Indu Navar Bingham (Los Altos, CA); Julián Peller (Buenos Aires, AR); Esteban Gabriel Roitberg (Provincia de Buenos Aires, AR); Ernest Samuel Fraenkel (Newton, MA)
Assignee: Peter Cohen Foundation
G10L25/48G10L15/26G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682916
App. No.
19/170,611
Granted
Jul 14, 2026
Kind
B1
Abstract

A non-transitory processor-readable medium stores instructions that, when executed by a processor, cause the processor to receive audio data and provide the audio data as input to a first machine learning model to generate (1) transcription data, (2) timestamp data associated with the transcription data, and (3) a confidence metric associated with the transcription data. A speaking metric is calculated based on the transcription data and the timestamp data. The speaking metric and the confidence metric are provided as input to a second machine learning model to predict a listener effort metric.

Claims (43)

1 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:

receive audio data associated with a person;

provide the audio data as input to a first machine learning model to generate (1) transcription data, (2) timestamp data associated with the transcription data, and (3) a confidence metric that represents an accuracy of the transcription data relative to the audio data;

calculate a speaking metric based on the transcription data and the timestamp data; and

predict, by providing the speaking metric and the confidence metric as input to a second machine learning model, a listener effort metric that is from a perspective other than that of the person and that indicates a level of dysarthria associated with the person.

2 . The non-transitory, processor-readable medium of claim 1 , further storing instructions that cause the processor to provide the listener effort metric as input to a third machine learning model to predict a progression rate associated with the dysarthria of the person.

3 . The non-transitory, processor-readable medium of claim 2 , wherein the progression rate is associated with a change in the listener effort metric during a predefined time period.

4 . The non-transitory, processor-readable medium of claim 2 , wherein the progression rate is associated with a disease quantification for at least one of amyotrophic lateral sclerosis (ALS), Alzheimer's disease, Huntington's disease, or Parkinson's disease.

5 . The non-transitory, processor-readable medium of claim 1 , wherein:

the audio data includes speech data; and

the first machine learning model is a speech recognition model.

6 . The non-transitory, processor-readable medium of claim 1 , wherein the instructions to cause the processor to predict include instructions to cause the processor to predict the listener effort metric by providing as input to the second machine learning model, at least one of a fundamental frequency, a jitter metric, a shimmer metric, a formants variation metric, a Wiener entropy metric, or a Cepstral peak prominence metric, associated with the audio data.

7 . The non-transitory, processor-readable medium of claim 1 , wherein the second machine learning model is at least one of a Lasso regression model, a neural network, a random forest model, or a nearest neighbor model.

8 . The non-transitory, processor-readable medium of claim 1 , wherein the second machine learning model is trained based on a perceived listener effort metric received via a user interface.

9 . The non-transitory, processor-readable medium of claim 1 , wherein the confidence metric is a collective confidence metric, the instructions to cause the processor to generate the collective confidence metric include instructions to cause the processor to generate the collective confidence metric based on a confidence metric associated with each word from a plurality of words included in the transcription data.

10 . The non-transitory, processor-readable medium of claim 1 , wherein the listener effort metric is from a plurality of listener effort metrics associated with the person over a period of time, the non-transitory, processor-readable medium further storing instructions that cause the processor to:

calculate, using the plurality of listener effort metrics, a progression trend; and

send a signal to display to a user a representation of the progression trend.

11 . The non-transitory, processor-readable medium of claim 1 , further storing instructions that cause the processor to send a signal to display to a user the listener effort metric.

12 . A method, comprising:

receiving speech data that represents a set of words spoken by a person;

providing the speech data as input to a first machine learning model to produce (1) transcription data that represents the set of words, (2) a set of confidence metrics for the set of words, and (3) timestamp data representing a speech rate;

providing the speech data, the transcription data, the set of confidence metrics, and the timestamp data as input to a second machine learning model to produce a listener effort metric that quantifies an effort to listen to the person; and

providing the listener effort metric as input to a third machine learning model to forecast a progression rate of a condition for the person.

13 . The method of claim 12 , wherein:

the first machine learning model includes an encoder-decoder model.

14 . The method of claim 12 , wherein:

the second machine learning model is a Least Absolute Shrinkage and Selection Operator (Lasso) regression model.

15 . The method of claim 12 , wherein:

the third machine learning model includes a mixture of Gaussian processes (MoGP) model.

16 . The method of claim 12 , further comprising:

providing the speech data as input to a fourth machine learning model to produce at least one of a dysarthria severity metric, a voice strain metric, a consistency metric, an intelligibility metric, an articulatory precision metric, a dysphonia severity metric, a hypernasality metric, a breath support metric, or a prosody metric.

17 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:

receive speech data for a person who has a condition;

perform a frequency decomposition on the speech data to produce spectrogram data;

provide the spectrogram data as input to a first machine learning model to produce acoustical feature data; and

provide the acoustical feature data as input to a second machine learning model to produce a listener effort metric (1) from a perspective other than that of the person and (2) that quantifies the condition of the person.

18 . The non-transitory, processor-readable medium of claim 17 , wherein:

the first machine learning model includes a convolutional neural network (CNN).

19 . The non-transitory, processor-readable medium of claim 17 , wherein:

the acoustical feature data represents at least one of a sound envelope metric, a fundamental frequency, a jitter metric, a shimmer metric, a pitch metric, a formants metric, a formants variation metric, a harmonic-to-noise ratio (HNR), a Wiener entropy metric, or a Cepstral peak prominence (CPP) metric.

20 . The non-transitory, processor-readable medium of claim 17 , further storing instructions to cause the processor to:

provide the listener effort metric as in input to a mixture of Gaussian processes (MoGP) model to predict a progression rate associated with the condition.