IP Library › Granted Patent US 12,433,530
Granted Patent B2
US 12,433,530 · App. 18/004,934 · Granted Oct 7, 2025

Voice characteristic-based method and device for predicting alzheimer's disease

Inventors: Jun-Young Lee (Seoul, KR); Hyunwoong Ko (Seoul, KR)
Assignee: EMOCOG CO., LTD.
A61B5/4088A61B5/7275G10L25/66
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,433,530
App. No.
18/004,934
Granted
Oct 7, 2025
Kind
B2
Abstract

A method and device for predicting Alzheimer's disease based on voice characteristics are provided. The device for predicting Alzheimer's disease according to an embodiment includes: a voice input unit configured to generate a voice sample by recording a voice of a subject; a data input unit configured to receive demographic information of the subject; a voice characteristic extraction unit configured to extract voice characteristics from the generated voice sample; and a prediction model that is pre-trained to predict presence or absence of Alzheimer's disease in the subject, based on the voice characteristics and the demographic information.

Claims (82)

1. A device for predicting Alzheimer's disease, the device comprising:

a voice input unit configured to generate a voice sample by recording a voice of a subject;

a data input unit configured to receive demographic information of the subject;

a voice characteristic extraction unit configured to extract voice characteristics from the generated voice sample; and

a prediction model that is pre-trained to predict presence or absence of Alzheimer's disease in the subject, based on the voice characteristics and the demographic information,

wherein the prediction model comprises a multivariate logistic regression model that comprises at least one term determined by a combination of a plurality of first independent variables representing the voice characteristics, a plurality of second independent variables representing the demographic information, and a plurality of regression coefficients,

wherein the plurality of regression coefficients are trained to determine a decision boundary of the multivariate logistic regression model, and

wherein the multivariate logistic regression model is configured to output a dementia risk. probability value according to the plurality of first independent variables and the plurality of second independent variables using the decision boundary, and output state information that is determined as Alzheimer's disease or normal cognitive function based on the dementia risk probability value.

2. The device of claim 1 , wherein the demographic information comprises age, gender, and years of education of the subject.

3. The device of claim 1 , wherein the voice characteristic extraction unit is further configured to extract, as the voice characteristics, at least one of a fundamental frequency, a speech rate, a speech time, a speech length, a pause degree, a number of pauses, a pause interval length, a shimmer, a jitter, a formant, a harmonic-to-noise ratio, a loudness, a spectral centroid, Mel-frequency cepstral coefficients (MFCCs), an identity vector (i-vector), an articulation rate, a zero-crossing rate (ZCR), a voicing probability (VP), line spectral pairs (LSP), a period perturbation, an amplitude perturbation quotient (APQ), stiffness, energy, an intensity (volume), and an entropy of a voice.

4. The device of claim 3 , wherein the voice characteristic extraction unit comprises an artificial neural network model configured to perform preprocessing to select a human voice from the voice sample, and

the voice characteristic extraction unit is further configured to extract the voice characteristics from a preprocessed voice sample.

5. The device of claim 1 , wherein the prediction model further comprises at least one analysis model among a linear regression model, a machine learning model, and a neural network model.

6. The device of claim 5 , wherein

the multivariate logistic regression model is configured based on Equation 1:

log

⁡

(

p

⁡

(

X

)

1

-

p

⁡

(

X

)

)

=

β

0

+

β

1

⁢

X

1

+

β

2

⁢

X

2

+

β

3

⁢

X

3

+

…

+

β

p

⁢

X

p

[

Equation

⁢

1

]

wherein, X 1 to X p are independent variables that are input values input to the prediction model and correspond to p voice characteristics and pieces of the demographic information, respectively, β 1 to β p correspond to constant values that are regression coefficients of the independent variables, β 0 corresponds to an initial constant value, and p(X) corresponds to the dementia risk probability value.

7. A method of predicting Alzheimer's disease based on voice characteristics, the method comprising:

generating a voice sample by recording a voice of a subject;

receiving demographic information of the subject;

extracting voice characteristics from the generated voice sample; and

predicting presence or absence of Alzheimer's disease in the subject by inputting the voice characteristics and the demographic information into a pre-trained prediction model,

wherein the prediction model comprises a multivariate logistic regression model that comprises at least one term determined by a combination of a plurality of first independent variables representing the voice characteristics, a plurality of second independent variables representing the demographic information, and a plurality of regression coefficients,

wherein the plurality of regression coefficients are trained to determine a decision boundary of the multivariate logistic regression model, and

wherein the multivariate logistic regression model is configured to output a dementia risk probability value according to the plurality of first independent variables and the plurality of second independent variables using the decision boundary, and output state information that is determined as Alzheimer's disease or normal cognitive function based on the dementia risk probability value.

8. A computer-readable recording medium having recorded thereon computer-readable instructions, the computer-readable instructions, when executed by at least one processor, causing the at least one processor to:

generate a voice sample by recording a voice of a subject;

receive demographic information of the subject;

extract voice characteristics from the generated voice sample; and

predict presence or absence of Alzheimer's disease in the subject by inputting the voice characteristics and the demographic information into a pre-trained prediction model,

wherein the prediction model comprises a multivariate logistic regression model that comprises at least one term determined by a combination of a plurality of first independent variables representing the voice characteristics, a plurality of second independent variables representing the demographic information, and a plurality of regression coefficients,

wherein the plurality of regression coefficients are trained to determine a decision boundary of the multivariate logistic regression model, and

wherein the multivariate logistic regression model is configured to output a dementia risk probability value according to the plurality of first independent variables and the plurality of second independent variables using the decision boundary, and output state information that determined as Alzheimer's disease or normal cognitive function based on the dementia risk probability value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2023
From: LEE, JUN-YOUNG; KO, HYUNWOONG
To: EMOCOG CO., LTD.
Reel/Frame 063032/0001 →
Priority Claims (2)
KR 10-2020-0085449 · Jul 10, 2020 · national
KR 10-2021-0089014 · Jul 7, 2021 · national
Continuity (1)
Related Publication 20230233136A1 · Jul 27, 2023
References Cited (20)
US 20170053665A1 · Quatieri, Jr. · 2017 [cited by examiner]
US 20180322961A1 · Kim et al. · 2018 [cited by applicant]
US 20190043618A1 · Vaughan et al. · 2019 [cited by applicant]
US 20190311815A1 · Kim et al. · 2019 [cited by applicant]
US 20200027557A1 · Karow et al. · 2020 [cited by applicant]
JP 2007532219A · 2007 [cited by applicant]
JP 2013255786A · 2013 [cited by applicant]
JP 201746659A · 2017 [cited by applicant]
JP 2019084249A · 2019 [cited by applicant]
JP 202064051A · 2020 [cited by applicant]
KR 1020120070668A · 2012 [cited by applicant]
KR 1020140119486A · 2014 [cited by applicant]
KR 101881731B1 · 2018 [cited by applicant]
KR 1020190063275A · 2019 [cited by applicant]
KR 1020190081626A · 2019 [cited by applicant]
KR 1020200073156A · 2020 [cited by applicant]
WO 2018204934A1 · 2018 [cited by applicant]
WO 2019086555A1 · 2019 [cited by applicant]
WO 2019221252A1 · 2019 [cited by applicant]
International Search Report of PCT/KR2021/008710 dated Oct. 25, 2021 [PCT/ISA/210]. [cited by applicant]