IP Library › Granted Patent US 12,426,828
Granted Patent B2
US 12,426,828 · App. 17/644,147 · Granted Sep 30, 2025

Detection of cognitive impairment using speech feature distribution

Inventors: Kaoru Shinkawa (Hachioji, JP); Yasunori Yamada (Kawaguchi, JP); Masatomo Kobayashi (Tokyo, JP)
Assignee: International Business Machines Corporation
A61B5/4088A61B5/4803A61B5/742A61B5/7475G10L15/02G10L15/04G10L15/063G10L15/22G10L25/66G16H50/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,426,828
App. No.
17/644,147
Granted
Sep 30, 2025
Kind
B2
Abstract

A method, computer system, and a computer program product for speech feature distribution is provided. The present invention may include receiving two or more speech samples from a user. The present invention may include dividing the two or more speech samples into a a plurality of pause segments. The present invention may include determining a pause duration for each of the plurality of pause segments. The present invention may include determining a distribution of pause durations. The present invention may include determining a distance between the distribution of pause durations.

Claims (67)

1. A method for speech feature distribution, the method comprising:

receiving two or more speech samples from a user, wherein a first speech sample corresponds to a first task, wherein a second speech sample corresponds to a second task, and wherein the first task and the second task are each one of at least two or more tasks selected by a medical professional within a speech user interface, wherein each of the at least two or more tasks are retrieved from a knowledge corpus and are of different cognitive difficulty levels;

dividing at least the first speech sample and the second speech sample into a plurality of pause segments using a Voice Activity Detection (VAD) feature of one or more speech-to-text engines, wherein the plurality of pause segments are below a sound intensity level threshold set or adjusted by the medical professional in the speech user interface;

determining a first pause duration distribution from the plurality of pause segments of the first speech sample;

determining a second pause duration distribution from the plurality of pause segments of the second speech sample;

determining a distance between the first pause duration distribution and the second pause duration distribution; and

comparing the distance with a corresponding distance between pause duration distributions from a second user for a same at least two or more tasks, wherein the second user has a cognitive impairment.

2. The method of claim 1 , further comprising:

determining a probability in which the user is suffering from the cognitive impairment of the second user based on the comparing of the distance between each of the distributions of the pause durations of the two or more speech samples received for the user and the second user, wherein the distributions of the pause durations for the second user are retrieved from the knowledge corpus.

3. The method of claim 2 , further comprising:

displaying one or more recommendations to the medical professional within the speech user interface based on the probability in which the user is suffering from the cognitive impairment, wherein the one or more recommendations are based on an efficacy of treatments for the second user stored in the knowledge corpus.

4. The method of claim 1 , wherein selecting the first task and the second task from the knowledge corpus further comprises:

retrieving at least two lists of similar individuals who have performed a same task, based on data input received from the medical professional with respect to the user, the data input being entered by the user in the speech user interface;

displaying the at least two lists of similar individuals to the medical professional in the speech user interface;

displaying the first task and the second task to the medical professional in the speech user interface based on the similar individuals selected by the medical professional from the at least two lists; and

adjusting the first task and the second task based on a modification of the medical professional.

5. The method of claim 1 , further comprising:

determining a probability in which the user is suffering from at least one cognitive impairment of a plurality of cognitive impairments using one or more machine learning models, wherein the one or more machine learning models are trained according to a holdout or cross validation method and are based on a plurality of speech features and plurality of data stored in the knowledge corpus, wherein an input received by the one or more machine learning models is selected from the group consisting of the distance between the distribution of the pause durations for the user, a plurality of additional speech features determined from the two or more speech samples of the user, and a plurality of user data inputted by the medical professional.

6. The method of claim 5 , further comprising:

determining, using the one or more machine learning models, which of the plurality of cognitive impairments the user is suffering from;

determining, using the one or more machine learning models, a stage of the at least one cognitive impairment; and

providing, in the speech user interface, one or more recommendations to the user for the at least one cognitive impairment, wherein the one or more recommendations are determined using the one or more machine learning models utilizing as input at least, one or more of, which of the plurality of cognitive impairments the user is suffering from, the stage of the at least one cognitive impairment, and data stored in the knowledge corpus.

7. The method of claim 1 , wherein the distance between each of the distributions for each of the pause durations is determined utilizing one or more distribution techniques, wherein the one or more distribution techniques are selected based on the tasks selected by the medical professional, a highest area under a curve (AUC), and a lowest p value.

8. A computer system for speech feature distribution, comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method comprising:

receiving two or more speech samples from a user, wherein a first speech sample corresponds to a first task, wherein a second speech sample corresponds to a second task, and wherein the first task and the second task are each one of at least two or more tasks selected by a medical professional within a speech user interface, wherein each of the at least two or more tasks are retrieved from a knowledge corpus and are of different cognitive difficulty levels;

dividing at least the first speech sample and the second speech sample into a plurality of pause segments using a Voice Activity Detection (VAD) feature of one or more speech-to-text engines, wherein the plurality of pause segments are below a sound intensity level threshold set or adjusted by the medical professional in the speech user interface;

determining a first pause duration distribution from the plurality of pause segments of the first speech sample;

determining a second pause duration distribution from the plurality of pause segments of the second speech sample;

determining a distance between the first pause duration distribution and the second pause duration distribution; and

comparing the distance with a corresponding distance between pause duration distributions from a second user for a same at least two or more tasks, wherein the second user has a cognitive impairment.

9. The computer system of claim 8 , further comprising:

determining a probability in which the user is suffering from the cognitive impairment of the second user based on the comparing of the distance between each of the distributions of the pause durations of the two or more speech samples received for the user and the second user, wherein the distributions of the pause durations for the second user are retrieved from the knowledge corpus.

10. The computer system of claim 9 , further comprising:

displaying one or more recommendations to the medical professional within the speech user interface based on the probability in which the user is suffering from the cognitive impairment, wherein the one or more recommendations are based on an efficacy of treatments for the second user stored in the knowledge corpus.

11. The computer system of claim 8 , wherein selecting the first task and the second task from the knowledge corpus further comprises:

retrieving at least two lists of similar individuals who have performed a same task, based on data input received from the medical professional with respect to the user, the data input being entered by the user in the speech user interface;

displaying the at least two lists of similar individuals to the medical professional in the speech user interface;

displaying the first task and the second task to the medical professional in the speech user interface based on the similar individuals selected by the medical professional from the at least two lists; and

adjusting the first task and the second task based on a modification of the medical professional.

12. The computer system of claim 8 , further comprising:

determining a probability in which the user is suffering from at least one of a plurality of cognitive impairments using one or more machine learning models, wherein the one or more machine learning models are trained according to a holdout or cross validation method and are based on a plurality of speech features and plurality of data stored in the knowledge corpus, wherein an input received by the one or more machine learning models is selected from the group consisting of the distance between the distribution of the pause durations for the user, a plurality of additional speech features determined from the two or more speech samples of the user, and a plurality of user data inputted by the medical professional.

13. The computer system of claim 12 , further comprising:

determining, using the one or more machine learning models, which of the plurality of cognitive impairments the user is suffering from;

determining, using the one or more machine learning models, a stage of the at least one cognitive impairment; and

providing, in the speech user interface, one or more recommendations to the user for the at least one cognitive impairment, wherein the one or more recommendations are determined using the one or more machine learning models utilizing as input at least, one or more of, which of the plurality of cognitive impairments the user is suffering from, the stage of the at least one cognitive impairment, and data stored in the knowledge corpus.

14. The computer system of claim 8 , wherein the distance between each of the distributions for each of the pause durations is determined utilizing one or more distribution techniques, wherein the one or more distribution techniques are selected based on the tasks selected by the medical professional, a highest area under a curve (AUC), and a lowest p value.

15. A computer program product for speech feature distribution, comprising:

one or more non-transitory computer-readable storage media and program instructions stored on at least one of the one or more tangible storage media, the program instructions executable by a processor to cause the processor to perform a method comprising:

receiving two or more speech samples from a user, wherein a first speech sample corresponds to a first task, wherein a second speech sample corresponds to a second task, and wherein the first task and the second task are each one of at least two or more tasks selected by a medical professional within a speech user interface, wherein each of the at least two or more tasks are retrieved from a knowledge corpus and are of different cognitive difficulty levels;

dividing at least the first speech sample and the second speech sample into a plurality of pause segments using a Voice Activity Detection (VAD) feature of one or more speech-to-text engines, wherein the plurality of pause segments are below a sound intensity level threshold set or adjusted by the medical professional in the speech user interface;

determining a first pause duration distribution from the plurality of pause segments of the first speech sample;

determining a second pause duration distribution from the plurality of pause segments of the second speech sample;

determining a distance between the first pause duration distribution and the second pause duration distribution; and

comparing the distance with a corresponding distance between pause duration distributions from a second user for a same at least two or more tasks, wherein the second user has a cognitive impairment.

16. The computer program product of claim 15 , further comprising:

determining a probability in which the user is suffering from the cognitive impairment of the second user based on the comparing of the distance between each of the distributions of the pause durations of the two or more speech samples received for the user and the second user, wherein the distributions of the pause durations for the second user are retrieved from the knowledge corpus.

17. The computer program product of claim 16 , further comprising:

displaying one or more recommendations to the medical professional within the speech user interface based on the probability in which the user is suffering from the cognitive impairment, wherein the one or more recommendations are based on an efficacy of treatments for the second user stored in the knowledge corpus.

18. The computer program product of claim 15 , wherein selecting the first task and the second task from the knowledge corpus further comprises:

retrieving at least two lists of similar individuals who have performed a same task, based on data input received from the medical professional with respect to the user, the data input being entered by the user in the speech user interface;

displaying the at least two lists of similar individuals to the medical professional in the speech user interface;

displaying the first task and the second task to the medical professional in the speech user interface based on the similar individuals selected by the medical professional from the at least two lists; and

adjusting the first task and the second task based on a modification of the medical professional.

19. The computer program product of claim 15 , further comprising:

determining a probability in which the user is suffering from at least one of a plurality of cognitive impairments using one or more machine learning models, wherein the one or more machine learning models are trained according to a holdout or cross validation method and are based on a plurality of speech features and plurality of data stored in the knowledge corpus, wherein an input received by the one or more machine learning models is selected from the group consisting of the distance between the distribution of the pause durations for the user, a plurality of additional speech features determined from the two or more speech samples of the user, and a plurality of user data inputted by the medical professional.

20. The computer program product of claim 15 , wherein the distance between each of the distributions for each of the pause durations is determined utilizing one or more distribution techniques, wherein the one or more distribution techniques are selected based on the tasks selected by the medical professional, a highest area under a curve (AUC), and a lowest p value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2021
From: SHINKAWA, KAORU; YAMADA, YASUNORI; KOBAYASHI, MASATOMO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058383/0492 →
Continuity (1)
Related Publication 20230181093A1 · Jun 15, 2023
References Cited (20)
US 20190304484A1 · Shinkawa · 2019 [cited by applicant]
US 20200229752A1 · Sumi · 2020 [cited by examiner]
US 20200327882A1 · Vairavan · 2020 [cited by applicant]
US 20220039741A1 · Gosztolya · 2022 [cited by examiner]
US 20220369976A1 · Abbas · 2022 [cited by examiner]
US 20230233136A1 · Lee · 2023 [cited by examiner]
JP 6804779B2 · 2020 [cited by applicant]
WO 2012045774A1 · 2012 [cited by applicant]
WO 2020128542A1 · 2020 [cited by applicant]
Rohanian, Morteza, Julian Hough, and Matthew Purver. “Alzheimer's dementia recognition using acoustic, lexical, disfluency and speech pause features robust to noisy inputs.” arXiv preprint arXiv:2106.15684 (Jun. 2021). … [cited by examiner]
Yamada, Yasunori, et al. “Tablet-based automatic assessment for early detection of Alzheimer's disease using speech responses to daily life questions.” Frontiers in Digital Health 3 (Mar. 2021): 653904. (Year: 2021). [cited by examiner]
Kobayashi, Masatomo, et al. “Effects of age-related cognitive decline on elderly user interactions with voice-based dialogue systems.” Human-Computer Interaction—INTERACT 2019: 17th IFIP TC 13 International Conference, … [cited by examiner]
Yamada, “Daily chats with AI could help spot early signs of Alzheimer's”, [online] https://research.ibm.com/blog/ai-chats-spot-alzheimers; published in 2021. (Year: 2021). [cited by examiner]
Yamada, Yasunori, et al. “A mobile application using automatic speech analysis for classifying Alzheimer's disease and mild cognitive impairment.” Computer Speech & Language 81 (2023): 101514. (Year: 2023). [cited by examiner]
Hall, “Using tablet-based assessment to characterize speech for individuals with dementia and mild cognitive impairment: preliminary results” (Year: 2019). [cited by examiner]
Hall, et al., “Using Tablet-Based Assessment to Characterize Speech for Individuals with Dementia and Mild Cognitive Impairment: Preliminary Results,” AMIA Jt Summits Transl Sci Proceedings, 2019, pp. 34-43, Retrieved f… [cited by applicant]
König, et al., “Automatic speech analysis for the assessment of patients with predementia and Alzheimer's disease,” Alzheimer's & Dementia: Diagnosis, Assessment & Disease Monitoring, Mar. 2015, pp. 112-124, doi: 10.101… [cited by applicant]
Mell, et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Roark, et al., “Spoken Language Derived Measures for Detecting Mild Cognitive Impairment,” IEEE Transactions on Audio, Speech, and Language Processing, Feb. 7, 2011, pp. 2081-2090, vol. 19, Issue 7, DOI: 10.1109/TASL.20… [cited by applicant]
Toth, et al., Automatic Detection of Mild Cognitive Impairment from Spontaneous Speech using ASR, Interspeech [conference paper], Sep. 2015, 6 pages, DOI:10.21437/Interspeech.2015-568, Retrieved from the Internet: <URL:… [cited by applicant]