IP Library › Granted Patent US 12,300,227
Granted Patent B2
US 12,300,227 · App. 17/724,320 · Granted May 13, 2025

Customizing computer generated dialog for different pathologies

Inventors: Jackson Liscombe (New Marlborough, MA); Hardik Kothare (Burlingame, CA); Doug Habberstad (Savannah, GA); Andrew Cornish (Gore, NZ); Oliver Roesler (Weyhe, DE); Michael Neumann (Waiblingen, DE); David Pautler (San Francisco, CA); David Suendermann-Oeft (San Francisco, CA); Vikram Ramanarayanan (San Francisco, CA)
Assignee: Modality.AI
G10L15/22G10L25/66G10L25/84G10L25/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,300,227
App. No.
17/724,320
Granted
May 13, 2025
Kind
B2
Abstract

A computer-generated dialog session is customized for a user having a pathology characterized at least in part by a speech pathology. The user's speech is analyzed for spans of speech in which the starts and ends of the spans satisfy predetermined thresholds of time. Customization occurs by altering at least one of the following configurable parameters: (a) a threshold minimum signal strength of speech (dB) to consider as the start of the span of speech; (b) an adjustment factor by which signal strengths of background noise increases between consecutive spans of speech; (c) a threshold between signal strength during the span of speech and signal strength during the span of non-speech; (d) a start speech time threshold; and (e) an end speech time threshold.

Claims (22)

1. A method of customizing a computer-generated dialog session for a user having a speech pathology, comprising:

eliciting speech from a user and corresponding facial movements from a user from a plurality of tasks during the computer-generated dialog session;

identifying, by a computer device, (a) a span of speech and (b) a span of non-speech in an audio stream of the user's speech, wherein the identification of the span of speech and span of non-speech comprises:

calculating, by the computing device, a value of a frame within the stream based on the logarithm of the root mean squared of the energy of the frame;

comparing, by the computing device, the value to a threshold;

marking, by the computing device, the frame as one of:

speech if the value meets the threshold; and

non-speech if the value is below the threshold;

altering, by the computing device, a plurality of the following configurable parameters: (a) a threshold minimum signal strength of the user's speech (dB) to consider as the start of the span of the user's speech; (b) an adjustment factor by which signal strengths of background noise increases between consecutive spans of the user's speech; (c) a threshold between signal strength during the span of speech and a signal strength during the span of non-speech; and

presenting, by the computing device, the customized dialog session by applying one of the plurality of altered parameters;

wherein the plurality of tasks comprise of: an open-ended question, a sustained vowel phonation, oral diadochokinesis alternating motion rate or a repetition of a pre-selected syllable, a speech intelligibility test sentence, a spontaneous speech while describing a pre-selected picture, or any combination thereof.

2. The method of claim 1 , further comprising

identifying, by the computing device, a start of speech in the audio stream as a function of the span of speech continuing for a first threshold period of time;

identifying, by the computing device, an end of speech in the audio stream as a function of the span of non-speech continuing for a second threshold period of time.

3. The method of claim 1 , further comprising beginning, by the computing device, the dialog session with a microphone check for speech and background noise.

4. The method of claim 1 , further identifying, by the computing device, a user as having a particular speech pathology, and using the identification to set a plurality of the configurable parameters.

5. The method of claim 4 , further comprising using, by the computing device, multiple questionnaires to identify progression of the particular speech pathology in the user.

6. The method of claim 1 , further comprising customizing, by the computing device, the configurable parameters as a function of a plurality of the following linguistic features: prosody, voice quality, articulation, acoustics, respiration, and cognitive/mental/emotional state.

7. The method of claim 1 , wherein the step of identifying the spans of speech and non-speech further comprises iterative re-estimation of speech sounds and background noise.

8. The method of claim 1 , wherein the step of identifying the spans of speech and non-speech further comprises determining, by the computing device, spans of speech to be those in which an average signal level of speech sounds (dB) exceeds an average signal level of the background noise (dB) by a threshold amount.

9. The method of claim 1 , further comprising calculating, by the computing device, a weighted penalty of proportions of false positive and false negative times, when compared to a hand annotation of actual speech in the audio stream.

10. The method of claim 1 , further comprising using a questionnaire to ascertain scores for at least three different domains affected by the speech pathology, bulbar, limb, or respiratory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2022
From: LISCOMBE, JACKSON; KOTHARE, HARDIK; HABBERSTAD, DOUG; CORNISH, ANDREW; ROESLER, OLIVER; NEUMANN, MICHAEL; PAUTLER, DAVID; SUENDERMANN-OEFT, DAVID; RAMANARAYANAN, VIKRAM
To: MODALITY.AI
Reel/Frame 059932/0358 →
Continuity (2)
Provisional Application 63176626 · Apr 19, 2021
Related Publication 20220335939A1 · Oct 20, 2022
References Cited (56)
US 5357427A · Langen et al. · 1994 [cited by applicant]
US 9311932B2 · Carter · 2016 [cited by examiner]
US 10592733B1 · Ramanarayanan et al. · 2020 [cited by applicant]
US 20070033042A1 · Marcheret · 2007 [cited by examiner]
US 20120144336A1 · Pinter et al. · 2012 [cited by applicant]
US 20130211832A1 · Talwar · 2013 [cited by examiner]
US 20140278386A1 · Konchitsky · 2014 [cited by examiner]
US 20150051906A1 · Dickins · 2015 [cited by examiner]
US 20150058013A1 · Pakhomov · 2015 [cited by examiner]
US 20150126888A1 · Patel et al. · 2015 [cited by applicant]
US 20150356982A1 · Chesney · 2015 [cited by examiner]
US 20180090127A1 · Hofer · 2018 [cited by examiner]
US 20180214061A1 · Knoth et al. · 2018 [cited by applicant]
US 20180310866A1 · Wrenn · 2018 [cited by applicant]
US 20190074028A1 · Howard · 2019 [cited by applicant]
US 20190198043A1 · Crespi · 2019 [cited by examiner]
US 20190325898A1 · O'Hart Kinney · 2019 [cited by examiner]
US 20190348065A1 · Talwar · 2019 [cited by examiner]
US 20190378537A1 · Li · 2019 [cited by examiner]
US 20190385711A1 · Shriberg et al. · 2019 [cited by applicant]
US 20200296513A1 · Littrell · 2020 [cited by examiner]
US 20200349938A1 · Hwang et al. · 2020 [cited by applicant]
US 20200365275A1 · Barnett et al. · 2020 [cited by applicant]
US 20210098110A1 · Periyasamy et al. · 2021 [cited by applicant]
US 20210118329A1 · Medan · 2021 [cited by examiner]
US 20210121124A1 · Shenhar · 2021 [cited by examiner]
US 20210158834A1 · Medan · 2021 [cited by examiner]
US 20220148570A1 · Weissberg · 2022 [cited by examiner]
US 20230018524A1 · Ramanarayanan · 2023 [cited by examiner]
US 20230386504A1 · Lee · 2023 [cited by examiner]
WO 2016154139 · 2016 [cited by applicant]
Cedarbaum, et al. “The ALSFRS-R: a revised ALS functional rating scale that incorporates assessments of respiratory function,” J.M. Cedarbaum et al. / Journal of the Neurological Sciences 169 (1999) 13-21. 5 pages. [cited by applicant]
Bombaci, et al. “Telemedicine for management of patients with amyotrophic lateral sclerosis through COVID-19 tail,” Neurological Sciences (2021) 42:9-13. 5 pages. [cited by applicant]
Victoria Young MHSc & Alex Mihailidis PhD (2010) Difficulties in Automatic Speech Recognition of Dysarthric Speakers and Implications for Speech-Based Applications Used by the Elderly: A Literature Review, Assistive Tec… [cited by applicant]
Stegmann, et al. “Earyl detection and tracking of bulbar changes in ALS via frequent and remote speech analysis,” npj Digital Medicine (2020) 3:132. 5 pages. [cited by applicant]
Suendermann-Oeft, et al. “NEMSI: A Multimodal Dialog System for Screening of Neurological or Mental Conditions,” ACM International Conference on Intelligent Virtual Agents (IVA '19), Jul. 2-5, 2019. 3 pages. [cited by applicant]
Green, et al. “Algorithmic Estimation of Pauses in Extended Speech Samples of Dysarthric and Typical Speech,” J Med Speech Lang Pathol. Dec. 2004 ; 12(4): 149-154. 9 pages. [cited by applicant]
Byers, et al. “2017 Pilot Open Speech Analytic Technologies Evaluation,” NISTIR. 27 pages. [cited by applicant]
Kodrasi, et al. “Spectro-Temporal Sparsity Characterization for Dysarthric Speech Detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, 2020. 13 pages. [cited by applicant]
Janbakhshi, et al. “Subspace-Based Learning for Automatic Dysarthric Speech Detection,” IEEE Signal Processing Letters, vol. 28, 2021. 5 pages. [cited by applicant]
Lee, et al. “Vowel-Specific Intelligibility and Acoustic Patterns in Individuals With Dysarthria Secondary to Amyotrophic Lateral Sclerosis,” Journal of Speech, Language, and 34 Hearing Research ⋅ vol. 62 ⋅ 34-59 ⋅ Jan.… [cited by applicant]
Goetz, et al. “MDS-UPDRS: The MDS-sponsored Revision of the Unified Parkinson's Disease Rating Scale,” International Parkinson and Movement Disorder Society. 2018. 33 pages. [cited by applicant]
Boersma, et al. “Praat: doing phonetics by computer,” file:///S:/M/Modality.AI/103678.0005US%20Multimodal%20Dialog%20Based%20Remote%20Patient%20Monitoring/References,%20cited/Praat_%20doing%20Phonetics%20by%20Computer.h… [cited by applicant]
Kothare, et al. “Speech, Facial and Fine Motor Features for Conversation-Based Remote Assessment and Monitoring of Parkinson's Disease,” 4 pages. [cited by applicant]
Arora, et al. “Detecting and monitoring the symptoms of Parkinson's disease using smartphones: a pilot study,” 18 pages. [cited by applicant]
Roesler, et al. “Multimodal Dialog Based Remote Patient Monitoring of Motor Function in Parkinson's Disease and Other Movement Disorders,” Modalitiy.AI and U of CA, San Francisco, 6 pages. [cited by applicant]
Getz, Lindsey. “MMSE vs. MoCA: What You Should Know,” Today's Geriatric Medicine, 2 pages. [cited by applicant]
Jia, et al. “A comparison of the Mini-Mental State Examination (MMSE) with the Montreal Cognitive Assessment (MoCA) for mild cognitive impairment screening in Chinese middle-aged and older population: a crosssectional s… [cited by applicant]
Naumann, et al. “Investigating the Utility of Multimodal Conversational Technology and Audiovisual Analytic Measures for the Assessment and Monitoring of Amyotrophic Lateral Sclerosis at Scale,” Interspeech, 2021. 5 pag… [cited by applicant]
Vasquez-Correa, et al. “Towards an automatic evaluation of the dysarthria level of patients with Parkinson's disease,” Journal of Communicaiton Orders. 76 (2018) 21-36, 16 pages. [cited by applicant]
Mundt, et al. “Voice acoustic measures of depression severity and treatment response collected via interactive voice response (IVR) technology,” J Neurolinguistics. Jan. 2007 ; 20(1): 50-64, 17 pages. [cited by applicant]
Yunusova, Y., Green, J.R., Wang, J., Pattee, G., Zinman, L. A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS). J. Vis. Exp. (48), e2422, doi:10.3791/2422 (2011, 5 pages. [cited by applicant]
Lisetti, et al. “Now All Together: Overview of Virtual Health Assistants Emulating Face-to-Face Health Interview Experience.” 11 pages. [cited by applicant]
Morales, et al. “Modelling Errors in Automatic Speech Recognition for Dysarthric Speakers,” EURASIP Journal on Advances in Signal Processing vol. 2009, Article ID 308340, 14 pages. [cited by applicant]
Neumann et al., “On the Utility of Audiovisual Dialog Technologies and Signal Analytics for Real-time Remote Monitoring of Depression Biomarkers,” Modality.ai, Inc., 6 pages. [cited by applicant]
Ivascu et al., “A multi-agent Architecture for Ontology-based Diagnosis of Mental Disorders,” 17 International Symposium on Symbolic and Numeric Algorithms for Scientific Computing, 8 pages. [cited by applicant]