IP Library Granted Patent US 12,537,018
Granted Patent B2
US 12,537,018 · App. 18/217,880 · Granted Jan 27, 2026

Method and system for predicting a mental condition of a speaker

Inventors: Daniel Sand (Tel Aviv-Jaffa, IL); Samuel Jefroykin (Tel Aviv-Jaffa, IL); Ilan Kahan (Tel Aviv-Jaffa, IL); Tal Simon (Tel Aviv-Jaffa, IL); Alon Rabinovich (Tel Aviv-Jaffa, IL); Shiri Sadeh-Sharvit (Tel Aviv-Jaffa, IL)
Assignee: ELEOS MENTAL SYSTEMS LTD.
G10L25/63G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,537,018
App. No.
18/217,880
Granted
Jan 27, 2026
Kind
B2
Abstract

Systems and methods of the present invention may relate to prediction of a mental health score, representing a mental condition of a speaker. Embodiments of the invention may analyze speech by receiving an audio data element representing a discussion; extracting a first set of audio segments pertaining to speech of a first speaker in the discussion; analyzing the first set of audio segments to produce a set of audio features; and applying a machine-learning (ML) model on the set of audio features, to predict a mental health score, representing a mental condition of the first speaker. Embodiments of the invention may provide technical means for supporting psychiatric assessment of mental disorders such as depression and anxiety in an automatic, nondisruptive manner, e.g., without direct patient's collaboration, thereby increasing validity of the assessment.

Claims (59)

1 . A method of analyzing speech by at least one processor, the method comprising:

receiving an audio data element representing a discussion;

extracting a first set of audio segments pertaining to speech of a first speaker in the discussion;

analyzing the first set of audio segments by:

applying a band-pass filter on an audio segment of the first set of audio segments, to obtain a filtered version of the audio segment;

applying a Hilbert transform on the filtered version of the audio segment, to obtain a Hilbert envelope;

determining a burst threshold value;

defining at least one audio burst based on the Hilbert envelope and the burst threshold value; and

analyzing the at least one audio burst, to calculate one or more audio features comprising an audio burst Area Under Curve (AUC) defined as an integral between a positive edge of the Hilbert envelope and the burst threshold value; and

applying a machine-learning (ML) model on said one or more audio features, to predict a mental health score, representing a mental condition of the first speaker.

2 . The method of claim 1 , wherein the mental health score represents an expected response of the first speaker in a mental health questionnaire.

3 . The method of claim 1 , further comprising analyzing the first set of audio segments to produce one or more textual n-grams;

analyzing the one or more textual n-grams to produce at least one textual feature, representing a mental condition of the first speaker; and

further applying the ML model on the at least one textual feature, to predict the mental health score.

4 . The method of claim 1 , wherein

said one or more audio features further comprise features selected from a list consisting of: audio burst amplitudes, audio burst duration, and audio burst coefficient of variation (CV).

5 . The method of claim 1 , further comprising:

monitoring the predicted mental health score over time; and

identifying one or more timestamps corresponding to points in the discussion, where the first speaker is suspected to have been in a predefined mental condition, based on said monitoring.

6 . The method of claim 5 , further comprising:

producing a reference data element, comprising a plurality of identified timestamps, corresponding to one or more discussions, and respective predictions of mental health score; and

providing a notification of previous at least one case in which the first speaker is suspected to have been in the predefined mental condition, based on the reference data element.

7 . The method of claim 6 , further comprising:

analyzing the reference data element, in relation to the plurality of identified timestamps; and

producing a recommendation data element, representing a recommendation of treatment, based on said analysis.

8 . The method of claim 6 , further comprising:

extracting a second set of audio segments pertaining to speech of a second speaker in the discussion;

analyzing the second set of audio segments to produce one or more second textual n-grams; and

associating the one or more second textual n-grams to at least one identified timestamp in the reference data element.

9 . A system for analyzing speech, the system comprising: a non-transitory memory device, wherein modules of instruction code are stored, and at least one processor associated with the memory device, and configured to execute the modules of instruction code, whereupon execution of said modules of instruction code, the at least one processor is configured to:

receive an audio data element representing a discussion;

extract a first set of audio segments pertaining to speech of a first speaker in the discussion;

analyze the first set of audio segments by:

applying a band-pass filter on an audio segment of the first set of audio segments, to obtain a filtered version of the audio segment;

applying a Hilbert transform on the filtered version of the audio segment, to obtain a Hilbert envelope;

determining a burst threshold value;

defining at least one audio burst based on the Hilbert envelope and the burst threshold value; and

analyzing the at least one audio burst, to calculate one or more audio features comprising an audio burst Area Under Curve (AUC) defined as an integral between a positive edge of the Hilbert envelope and the burst threshold value; and

apply a machine-learning (ML) model on said one or more audio features, to predict a mental health score, representing a mental condition of the first speaker.

10 . The system of claim 9 , wherein the mental health score represents an expected response of the first speaker in a mental health questionnaire.

11 . The system of claim 9 , wherein the at least one processor is further configured to:

analyze the first set of audio segments to produce one or more textual n-grams;

analyze the one or more textual n-grams to produce at least one textual feature, representing a mental condition of the first speaker; and

apply the ML model on the at least one textual feature, to predict the mental health score.

12 . The system of claim 9 , wherein

said one or more audio features further comprise features selected from a list consisting of: audio burst amplitudes, audio burst duration, and audio burst coefficient of variation (CV).

13 . The system of claim 9 , wherein the at least one processor is further configured to:

monitor the predicted mental health score over time; and

identify one or more timestamps corresponding to points in the discussion, where the first speaker is suspected to have been in a predefined mental condition, based on said monitoring.

14 . The system of claim 13 , wherein the at least one processor is further configured to:

produce a reference data element, comprising a plurality of identified timestamps, corresponding to one or more discussions, and respective predictions of mental health score; and

provide a notification of previous at least one case in which the first speaker is suspected to have been in the predefined mental condition, based on the reference data element.

15 . The system of claim 14 , wherein the at least one processor is further configured to:

perform an analysis of the reference data element, in relation to the plurality of identified timestamps; and

produce a recommendation data element, representing a recommendation of treatment, based on said analysis.

16 . The system of claim 14 , wherein the at least one processor is further configured to:

extract a second set of audio segments pertaining to speech of a second speaker in the discussion;

analyze the second set of audio segments to produce one or more second textual n-grams; and

associate the one or more second textual n-grams to at least one identified timestamp in the reference data element.

Assignments (2)
SECURITY INTEREST Recorded Mar 26, 2026
From: ELLEOS MENTAL SYSTEMS LTD
To: HSBC BANK PLC
Reel/Frame 074193/0296 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 3, 2023
From: SAND, DANIEL; JEFROYKIN, SAMUEL; KAHAN, ILAN; SIMON, TAL; RABINOVICH, ALON; SADEH-SHARVIT, SHIRI
To: ELEOS MENTAL SYSTEMS LTD.
Reel/Frame 064140/0867 →
Continuity (2)
Provisional Application 63402540 · Aug 31, 2022
Related Publication 20240071412A1 · Feb 29, 2024
References Cited (10)
US 10580435B2 · Ashoori · 2020 [cited by examiner]
US 11887622B2 · Hu · 2024 [cited by examiner]
US 11922968B2 · Stojancic · 2024 [cited by examiner]
US 20170206915A1 · Prasad · 2017 [cited by examiner]
US 20190385711A1 · Shriberg · 2019 [cited by examiner]
US 20230352194A1 · Kim · 2023 [cited by examiner]
US 20240374187A1 · Liyanage · 2024 [cited by examiner]
JP 2009003162A · 2009 [cited by examiner]
Patient Health Questionnaire-9 (PHQ-9)—Mental health screening—National HIV curriculum. (2024). Available online: [https://www.hiv.uw.edu/page/mental-health-screening/phq-9]. [cited by applicant]
Generalized Anxiety Disorder 7-Item (GAD-7)—Mental health screening—National HIV curriculum. (2024). Available online: [https://www.hiv.uw.edu/page/mental-health-screening/gad-7]. [cited by applicant]