IP Library Granted Patent US 11,727,954
Granted Patent B2
US 11,727,954 · App. 17/233,487 · Granted Aug 15, 2023

Diagnostic techniques based on speech-sample alignment

Inventor: Ilan D. Shallom (Gedera, IL)
Assignee: CORDIO MEDICAL LTD.
G10L25/66G10L15/22G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,727,954
App. No.
17/233,487
Granted
Aug 15, 2023
Kind
B2
Abstract

A method includes obtaining a first sequence of reference-sample feature vectors that quantify acoustic features of different respective portions of at least one reference speech sample, which was produced by a subject at a first time while a physiological state of the subject was known, and a second sequence of test-sample feature vectors that quantify the acoustic features of different respective portions of at least one test speech sample, which was produced by the subject at a second time while the physiological state of the subject was unknown. The test-sample feature vectors are mapped to respective ones of the reference-sample feature vectors, under predefined constraints, such that a total distance between the test-sample feature vectors and the respective ones of the reference-sample feature vectors is minimized. In response to the mapping, an output indicating the physiological state of the subject at the second time is generated.

Claims (47)

1. A method, comprising:

obtaining a first sequence of reference-sample feature vectors that quantify acoustic features of different respective portions of at least one reference speech sample, which was produced by a subject at a first time while a physiological state of the subject was known;

obtaining a second sequence of test-sample feature vectors that quantify the acoustic features of different respective portions of at least one test speech sample, which was produced by the subject at a second time while the physiological state of the subject was unknown;

aligning the test speech sample with the reference speech sample, by mapping the test-sample feature vectors to respective ones of the reference-sample feature vectors, under predefined constraints, such that a total distance between the second sequence and the first sequence is minimized; and

in response to aligning the test speech sample with the reference speech sample, generating an output indicating the physiological state of the subject at the second time.

2. The method according to claim 1 , further comprising receiving the test speech sample, wherein obtaining the test-sample feature vectors comprises obtaining the test-sample feature vectors by computing the test-sample feature vectors based on the test speech sample.

3. The method according to claim 1 , wherein the total distance is derived from respective local distances between the test-sample feature vectors and the respective ones of the reference-sample feature vectors.

4. The method according to claim 3 , wherein the total distance is a weighted sum of the local distances.

5. The method according to claim 4 , wherein aligning the test speech sample with the reference speech sample comprises aligning the test speech sample with the reference speech sample using a dynamic time warping (DTW) algorithm.

6. The method according to claim 1 , wherein generating the output comprises:

comparing the total distance to a predetermined threshold; and

generating the output in response to the comparison.

7. The method according to claim 1 , wherein the reference speech sample was produced while the physiological state of the subject was stable with respect to a particular physiological condition.

8. The method according to claim 1 , wherein the reference speech sample was produced while the physiological state of the subject was unstable with respect to a particular physiological condition.

9. The method according to claim 1 , wherein the reference speech sample and the test speech sample include the same predetermined utterance.

10. The method according to claim 1 , wherein the reference speech sample includes free speech of the subject, and wherein the test speech sample includes a plurality of speech units that are included in the free speech.

11. Apparatus, comprising:

a network interface; and

a processor, configured to:

obtain a first sequence of reference-sample feature vectors that quantify acoustic features of different respective portions of at least one reference speech sample, which was produced by a subject at a first time while a physiological state of the subject was known,

obtain, via the network interface, a second sequence of test-sample feature vectors that quantify the acoustic features of different respective portions of at least one test speech sample, which was produced by the subject at a second time while the physiological state of the subject was unknown,

align the test speech sample with the reference speech sample, by mapping the test-sample feature vectors to respective ones of the reference-sample feature vectors, under predefined constraints, such that a total distance between the second sequence and the first sequence is minimized, and

in response to aligning the test speech sample with the reference speech sample, generate an output indicating the physiological state of the subject at the second time.

12. The apparatus according to claim 11 , wherein the processor is configured to obtain the second sequence by:

receiving the test speech sample, and

computing the test-sample feature vectors based on the test speech sample.

13. The apparatus according to claim 11 , wherein the total distance is derived from respective local distances between the test-sample feature vectors and the respective ones of the reference-sample feature vectors.

14. The apparatus according to claim 11 , wherein the processor is configured to generate the output by:

comparing the total distance to a predetermined threshold, and

generating the output in response to the comparison.

15. The apparatus according to claim 11 , wherein the reference speech sample was produced while the physiological state of the subject was stable with respect to a particular physiological condition.

16. A system, comprising:

an analog-to-digital (A/D) converter; and

one or more processors, configured to cooperatively carry out a process that includes:

obtaining a first sequence of reference-sample feature vectors that quantify acoustic features of different respective portions of at least one reference speech sample, which was produced by a subject at a first time while a physiological state of the subject was known,

receiving, via the A/D converter, at least one test speech sample that was produced by the subject at a second time while the physiological state of the subject was unknown,

computing a second sequence of test-sample feature vectors that quantify the acoustic features of different respective portions of the test speech sample,

aligning the test speech sample with the reference speech sample, by mapping the test-sample feature vectors to respective ones of the reference-sample feature vectors, under predefined constraints, such that a total distance between the second sequence and the first sequence is minimized, and

in response to aligning the test speech sample with the reference speech sample, generating an output indicating the physiological state of the subject at the second time.

17. The system according to claim 16 , wherein the process further includes receiving the reference speech sample, and wherein obtaining the first sequence of reference-sample feature vectors includes obtaining the first sequence of reference-sample feature vectors by computing the reference-sample feature vectors based on the reference speech sample.

18. The system according to claim 16 , wherein the reference speech sample and the test speech sample include the same predetermined utterance.

19. A computer software product comprising a tangible non-transitory computer-readable medium in which program instructions are stored, which instructions, when read by a processor, cause the processor to:

obtain a first sequence of reference-sample feature vectors that quantify acoustic features of different respective portions of at least one reference speech sample, which was produced by a subject at a first time while a physiological state of the subject was known,

obtain a second sequence of test-sample feature vectors that quantify the acoustic features of different respective portions of at least one test speech sample, which was produced by the subject at a second time while the physiological state of the subject was unknown,

align the test speech sample with the reference speech sample, by mapping the test-sample feature vectors to respective ones of the reference-sample feature vectors, under predefined constraints, such that a total distance between the second sequence and the first sequence is minimized, and

in response to aligning the test speech sample with the reference speech sample, generate an output indicating the physiological state of the subject at the second time.

20. The computer software product according to claim 19 , wherein the instructions further cause the processor to receive the test speech sample, and wherein the instructions cause the processor to obtain the second sequence of test-sample feature vectors by computing the test-sample feature vectors based on the test speech sample.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2021
From: SHALLOM, ILAN D.
To: CORDIO MEDICAL LTD.
Reel/Frame 055951/0225 →
Continuity (2)
Continuation 16299178 · Mar 12, 2019
Related Publication 20210256992A1 · Aug 19, 2021
Cited By (2)
US 12,336,840 US 12,367,876