IP Library Granted Patent US 8,457,967
Granted Patent B2
US 8,457,967 · App. 12/541,927 · Granted Jun 4, 2013

Automatic evaluation of spoken fluency

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,457,967
App. No.
12/541,927
Granted
Jun 4, 2013
Kind
B2
Abstract

A procedure to automatically evaluate the spoken fluency of a speaker by prompting the speaker to talk on a given topic, recording the speaker's speech to get a recorded sample of speech, and then analyzing the patterns of disfluencies in the speech to compute a numerical score to quantify the spoken fluency skills of the speakers. The numerical fluency score accounts for various prosodic and lexical features, including formant-based filled-pause detection, closely-occurring exact and inexact repeat N-grams, normalized average distance between consecutive occurrences of N-grams. The lexical features and prosodic features are combined to classify the speaker with a C-class classification and develop a rating for the speaker.

Claims (51)

1. A method of evaluating spoken fluency of a speaker, comprising:

capture a speech sample of the speaker;

process the speech sample to compute prosodic features related to fluency evaluation;

convert the speech sample to text using an automatic speech recognizer;

compute lexical features from an output of the automatic speech recognizer; and

combine the lexical features and the prosodic features to classify the speaker and develop a rating for the speaker;

wherein the capturing of the speech sample comprises prompting the speaker to speak on a first topic, prompting the speaker to speak on a second topic, and prompting the speaker to speak on a third topic; and

wherein the first topic is more familiar to the speaker than the second and third topics, and the second topic is more familiar to the speaker than the third topic.

2. The method of claim 1 , wherein the combining of the lexical features and the prosodic features further comprises:

perform a C-class classification to classify the speaker.

3. The method of claim 1 , wherein the prosodic features include filled-pause features and amount of silence based features.

4. The method of claim 3 , wherein the filled-pauses features are detected using measures based on stability of the formants of the speech signal.

5. The method of claim 1 , wherein the lexical features include features selected from a group consisting of a count of total word repetitions, a count of closely repeated exact and inexact N-grams, and a normalized average distance between consecutive occurrences of N-grams.

6. The method of claim 1 , wherein the combination of lexical and prosodic features is hierarchical, the method further comprising:

using the prosodic features to either validate or invalidate a disfluency hypothesis made by the lexical features, or using the lexical features to either validate or invalidate a disfluency hypothesis made by the prosodic features.

7. The method of claim 1 , further comprising:

detecting and identifying different disfluency characteristics present in the speech sample.

8. The method of claim 7 , further comprising:

provide feedback including a list of the different disfluency characteristics and an indication of a relative proportion of the different disfluency characteristics.

9. The method of claim 7 , further comprising:

provide feedback including an indication of locations of the different disfluency characteristics within the speech sample, and an indication of a type of disfluency characteristic at each of said locations.

10. A software product comprising a program of instructions stored on a machine readable device for evaluating spoken fluency of a speaker, wherein the program of instructions upon being executed on a computer causes the computer to perform activities comprising: capturing a speech sample of the speaker;

processing the speech sample to compute prosodic features related to fluency evaluation;

converting the speech sample to text using an automatic speech recognizer;

computing lexical features from an output of the automatic speech recognizer; and

combining the lexical features and the prosodic features to classify the speaker and develop a rating for the speaker;

wherein the capturing of the speech sample comprises prompting the speaker to speak on a first topic, prompting the speaker to speak on a second topic, and prompting the speaker to speak on a third topic; and

wherein the first topic is more familiar to the speaker than the second and third topics, and the second topic is more familiar to the speaker than the third topic.

11. The software product of claim 10 , wherein the combining of the lexical features and the prosodic features further comprises:

perform a C-class classification to classify the speaker.

12. The software product of claim 10 , wherein the prosodic features include filled-pause features and amount of silence based features; and

wherein the filled-pauses features are detected using measures based on stability of the formants of the speech signal.

13. The software product of claim 10 , wherein the lexical features include features selected from a group consisting of a count of total word repetitions, a count of closely repeated exact and inexact N-grams, and a normalized average distance between consecutive occurrences of N-grams.

14. The software product of claim 10 , wherein the combination of lexical and prosodic features is hierarchical, the activities further comprising:

using the prosodic features to either validate or invalidate a disfluency hypothesis made by the lexical features, or using the lexical features to either validate or invalidate a disfluency hypothesis made by the prosodic features.

15. The software product of claim 10 , further comprising:

detecting and identifying different disfluency characteristics present in the speech sample.

16. The software product of claim 15 , further comprising:

provide feedback including a list of the different disfluency characteristics and an indication of a relative proportion of the different disfluency characteristics; wherein

said feedback includes an indication of locations of the different disfluency characteristics within the speech sample, and an indication of a type of disfluency characteristic at each of said locations.

17. The software product of claim 10 , wherein the first topic comprises asking the speaker questions and recording the speaker's answers to the questions to include in the sample, and wherein the second and third topics are selected based on content from the speaker's answers that indicates a familiarity with one or more topics.

18. A system configured to evaluate spoken fluency of a speaker, the system comprising:

a recording device configured to capture a speech sample of the speaker;

a processor configured to process the speech sample to compute prosodic features related to fluency evaluation;

an automatic speech recognition module configured to convert the speech sample to text;

a memory configured to store instructions for computing lexical features from an output of the automatic speech recognizer; and

a display device configured to display feedback created by combining the lexical features and the prosodic features to classify the speaker and develop a rating for the speaker;

wherein the capturing of the speech sample by the recording device comprises prompting the speaker to speak on a first topic, prompting the speaker to speak on a second topic, and prompting the speaker to speak on a third topic; and

wherein the first topic is more familiar to the speaker than the second and third topics, and the second topic is more familiar to the speaker than the third topic.

19. The method of claim 1 , wherein the first topic comprises asking the speaker questions and recording the speaker's answers to the questions to include in the sample, and wherein the second and third topics are selected based on content from the speaker's answers that indicates a familiarity with one or more topics.

20. The system of claim 18 , wherein the first topic comprises asking the speaker questions and recording the speaker's answers to the questions to include in the sample, and wherein the second and third topics are selected based on content from the speaker's answers that indicates a familiarity with one or more topics.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065566/0013 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2013
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 030323/0965 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2009
From: AUDHKHASI, KARTIK; DESHMUKH, OM D.; KANDHWAY, KUNDAN; VERMA, ASHISH
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 023103/0888 →