IP Library Granted Patent US 9,799,228
Granted Patent B2
US 9,799,228 · App. 14/152,178 · Granted Oct 24, 2017

Systems and methods for natural language processing for speech content scoring

Inventors: Lei Chen (Pennington, NJ); Klaus Zechner (Princeton, NJ); Anastassia Loukina (Princeton, NJ)
Assignee: Educational Testing Service
G09B7/02G10L15/01G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,799,228
App. No.
14/152,178
Granted
Oct 24, 2017
Kind
B2
Abstract

Computer-implemented systems and methods are provided for scoring content of a spoken response to a prompt. A scoring model is generated for a prompt, where generating the scoring model includes generating a transcript for each of a plurality of training responses to the prompt, dividing the plurality of training responses into clusters based on the transcripts of the training responses, selecting a subset of the training responses in each cluster for scoring, scoring the selected subset of training responses for each cluster, and generating content training vectors using the transcripts from the scored subset. A transcript is generated for a received spoken response to be scored, and a similarity metric is computed between the transcript of the spoken response to be scored and the content training vectors. A score is assigned to the spoken response based on the determined similarity metric.

Claims (50)

1. A computer-implemented method of scoring content of a recording of a spoken response to a prompt, comprising:

performing automatic speech recognition on a plurality of recordings of training responses to generate a scoring model for a prompt, wherein generating the scoring model comprises:

generating a transcript for each of a plurality of training responses to the prompt using a trained acoustic model;

dividing the plurality of training responses into clusters based on the transcripts of the training responses based on one or more criteria of the training responses;

selecting a subset of the training responses in each cluster for scoring;

scoring the selected subset of training responses for each cluster; and

generating content training vectors using the transcripts from the scored subset, wherein each content training vector is associated with a scored subset, wherein each content training vector identifies vectors of words that appear in training responses of the scored subset associated with that content training vector;

generating, using the trained acoustic model, a transcript for a received spoken response to be scored;

computing a similarity metric between words in the transcript of the spoken response to be scored and words in the vectors of words of the content training vectors; and

assigning a score to the spoken response based on the determined similarity.

2. The method of claim 1 , wherein generating the scoring model for the prompt further comprises:

determining a score for each cluster based on scores for the selected subset for each cluster; and

wherein the content training vectors are generated based on these cluster scores and all transcripts of training responses.

3. The method of claim 1 , wherein the subset of training responses for a particular cluster is selected as a subset of m responses in the particular cluster that are closest to a center of the particular cluster.

4. The method of claim 1 , wherein the subset of training responses for a particular cluster is selected as a random sample of m responses in the particular cluster.

5. The method of claim 1 , wherein spoken responses to the prompt are scored on a scale of 1-n, wherein the plurality of training responses are divided into n clusters.

6. The method of claim 1 , wherein spoken responses to the prompt are scored on a scale of 1-n, wherein the plurality of training responses are divided into more than n clusters, wherein the more than n clusters are combined to form n clusters.

7. The method of claim 1 , wherein the plurality of training responses are divided into clusters using a latent dirichlet allocation or a latent semantic analysis procedure.

8. The method of claim 1 , wherein the spoken response to be scored is determined to belong to a cluster using a content vector analysis procedure.

9. The method of claim 8 , wherein generating the scoring model further comprises determining a word vector for each cluster based on transcripts of training responses in each cluster.

10. The method of claim 9 , further comprising generating a word vector for the spoken response, wherein the word vector for the spoken response to be scored is compared to word vectors for each vector to determine to which cluster the spoken response to be scored belongs.

11. The method of claim 1 , wherein the selected subset of training responses for each cluster are scored based on the content of those training responses.

12. The method of claim 1 , wherein the selected subset of training responses for each cluster are scored automatically.

13. The method of claim 1 , wherein the selected subset of training responses for each cluster are scored automatically based on non-content features of those training responses, wherein the score assigned to the spoken response to be scored is a content score.

14. The method of claim 13 , wherein the non-content features include one or more of fluency, prosody, grammar, pronunciation, and vocabulary.

15. The method of claim 1 , wherein the transcript for each training response is generated using automatic speech recognition or manual transcription.

16. A computer-implemented system for scoring content of a spoken response to a prompt, comprising:

one or more data processors;

one or more non-transitory computer-readable mediums containing instructions for commanding the one or more data processors to execute steps that include:

performing automatic speech recognition on a plurality of recordings of training responses to generate a scoring model for a prompt, wherein generating the scoring model comprises:

generating a transcript for each of a plurality of training responses to the prompt using a trained acoustic model;

dividing the plurality of training responses into clusters based on the transcripts of the training responses based on one or more criteria of the training responses;

selecting a subset of the training responses in each cluster for scoring; scoring the selected subset of training responses for each cluster; and

generating content training vectors using the transcripts from the scored subset, wherein each content training vector is associated with a scored subset; wherein each content training vector identifies vectors of words that appear in training responses of the scored subset associated with that content training vector;

generating, using the trained acoustic model, a transcript for a received spoken response to be scored;

computing a similarity metric between words in the transcript of the spoken response to be scored and words in the vectors of words of the content training vectors; and

assigning a score to the spoken response based on the determined similarity.

17. The system of claim 16 , wherein generating the scoring model for the prompt further comprises:

determining a score for each cluster based on scores for the selected subset for each cluster; and

wherein the content training vectors are generated based on these cluster scores and all transcripts of training responses.

18. The system of claim 16 , wherein the subset of training responses for a particular cluster is selected as a subset of m responses in the particular cluster that are closest to a center of the particular cluster.

19. The system of claim 16 , wherein the subset of training responses for a particular cluster is selected as a random sample of m responses in the particular cluster.

20. The system of claim 16 , wherein spoken responses to the prompt are scored on a scale of 1-n, wherein the plurality of training responses are divided into n clusters.

21. The system of claim 16 , wherein spoken responses to the prompt are scored on a scale of 1-n, wherein the plurality of training responses are divided into more than n clusters, wherein the more than n clusters are combined to form n clusters.

22. The system of claim 16 , wherein the plurality of training responses are divided into clusters using a latent dirichlet allocation or a latent semantic analysis procedure.

23. The system of claim 16 , wherein the selected subset of training responses for each cluster are scored manually by humans.

24. The system of claim 16 , wherein:

the spoken response to be scored is determined to belong to a cluster using a content vector analysis procedure;

the generating of the scoring model further comprises determining a word vector for each cluster based on transcripts of training responses in each cluster; and

the steps further include generating a word vector for the spoken response to be scored, wherein the word vector for the spoken response to be scored is compared to word vectors for each vector to determine to which cluster the spoken response to be scored belongs.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE STATE OF INCORPORATION INSIDE THE ASSIGNMENT DOCUMENT PREVIOUSLY RECORDED AT REEL: 032827 FRAME: 0103. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 1, 2015
From: CHEN, LEI; ZECHNER, KLAUS; LOUKINA, ANASTASSIA
To: EDUCATIONAL TESTING SERVICE
Reel/Frame 035799/0224 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2014
From: CHEN, LEI; ZECHNER, KLAUS; LOUKINA, ANASTASSIA
To: EDUCATIONAL TESTING SERVICE
Reel/Frame 032827/0103 →
Continuity (2)
Provisional Application 61751300 · Jan 11, 2013
Related Publication 20140199676A1 · Jul 17, 2014