IP Library › Granted Patent US 11,049,409
Granted Patent B1
US 11,049,409 · App. 14/974,721 · Granted Jun 29, 2021

Systems and methods for treatment of aberrant responses

Inventors: Mo Zhang (Plainsboro, NJ); Jing Chen (Princeton, NJ); Andre Alexander Rupp (Princeton, NJ); David Michael Williamson (Yardley, PA)
Assignee: Educational Testing Service
G09B7/02G06F40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,049,409
App. No.
14/974,721
Granted
Jun 29, 2021
Kind
B1
Abstract

Systems and methods are provided for automatically scoring a response and statistically revaluating whether it can be considered as aberrant. In one embodiment, a constructed response is evaluated via a pre-screening stage and a post-hoc screening stage. The pre-screening stage attempts to determine whether the constructed response is aberrant based on a variety of aberration metrics and criteria. If the constructed response is deemed not to be aberrant, then the post-hoc screening stage attempts to predict a discrepancy between what score an automated scoring system would assigned and what score a human rater would assign to the response. If the discrepancy is sufficiently low, then the constructed response may be scored by an automated scoring engine. On the other hand, if the constructed response failed to pass either of the two stages, then a flag may be raised to indicate that additional human review may be needed.

Claims (37)

1. A computer-implemented method of automatically scoring a constructed response and statistically evaluating whether it is an aberrant response, comprising:

collecting reference responses with known discrepancies between their automated scores and human scores;

extracting values relating to discrepancy features from the reference responses;

training a discrepancy prediction model with the automated scores of the reference responses, the human scores of the reference responses, and the values relating to discrepancy features extracted from the reference responses; wherein the discrepancy prediction model predicts a level of discrepancy between scores from an automated scoring engine and a human score from a human rater;

receiving a constructed response from a source; wherein the source is a server, database, local storage or memory, or examinee computer;

extracting, from the constructed response, values relating to aberration metrics associated with a plurality of aberration characteristics selected from the group consisting of excessive repetition, lack of key concept, irrelevance to a prompt, restatement of a prompt, off topic, overly brief, overly long, unidentifiable organization, excessive number of problems, and unexpected topic;

determining that the constructed response is not sufficiently aberrant based on the extracted aberration metrics values and a predetermined set of aberration criteria;

extracting, from the constructed response, values relating to discrepancy features used by the discrepancy prediction model; and

generating a discrepancy prediction using the discrepancy prediction model and the values relating to discrepancy features extracted from the constructed response;

extracting, from the constructed response, values relating to scoring features used by a scoring model; and

generating an automated score for the constructed response using the scoring model and the extracted scoring features values.

2. The method of claim 1 , wherein the training responses used for calibrating the discrepancy prediction model are empirically determined.

3. The method of claim 1 , wherein the discrepancy prediction model's discrepancy features include linguistic features of responses.

4. The method of claim 1 , wherein the discrepancy prediction model's discrepancy features include activity measures relating to processes in which responses are constructed.

5. The method of claim 1 , wherein the discrepancy prediction model's discrepancy features include data associated with demographic characteristics of response authors.

6. A computer-implemented system for automatically scoring a constructed response and statistically evaluating whether is an aberrant response, comprising:

a processing system;

one or more computer-readable mediums encoded with instructions for commanding the processing system to execute steps comprising:

extracting values relating to discrepancy features from a collection of reference responses with known discrepancies between the reference responses' automated scores and human scores;

training a discrepancy prediction model with the automated scores of the reference responses, the human scores of the reference responses, and the values relating to discrepancy features extracted from the reference responses; wherein the discrepancy prediction model predicts a level of discrepancy between scores from an automated scoring engine and a human score from a human rater;

receiving a constructed response from a source; wherein the source is a server, database, local storage or memory, or examinee computer;

extracting, from the constructed response, values relating to aberration metrics associated with a plurality of aberration characteristics selected from the group consisting of excessive repetition, lack of key concept, irrelevance to a prompt, restatement of a prompt, off topic, overly brief, overly long, unidentifiable organization, excessive number of problems, and unexpected topic;

determining that the constructed response is not sufficiently aberrant based on the extracted aberration metrics values and a predetermined set of aberration criteria;

extracting, from the constructed response, values relating to discrepancy features used by the discrepancy prediction model;

generating a discrepancy prediction using the discrepancy prediction model and the values relating to discrepancy features extracted from the constructed response;

extracting, from the constructed response, values relating to scoring features used by a scoring model; and

generating an automated score for the constructed response using the scoring model and the extracted scoring features values.

7. A non-transitory computer-readable medium encoded with instructions for commanding a processing system to execute steps for automatically scoring a constructed response and statistically evaluating whether it is an aberrant response, the steps comprising:

extracting values relating to discrepancy features from a collection of reference responses with known discrepancies between the reference responses' automated scores and human scores;

training a discrepancy prediction model with the automated scores of the reference responses, the human scores of the reference responses, and the values relating to discrepancy features extracted from the reference responses; wherein the discrepancy prediction model predicts a level of discrepancy between scores from an automated scoring engine and a human score from a human rater;

receiving a constructed response from a source; wherein the source is a server, database, local storage or memory, or examinee computer;

extracting, from the constructed response, values relating to aberration metrics associated with a plurality of aberration characteristics selected from the group consisting of excessive repetition, lack of key concept, irrelevance to a prompt, restatement of a prompt, off topic, overly brief, overly long, unidentifiable organization, excessive number of problems, and unexpected topic;

determining that the constructed response is not sufficiently aberrant based on the extracted aberration metrics values and a predetermined set of non fatal aberration criteria;

extracting, from the constructed response values relating to discrepancy features used by the discrepancy prediction model;

generating a discrepancy prediction using the discrepancy prediction model and the values relating to discrepancy features extracted from the constructed response;

extracting, from the constructed response, values relating to scoring features used by a scoring model; and

generating an automated score for the constructed response using the scoring model and the extracted scoring features values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 9, 2016
From: ZHANG, MO; CHEN, JING; RUPP, ANDRE ALEXANDER; WILLIAMSON, DAVID MICHAEL
To: EDUCATIONAL TESTING SERVICE
Reel/Frame 037928/0548 →
Continuity (1)
Provisional Application 62094368 · Dec 19, 2014
Cited By (1)
US 12,670,322