IP Library › Granted Patent US 12,046,155
Granted Patent B1
US 12,046,155 · App. 17/223,390 · Granted Jul 23, 2024

Automatic evaluation of argumentative writing by young students

Inventors: Debanjan Ghosh (Cambridge, MA); Beata Beigman Klebanov (Hopewell, NJ)
Assignee: Educational Testing Service
G09B7/02G06F40/20G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,046,155
App. No.
17/223,390
Filed
Apr 6, 2021
Granted
Jul 23, 2024
Kind
B1
Art Unit
2693
USPC
704/9
Abstract

Systems and methods are provided for automatic evaluation of argument critique essays written by young students in response to prompts. A transformer pre-trained for natural language processing is employed as a machine learning model, which is fine-tune with a first training dataset comprising unannotated argument critique essays written by college students, and then fine-tuned with a second training dataset comprising annotated argument critique essays written by middle school students, where each sentence in the second training dataset is annotated for the presence of valid critiques to prompts. The fine-tuned machine learning model is used to classify each sentence in an essay to be evaluated as either containing a valid critique or not.

Claims (38)

1. A computer-implemented method for automatically evaluating an argument critique essay, the method comprising:

fine-tune a machine learning model with a first training dataset, wherein the first training dataset comprises argument critique essays written by mature writers, and the argument critique essays in first training dataset are unannotated;

fine-tune the machine learning model with a second training dataset, wherein the second training dataset comprises argument critique essays written by young writers, and each sentence of each argument critique essay in the second training dataset is annotated for whether the sentence contains any valid critique;

receive the argument critique essay to be evaluated;

classify by the machine learning model every sentence in the argument critique essay to be evaluated as either (i) containing a valid critique, or (ii) not containing any valid critique.

2. The computer-implemented method of claim 1 , wherein the argument critique essay to be evaluated is written by a child, wherein the mature writers are college students, and wherein the young writers are middle school students.

3. The computer-implemented method of claim 1 , wherein the fine-tuning of the machine learning model with the second training dataset comprises:

fine-tune the machine learning model with a plurality of paired sentences in the second training dataset, wherein each pair in the plurality of paired sentences comprises (i) a first sentence that is annotated for whether the sentence contains any valid critique, and (ii) a second sentence that is an immediate next sentence following the first sentence in an argument critique essay in the second training dataset.

4. The computer-implemented method of claim 3 , wherein the second sentence is replaced with a special token when the first sentence is a last sentence in the argument critique essay in the second training dataset.

5. The computer implemented method of claim 1 , wherein the machine learning model is a transformer.

6. The computer implemented method of claim 5 , wherein the machine learning model is pre-trained for natural language processing.

7. The computer implemented method of claim 1 , wherein the method further comprises:

assigning a score to the argument critique essay to be evaluated based on a total number of sentences in the argument critique essay that are classify by the machine learning model as containing a valid critique.

8. The computer-implemented method of claim 1 , wherein the first training dataset comprises at least 5 million sentences, and the second training dataset comprises a smaller number of sentences than the first training dataset.

9. The computer-implemented method of claim 8 , wherein the second training dataset comprises no more than 3,000 sentences.

10. The computer-implemented method of claim 1 , wherein the second training dataset is annotated by a plurality of annotators, and an inter-annotator agreement among the plurality of annotators has a kappa value of at least 0.6.

11. The computer implemented method of claim 1 , wherein a valid critique comprises any criticism of a prompt on a ground of over-generalization, irrelevant example, misrepresentation of events, or neglecting negative side effects.

12. A system for automatically evaluating an argument critique essay written by a child in response to a prompt, the system comprising:

a processor; and

a computer-readable memory in communication with the processor, the computer-readable memory is encoded with instructions for commanding the processor to execute steps comprising:

receive the argument critique essay to be evaluated; and

classify by a machine learning model every sentence in the argument critique essay to be evaluated as either (i) containing a valid critique, or (ii) not containing any valid critique; wherein a valid critique comprises any criticism of the prompt on a ground of over-generalization, irrelevant example, misrepresentation of events, or neglecting negative side effects;

wherein the machine learning model is fine-tuned using:

a first training dataset, wherein the first training dataset comprises argument critique essays written by adults, and the argument critique essays in first training dataset are unannotated; and

a second training dataset, wherein the second training dataset comprises argument critique essays written by children, and each sentence of each argument critique essay in the second training dataset is annotated for whether the sentence contains any valid critique.

13. The system of claim 12 , wherein the machine learning model is a transformer pre-trained for natural language processing.

14. The system of claim 12 , wherein the machine learning model is Bidirectional Encoder Representations from Transformers (BERT).

15. The computer-implemented method of claim 12 , wherein the second training dataset comprises a smaller number of essays than the first training dataset.

16. The computer-implemented method of claim 12 , wherein the fine-tuning of the machine learning model with the second training dataset comprises:

fine-tune the machine learning model with a plurality of paired sentences in the second training dataset, wherein each pair in the plurality of paired sentences comprises (i) a first sentence that is annotated for whether the sentence contains any valid critique, and (ii) a second sentence that is an immediate next sentence following the first sentence in an argument critique essay in the second training dataset.

17. The computer-implemented method of claim 16 , wherein the second sentence is replaced with a special token when the first sentence is a last sentence in the argument critique essay in the second training dataset.

18. A method for configuring a machine learning model for evaluating argument critique essay written by children, the method comprising:

fine-tune the machine learning model with a first training dataset, wherein the first training dataset comprises argument critique essays written by mature writers, and the argument critique essays in first training dataset are unannotated;

fine-tune the machine learning model with a second training dataset, wherein the second training dataset comprises argument critique essays written by children, the argument critique essays written by children are arranged as a plurality of paired sentences, wherein each pair in the plurality of paired sentences comprises:

(i) a first sentence that is annotated for whether the sentence contains any valid critique, wherein a valid critique comprises any criticism of on a ground of over-generalization, irrelevant example, misrepresentation of events, or neglecting negative side effects; and

(ii) a second sentence that is an immediate next sentence following the first sentence in an argument critique essay in the second training dataset, wherein the second sentence is replaced with a special token when the first sentence is a last sentence in the argument critique essay.

19. The method of claim 18 , wherein the machine learning model is a transformer pre-trained for natural language processing.

20. The method of claim 18 , wherein the second training dataset comprises no more than 3,000 sentences, and the sentences in the second training dataset are annotated by a plurality of annotators, and an inter-annotator agreement among the plurality of annotators has a kappa value of at least 0.6.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2021
From: GHOSH, DEBANJAN; BEIGMAN KLEBANOV, BEATA
To: EDUCATIONAL TESTING SERVICE
Reel/Frame 055852/0476 →
Continuity (1)
Provisional Application 63006406 · Apr 7, 2020
Cited By (2)
US 12,596,889 US 12,718,433