Systems and methods for automated machine learning model training quality control
Systems and methods for automated custom training of a scoring model are disclosed herein. The method include: receiving a plurality of responses received from a plurality of students in response to providing of a prompt; identifying an evaluation model relevant to the provided prompt, which evaluation model can be a machine learning model trained to output a score relevant to at least portions of a response; generating a training indicator that provides a graphical depiction of the degree to which the identified evaluation model is trained; determining a training status of the model; receiving at least one evaluation input when the model is identified as insufficiently trained; updating training of the evaluation model based on the at least one received evaluation input; and controlling the training indicator to reflect the degree to which the evaluation model is trained subsequent to the updating of the training of the evaluation model.
1. A system for controlling training quality of a machine learning model, the system comprising:
a memory comprising: a content library database containing a plurality of prompts and evaluation data associated with each of the plurality of prompts; and a model database comprising a plurality of evaluation models for automated evaluation of received user responses, wherein the evaluation data of each of the plurality of prompts includes a pointer linking to the associated evaluation model; and
a processing unit, including at least one processor, the processing unit configured to:
receive a plurality of responses received from a plurality of users in response to providing of a prompt from the plurality of prompts;
identify, based at least on an associated pointer for the provided prompt, an evaluation model from the plurality of evaluation models relevant to the provided prompt, wherein the evaluation model comprises a machine learning model trained to output a score relevant to at least portions of a response;
generate a training indicator, wherein the training indicator provides a graphical depiction of a degree to which the evaluation model is trained;
determine a training status of the evaluation model;
control the training indicator to identify the training status of the evaluation model;
automatically evaluate the plurality of responses with the evaluation model when the evaluation model is identified as sufficiently trained;
after the plurality of responses have been auto-evaluated, generate output data characterizing results of the auto-evaluation of the plurality of responses;
generate, based on the output data, a graphical indicator for the evaluation model, wherein the graphical indicator characterizes performance of the auto-evaluation by the evaluation model; and
control a user interface to display the graphical indicator.
2. The system of claim 1 , wherein the graphical indicator indicates a confidence level of the evaluation model.
3. The system of claim 1 , wherein the graphical indicator identifies outlier scores based on historical user data.
4. The system of claim 3 , wherein identifying outlier scores based on historical user data comprises: retrieving historical data; comparing the historical data to the results of evaluating the plurality of responses; and indicating an outlier score when a discrepancy between the historical data and the results of evaluating the plurality of responses exceeds a threshold level.
5. The system of claim 4 , wherein the historical data comprises a historical evaluation result distribution and wherein the results of evaluating the plurality of responses comprises an evaluation result distribution.
6. The system of claim 4 , wherein the historical data comprises user historical data, wherein the user historical data relates to previously received evaluation results for at least one user.
7. The system of claim 3 , wherein the processing unit is further configured to determine acceptability of the auto-evaluation.
8. The system of claim 7 , wherein the acceptability of the auto-evaluation is determined based on the identified outlier scores.
9. The system of claim 8 , wherein the processing unit is further configured to further train the evaluation model when the auto-evaluation is unacceptable.
10. The system of claim 1 , wherein the processing unit is further configured to: receive a selection of at least one response for reevaluation; receive an evaluation input for the at least one response; and update training of the evaluation model based on the received evaluation input for the at least one response.
11. A method of controlling training quality of a machine learning model, the method comprising:
receiving a plurality of responses received from a plurality of users in response to providing of a prompt from a plurality of prompts included in a content library database;
identifying an evaluation model relevant to the provided at least one prompt from a plurality of evaluation models included in a model database, wherein the evaluation model comprises a machine learning model trained to output a score relevant to at least portions of a response;
generating a training indicator, wherein the training indicator provides a graphical depiction of a degree to which the evaluation model is trained;
determining a training status of the evaluation model;
controlling the training indicator to identify the training status of the evaluation model;
automatically evaluating the plurality of responses with the evaluation model when the evaluation model is identified as sufficiently trained;
after the plurality of responses have been auto-evaluated, generating output data characterizing results of the auto-evaluation of the plurality of responses;
generating, based on the output data, a graphical indicator for the evaluation model, wherein the graphical indicator characterizes performance of the auto-evaluation by the evaluation model; and
controlling a user interface to display the graphical indicator.
12. The method of claim 11 , wherein the graphical indicator indicates an accuracy level of the evaluation model.
13. The method of claim 11 , wherein the graphical indicator identifies at least one outlier score based on historical user data.
14. The method of claim 13 , wherein identifying the at least one outlier score includes: retrieving historical data; comparing the historical data to the results of evaluating the plurality of responses; and indicating the at least one outlier score when a discrepancy between the historical data and the results of evaluating the plurality of responses exceeds a threshold level.
15. The method of claim 14 , wherein the historical data comprises a historical evaluation result distribution and wherein the results of evaluating the plurality of responses comprises an evaluation result distribution.
16. The method of claim 14 , wherein the historical data comprises user historical data, wherein the user historical data relates to previously received evaluation results for at least one user.
17. The method of claim 13 , further comprising determining an acceptability of the auto-evaluation.
18. The method of claim 17 , wherein the acceptability of the auto-evaluation is determined based on the identified outlier scores.
19. The method of claim 18 , further comprising training the evaluation model when the auto-evaluation is unacceptable.
20. The method of claim 11 , further comprising: receiving a selection of at least one response for reevaluation; receiving an evaluation input for the at least one response; and updating training of the evaluation model based on the received evaluation input for the at least one response.