IP Library Granted Patent US 10,504,376
Granted Patent B2
US 10,504,376 · App. 12/429,818 · Granted Dec 10, 2019

System and method for improving the quality of computer generated exams

Inventors: Kjell Bjornar Nibe (Hamar, NO); Eirik Mikkelsen (Furnes, NO)
Assignee: RELIANT EXAMS AS
G09B7/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,504,376
App. No.
12/429,818
Granted
Dec 10, 2019
Kind
B2
Abstract

System and method for generating exams, questionnaires or similar, including a set of questions to be answered, the system comprising a data base including a number of questions in at least one topic, each question being associated with a data set related to answers given to the questions in previous uses of the question sets, the system further comprising selection means for randomly selecting a question within one or more predetermined topics from said data base, and an evaluation means for evaluating the selected questions relative to predetermined requirements to the selected questions and their corresponding answers in said data set, and discarding questions not fulfilling said requirements.

Claims (138)

1. A computer-based system for improving computer generated exams and generating exam question sets suited to separating test candidates of different skill levels, the system comprising:

a database including a number of questions in at least one predetermined topic, each question being associated with a data set comprising prior use information related to answers given to the question in previous uses of such question in question sets presented to test candidates;

said database is stored on at least one data storage device and is updated to reflect empirical data collected from previous exams to find selectivity, facility and reliability of questions and question sets presented to test candidates, wherein the empirical data for each question on the computer generated exams is collected to find the selectivity, reliability and facility of the questions and the database is updated to reflect the collected empirical data, said database further comprising unused questions including newly added questions;

the system further comprising selection means executing on a computer for randomly selecting questions within one or more of the predetermined topics from said database; and

an evaluation means executing on a computer for evaluating for possible inclusion in a new question set for presentation to test candidates the selected questions relative to predetermined requirements applied to such prior use information on the selected questions and their corresponding answers in said data set, and discarding from possible inclusion in the new question set selected questions not fulfilling said predetermined requirements, thereby to generate from the database of questions the new question set for presentation to test candidates, wherein selected questions that do not fit a chosen profile regarding selectivity and facility are discarded from the new question set, such that a controlled distribution of difficulty is obtained for the new question set;

wherein the evaluation means is adapted for discarding from the new question set the selected questions not fulfilling the predetermined requirements by fitting a number of correct answers per failure in the selected questions with a predetermined distribution, so as to ensure the controlled distribution of difficulty in the new question set; and

wherein the evaluation means is further adapted for:

finding a coefficient indicating a quality of questions in the database, the database including the questions associated with prior use information and the unused questions including newly added questions, by calculating an actual reliability of the new question set after the new question set has been in use and comparing the actual reliability of the new question set after it has been in use to a calculated predicted reliability to find the coefficient, based on:

constructing a set of probable question set answers, and, with the set of probable question set answers, calculating a measure for predicted reliability of the new question set, assuming the new question set to have a same distribution of characteristics as a known set of the questions having known reliability, and

correcting the predicted reliability of the new question set, relative to the actual reliability of the new question set after the new question set has been in use, wherein the empirical data are used to discriminate the questions and produce an exam question set with higher quality than a randomly selected set.

2. The system according to claim 1 , wherein the data set comprises information regarding quality of the corresponding answers to each question, and wherein the quality of the corresponding answers is related to whether generally good candidates of the test candidates providing an answer answered the question correctly.

3. The system according to claim 2 , wherein the information regarding quality of the corresponding answers comprises a deviation in said quality relative to a complete exam question set of which each answer was a part.

4. The system according to claim 1 , wherein the prior use information comprises a ratio of correct answers to each question.

5. The system according to claim 1 , wherein the evaluation means is further adapted to discard selected questions not fulfilling the predetermined requirements, said predetermined requirements comprising:

a correlation between quality of the answers given to the questions and the quality of an exam question set in which the answer is provided, a deviation in the quality of the answers given to the questions being above a predetermined value to ensure that the selected questions are suitable for distinguishing between test candidates of different skill levels.

6. The system according to claim 1 , wherein the system is arranged to use a Kuder-Richardson measure (KR 20) or a Cronbach's Alpha measure to measure overall test-quality of an exam question set comprising previously used questions in the database.

7. The system according to claim 1 , wherein the system is arranged to use a Spearman-Brown prediction formula to measure predicted reliability of the new question set according to:

ρ

xx

p

=

N

ρ

xx

1

+

(

N

-

1

)

ρ

xx

where N is a multiple of the number of items in a known good subset of the new questions set, found by

N

=

N

T

N

R

(where N T is the number of items in the new question set, while N R is the number of items in the subset), and ρ xx′ is reliability of the subset.

8. A computer-based method for improving computer-generated exams and generating exam question sets suited to separating test candidates of different skill levels, the method comprising the steps of:

storing a number of questions in at least one predetermined topic in a database, each question being associated with a data set comprising prior use information related to answers given to the question in previous tests presented to test candidates, said database being stored on at least one data storage device and updated to reflect empirical data collected from previous exams to find selectivity, facility and reliability of question and question sets presented to test candidates, wherein the empirical data for each question on the computer generated exams is collected to find the selectivity, reliability and facility of the questions and the database is updated to reflect the collected empirical data, said database further comprising unused questions including newly added questions;

randomly selecting by means executing on a computer questions within one or more of the predetermined topics from said database;

evaluating by means executing on a computer, for possible inclusion in a new question set for presentation to test candidates, the selected questions relative to predetermined requirements applied to such prior use information on the selected questions and their corresponding answer alternatives in said data set; and

discarding from possible inclusion in the new question set selected questions not fulfilling said predetermined requirements, thereby generating from the database of questions the new question set for presentation to test candidates, wherein selected questions that do not fit a chosen profile regarding selectivity and facility are discarded from the new question set, such that a controlled distribution of difficulty is obtained;

wherein the evaluating means is adapted for discarding the selected questions not fulfilling the predetermined requirements by fitting a number of correct answers per failure in the selected questions with a predetermined distribution, so as to ensure the controlled distribution of difficulty in the new question set; and

wherein the evaluating means is further adapted to find a coefficient indicating a quality of the questions in the database, the database including the questions associated with prior use information and the unused questions including newly added questions, by calculating an actual reliability of the new question set after the new question set has been in use and comparing the actual reliability of the new question set after it has been in use to a calculated predicted reliability to find the coefficient; based on:

constructing a set of probable question set answers, and, with the set of probable question set answers, calculating a measure for predicted reliability of the new question set, assuming the new question set to have a same distribution of characteristics as a known set of the questions having known reliability; and

correcting the predicted reliability of the new question set, relative to the actual reliability of the new question set after the new question set has been in use, wherein the empirical data are used to discriminate the questions and produce an exam question set with higher quality than a randomly selected set.

9. The method according to claim 8 , wherein the data set comprises information regarding quality of the corresponding answers to each question, and wherein the quality of the corresponding answers is related to whether generally good candidates of the test candidates providing an answer answered the question correctly.

10. The method according to claim 9 , wherein the information regarding quality of the corresponding answers comprises a deviation in said quality relative to a complete exam question set of which each answer was a part.

11. The method according to claim 8 , wherein the prior use information comprises a ratio of correct answers to each question.

12. The method according to claim 8 , wherein the evaluation means is further adapted to discard selected questions not fulfilling the predetermined requirements, said predetermined requirements comprising:

a correlation between quality of the answers given to the questions and the quality of an exam question set in which the answer is provided, a deviation in the quality of the answers given to the questions being above a predetermined value, so as to ensure that the selected questions are suitable for distinguishing between test candidates of different skill levels.

13. The method according to claim 8 , wherein the method uses a Kuder-Richardson measure (KR 20) or a Cronbach's Alpha measure to measure overall test-quality of a question set comprising previously used questions in the database.

14. The method according to claim 8 , wherein the method uses a Spearman-Brown prediction formula to measure predicted reliability of the new question set according to:

ρ

xx

p

=

N

ρ

xx

1

+

(

N

-

1

)

ρ

xx

where N is a multiple of the number of items in a known good subset of the new set, found by

N

=

N

T

N

R

(where N T is the number of items in the new question set, while N R is the number of items in the subset) and ρ xx′ is reliability of the subset.

15. A test generation method for improving computer-generated exams by generating exam question sets suited to separating test candidates of different skill levels, the method comprising:

storing a number of questions in a plurality of topics on a computer database, wherein each question is associated with a data set of prior use information related to answers given to the question in previous tests presented to test candidates and updating the database to reflect empirical data collected from previous exams to find selectivity, facility and reliability of question and question sets presented to test candidates, wherein the empirical data for each question on the computer generated exams is collected to find the selectivity, reliability and facility of the questions and the database is updated to reflect the collected empirical data, said database further comprising unused questions including newly added questions;

randomly selecting, with a computer system, questions within at least one of the topics from the database; and

evaluating, with the computer system, the selected questions for inclusion in a new question set relative to predetermined requirements applied to the prior use information on the selected questions and their corresponding answers in the data set;

wherein selected questions not fulfilling the predetermined requirements are discarded from the new question set and selected questions that do not fit a chosen profile regarding selectivity and facility are discarded, such that a controlled distribution of difficulty is obtained;

wherein the evaluation means is adapted to discard selected questions not fulfilling the predetermined requirements by fitting a number of correct answers per failure in the selected questions with a predetermined distribution, so as to ensure the controlled distribution of difficulty in the new question set; and

wherein the evaluation means is further adapted to find a coefficient indicating a quality of questions in the database, the database including the questions associated with prior use information and the unused questions including newly added questions, by calculating an actual reliability of the new question set and comparing the actual reliability of the new question set after it has been in use to a calculated predicted reliability to find the coefficient, based on

constructing a set of probable question answers, and, with the set of probable question set answers, calculating a measure for predicted reliability of the new question set, assuming the new question set to have a same distribution of characteristics as a known set of the questions having known reliability; and

correcting the predicted reliability of the new question set, relative to the actual reliability of the new question set after the new question set has been in use, wherein the empirical data are used to discriminate the questions and produce an exam question set with higher quality than a randomly selected set.

16. The method of claim 15 , wherein the data set of prior use information comprises information regarding quality of the corresponding answers given to each question, and wherein the quality of the corresponding answers is related to whether generally good candidates of the test candidates providing an answer answered the question correctly.

17. The method of claim 15 , wherein the prior use information comprises a ratio of correct answers to each question.

18. The method of claim 15 , further comprising using a Kuder-Richardson or Cronbach's Alpha measure to measure an overall test-quality of previously used questions in the database.

19. The method of claim 15 , further comprising using a Spearman-Brown prediction formula to predict reliability of the new question set according to:

ρ

xx

p

=

N

ρ

xx

1

+

(

N

-

1

)

ρ

xx

,

in which N is a multiple of the number of items in a known good subset of the new question set, found by:

N

=

N

T

N

R

(in which N T is the number of items in the new question set, while N R is the number of items in the subset), and ρ xx′ is a reliability of the subset.

20. The method of claim 15 , wherein the database contains information about who contributed the questions, and wherein an overall quality of questions from a workgroup is estimated based on origin.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2011
From: BJORNAR NIBE, KJELL; MIKKELSEN, EIRIK
To: RELIANT EXAMS AS
Reel/Frame 027100/0451 →
Priority Claims (1)
NO 2006 4831 · Oct 25, 2006 · national
Continuity (2)
Continuation PCTNO2007000366 · Oct 24, 2007
Related Publication 20100203492A1 · Aug 12, 2010
Cited By (2)
US 12,277,870 US 12,688,368