IP Library › Granted Patent US 12,548,587
Granted Patent B1
US 12,548,587 · App. 18/357,257 · Granted Feb 10, 2026

Systems and methods for modeling in-the-moment lexical experience for tracking oral reading fluency

Inventors: Beata Beigman Klebanov (Hopewell, NJ); Michael Suhan (Princeton, NJ)
Assignee: Educational Testing Service
G10L25/51G10L15/22G10L15/26G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,587
App. No.
18/357,257
Granted
Feb 10, 2026
Kind
B1
Abstract

Systems and methods are provided for modeling lexical experience for tracking of oral reading fluency. In embodiments, a background corpus and a text are received. A reading passage is selected from the text. A surprisal model is generated based on the background corpus and a portion of the text preceding the reading passage. An audio of a user reciting the reading passage is iteratively received, wherein a token of the audio from a plurality of tokens is received at a time. Oral reading fluency is evaluated based on the audio and the surprisal mode. The next reading passage is selected. An oral reading fluency report is stored in a computer readable medium.

Claims (58)

1 . A computer-implemented method for modeling lexical experience for tracking of oral reading fluency comprising:

receiving a background corpus and a text, wherein the background corpus comprises general background corpora and personalized corpora from passages that a user has read;

selecting a reading passage from the text, wherein the reading passage comprises a plurality of reading passage tokens;

generating a surprisal model based on the background corpus and a portion of the text preceding the reading passage to calculate a probability of a particular next token of the reading passage;

iteratively receiving audio of the user reciting the reading passage, wherein a token of audio from a plurality of audio tokens is received at a time;

evaluating oral reading fluency based on the audio, the reading passage, and the surprisal model by comparing the plurality of audio tokens to the plurality of reading passage tokens in view of the calculated probability of each token of the plurality of reading passage tokens;

selecting a next reading passage; and

storing an oral reading fluency report in a computer readable medium.

2 . The method of claim 1 , wherein evaluating oral reading fluency further comprises using a baseline model.

3 . The method of claim 2 , wherein the baseline model measures words read correctly per minute of oral reading, text complexity, genre, prosody, or local discourse structure.

4 . The method of claim 1 , further comprising:

processing the background corpus and the text by tokenizing the background corpus and the text.

5 . The method of claim 1 , further comprising:

processing the audio which comprises:

transcribing the audio into a transcription using a speech to text transcriber; and

tokenizing the transcription.

6 . The method of claim 1 , wherein the background corpus and the text are pre-processed to normalized spelling and handle contractions and hyphenation.

7 . The method of claim 1 , wherein the probability of the particular next token is continuously updated, token by token, as the user progresses through the reading passage.

8 . The method of claim 1 , wherein the surprisal model calculates a surprisal value of the particular next token based on a logarithm of an inverse of the probability of the particular next token.

9 . The method of claim 7 , wherein the surprisal value is precomputed for every token in the reading passage.

10 . The method of claim 1 , further comprising:

providing the selected reading passage and the selected next reading passage to the user via a user interface.

11 . A system for modeling lexical experience for tracking of oral reading fluency comprising:

a processing system comprising one or more data processors; and

a computer-readable medium encoded with instructions for commanding the processing system to execute steps comprising:

receiving a background corpus and a text, wherein the background corpus comprises general background corpora and personalized corpora from passages that a user has read;

selecting a reading passage from the text, wherein the reading passage comprises a plurality of reading passage tokens;

generating a surprisal model based on the background corpus and a portion of the text preceding the reading passage to calculate a probability of a particular next token of the reading passage;

iteratively receiving audio of the user reciting the reading passage, wherein a token of audio from a plurality of audio tokens is received at a time;

evaluating oral reading fluency based on the audio, the reading passage, and the surprisal model by comparing the plurality of audio tokens to the plurality of reading passage tokens in view of the calculated probability of each token of the plurality of reading passage tokens;

selecting a next reading passage; and

storing an oral reading fluency report in a computer readable medium.

12 . The system of claim 11 , wherein evaluating oral reading fluency further comprises using a baseline model.

13 . The system of claim 12 , wherein the baseline model measures words read correctly per minute of oral reading, text complexity, genre, prosody, or local discourse structure.

14 . The system of claim 11 , wherein the steps further comprise:

processing the background corpus and the text by tokenizing the background corpus and the text.

15 . The system of claim 11 , wherein the steps further comprise:

processing the audio which comprises:

transcribing the audio into a transcription using a speech to text transcriber; and

tokenizing the transcription.

16 . The system of claim 11 , wherein the background corpus and the text are pre-processed to normalize spelling and handle contractions and hyphenation.

17 . A non-transitory computer-readable medium encoded with instructions for commanding one or more data processors to execute steps of a method for modeling lexical experience for tracking of oral reading fluency comprising:

receiving a background corpus and a text, wherein the background corpus comprises general background corpora and personalized corpora from passages that a user has read;

selecting a reading passage from the text, wherein the reading passage comprises a plurality of reading passage tokens;

generating a surprisal model based on the background corpus and a portion of the text preceding the reading passage to calculate a probability of a particular next token of the reading passage;

iteratively receiving audio of the user reciting the reading passage, wherein a token of audio from a plurality of audio tokens is received at a time;

evaluating oral reading fluency based on the audio, the reading passage, and the surprisal model by comparing the plurality of audio tokens to the plurality of reading passage tokens in view of the calculated probability of each token of the plurality of reading passage tokens;

selecting a next reading passage; and

storing an oral reading fluency report in a computer readable medium.

18 . The non-transitory computer-readable medium of claim 17 , wherein evaluating oral reading fluency further comprises using a baseline model.

19 . The non-transitory computer-readable medium of claim 18 , wherein the baseline model measures words read correctly per minute of oral reading, text complexity, genre, prosody, or local discourse structure.

20 . The non-transitory computer-readable medium of claim 17 , wherein the steps further comprise;

processing the background corpus and the text by tokenizing the background corpus and the text.

21 . The non-transitory computer-readable medium of claim 17 , wherein the steps further comprise:

processing the audio which comprises:

transcribing the audio into a transcription using a speech to text transcriber; and

tokenizing the transcription.

22 . The non-transitory computer-readable medium of claim 17 , wherein the background corpus and the text are pre-processed to normalize spelling and handle contractions and hyphenation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2023
From: BEIGMAN KLEBANOV, BEATA; SUHAN, MICHAEL
To: EDUCATIONAL TESTING SERVICE
Reel/Frame 064665/0358 →
Continuity (1)
Provisional Application 63391899 · Jul 25, 2022
References Cited (7)
US 11024194B1 · Beigman Klebanov · 2021 [cited by examiner]
Monsalve et al, (“Lexical Surprisal as a General Predictor of Reading Time,”(2012), In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics, pp. 398-408, Avignon, F… [cited by examiner]
Breland et al.; The College Board Vocabulary Study; ETS Report No. 94-26; College Entrance Examination Board; pp. 1-51; 1994. [cited by applicant]
Landauer, Thomas, Foltz, Peter, Laham, Darrell; Introduction to Latent Semantic Analysis; Discourse Processes, 25; pp. 259-284; 1998. [cited by applicant]
Terzopoulos et al.; HelexKids: A word frequency database for Greek and Cypriot primary school children; Behavior Research Methods, 49(1); pp. 83-96; 2017. [cited by applicant]
Tribus, Myron; Information Theory as the Basis for Thermostatics and Thermodynamics; Journal of Applied Mechanics, 28(1); pp. 1-8; 1961. [cited by applicant]
Zeno, Susan, Ivens, Stephen, Millard, Robert, Duvvuri, Raj; The Educator's Word Frequency Guide; Brewster, NY: Touchstone Applied Science Associates; 1995. [cited by applicant]