IP Library Granted Patent US 10,325,517
Granted Patent B2
US 10,325,517 · App. 15/847,665 · Granted Jun 18, 2019

Systems and methods for extracting keywords in language learning

Inventors: Katharine Nielson (Richmond, VA); Na'im Tyson (Mount Vernon, NY); Andrew Breen (New York, NY); Kasey Kirkham (New York, NY)
Assignee: VOXY, Inc.
G09B19/06A61B5/048A61B5/162A61B5/165A61B5/411A61B5/486G06F17/274G06F17/2705G06F17/275G09B5/00G09B5/02G09B7/08A61B5/0002A61B5/0476A61B5/04842A61B5/1124A61B5/4064A61B5/4076A61B5/4088A61B5/4842
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,325,517
App. No.
15/847,665
Filed
Dec 19, 2017
Granted
Jun 18, 2019
Kind
B2
Art Unit
2658
USPC
704/9
Abstract

Systems, methods, and products for language learning may extract text from various resources having text, using various natural-language processing features, which can be combined with custom-designed learning activities to offer a needs-based, adaptive learning methodology. The system may receive a resource, extract keywords pedagogically valuable to non-native language learning and academic exercises. Metadata describing various aspects of resources from which keywords are extracted may be associated with keywords. Metadata describing various aspects of keywords may also be associated with keywords. Extracted keywords may be stored into a keyword store along with any metadata associated with keywords.

Claims (49)

1. A computer-implemented method for extracting keywords from a text comprising:

parsing, by a computer, a set of one or more potential keywords from the text of a resource containing text;

storing, by the computer, into a keyword store, each potential keyword in the set of the one or more potential keywords matching a term in a computer file containing a keyword whitelist and each potential keyword in the set of the one or more potential keywords matching a collocation in the keyword whitelist;

determining, by the computer, for one or more potential keywords in the set of one or more potential keywords, a word frequency score associated with each of the one or more potential keywords with keys of n-grams according to one or more co-occurrence statistics based on one or more scoring rules associated with measuring each of the one or more potential keywords frequency in the text;

determining, by the computer, a pedagogical value threshold based upon a proficiency level of a learner and a resource difficulty score of the resource;

tagging, by the computer, each potential keyword having a determined word frequency score that satisfies the pedagogical value threshold with a metadata tag indicating the keyword attributes associated with the potential keyword;

storing, by the computer, into the keyword store, the metadata tag associated with each potential keyword having the determined word frequency score that satisfies the pedagogical value threshold;

generating, by the computer, potential keywords by assigning a difficulty score to each potential keyword associated with an esoteric or unique definition based on a number of letters and syllables; and

generating, by the computer, a learning activity executed by the computer for the learner using distractors that are configured to test the potential keywords and based on attributes of the learner.

2. The method according to claim 1 , further comprising discarding, by the computer, from the set of one or more potential keywords each potential keyword matching a filtered word in a file containing a stop word list.

3. The method according to claim 2 , wherein the filtered word in the stop word list is selected from the group consisting of:

a proper noun, an ordinal number, a number, a preposition, and a conjunction.

4. The method according to claim 1 , further comprising ranking, by the computer, the potential keywords in the set of one or more potential keywords according to an extraction score determined for each respective potential keyword.

5. The method according to claim 1 , further comprising discarding, by the computer, from the set of one or more potential keywords each potential keyword not satisfying the keyword frequency score threshold.

6. The method according to claim 1 , further comprising:

calculating, by the computer, a word difficulty score for a potential keyword; and

associating, by the computer, with metadata of the potential keyword containing the word difficulty score of the potential keyword.

7. The method according to claim 1 , further comprising

identifying, by the computer, one or more word attributes associated with a potential keyword; and

associating, by the computer, with metadata of the potential keyword indicating each of the one or more word attributes of the potential keyword.

8. The method according to claim 7 , wherein a word attribute in the one or more word attributes is selected from the group consisting of: a word length, a frequency of use of the word in the text, a part-of-speech, a number of syllables, and a word spelling.

9. The method according to claim 7 , further comprising identifying, by the computer, a term frequency-inverse document frequency (TF-IDF) score for the potential keyword, wherein a word attribute in the one or more word attributes is the identified TF-IDF score of the potential keyword.

10. The method according to claim 1 further comprising transmitting, by the computer, each of the potential keywords stored into the keyword store to a computing device of a content curator.

11. A system comprising a processor and non-transitory machine-readable storage containing a keyword extractor module instructing the processor to execute the steps of:

parsing text from a resource into a set of one or more potential keywords;

identifying one or more collocations in the set of one or more potential keywords matching a collocation in a file containing a keyword whitelist;

determining for each potential keyword with key of n-grams in the set of one or more potential keywords, a word frequency score according to one or more co-occurrence statistics based upon one or more scoring rules associated with measuring the potential keyword frequency in the text;

determining a pedagogical value threshold based upon a proficiency level of a learner and a resource difficulty score of the resource;

storing a set of one or more extracted keywords into a keyword store, wherein the set of one or more extracted keywords comprises each potential keyword having a word frequency score satisfying the pedagogical value threshold and each identified collocation;

generate potential keywords by assigning a difficulty score to each potential keyword associated with an esoteric or unique definition based on a number of letters and syllables; and

generate a learning activity executed by the computer for the learner using distractors that are configured to test the potential keywords and based on attributes of the learner.

12. The system according to claim 11 , further comprising:

extracting and parsing the potential keywords from the text of a document using natural language processing;

identifying one or more word attributes associated with each potential keyword; and

storing the one or more word attributes in the keyword store.

13. The system according to claim 12 , further comprising:

re-scoring each potential keyword according to the one or more co-occurrence statistics; and

removing each potential keyword falling below a usage frequency threshold from the set of potential keywords.

14. The system according to claim 12 , further comprising:

tagging each of the potential keywords with one or more metadata tags indicating the word attributes associated with the potential keyword; and

storing the metadata tags in the keyword store.

15. The system according to claim 11 , further comprising:

identifying content of the text that is associated with a potential keyword; and

storing the identified content in the keyword store.

16. The system according to claim 11 , wherein the one or more collocations in the keyword whitelist are grouped into one or more difficulty levels.

17. The system according to claim 11 , further comprising determining a word difficulty score for each potential keyword.

18. The system according to claim 11 , further comprising discarding each potential keyword from the set of potential keywords matching a filtered word in a file having a stop word list.

19. The system according to claim 18 , wherein the filtered word in the stop word list is selected from the group consisting of: an ordinal number, a number, a proper noun, an article, a preposition, and a conjunction.

20. The system according to claim 11 , wherein n-gram is a sequence of one or more words, n, recited within a unit of meaning in the text, and wherein the n-gram is selected from the group consisting of: unigrams, bigrams, and trigrams.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2025
From: ESPRESSO CAPITAL LTD
To: VOXY HOLDINGS LTD.
Reel/Frame 073002/0614 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 21, 2022
From: VOXY, INC.
To: ESPRESSO CAPITAL LTD.
Reel/Frame 059747/0542 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2017
From: NIELSON, KATHARINE; TYSON, NA'IM; BREEN, ANDREW; KIRKHAM, KASEY
To: VOXY, INC.
Reel/Frame 044441/0398 →
Continuity (4)
Continuation 15042427 · Feb 12, 2016
Continuation 14180885 · Feb 14, 2014
Provisional Application 61765105 · Feb 15, 2013
Related Publication 20180108273A1 · Apr 19, 2018
Cited By (2)
US 12,321,694 US 12,670,196