IP Library Granted Patent US 9,262,935
Granted Patent B2
US 9,262,935 · App. 14/180,885 · Granted Feb 16, 2016

Systems and methods for extracting keywords in language learning

Inventors: Katharine Nielson (Richmond, VA); Kasey Kirkham (New York, NY); Na'im Tyson (Mount Vernon, NY); Andrew Breen (New York, NY)
Assignee: VOXY, Inc.
G09B5/00G06F17/275G09B5/02G09B7/08G09B19/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,262,935
App. No.
14/180,885
Granted
Feb 16, 2016
Kind
B2
Abstract

Systems, methods, and products for language learning that may extract text from various resources having text, using various natural-language processing features, which can be combined with custom-designed learning activities to offer a needs-based, adaptive learning methodology. The system may receive a resource, extract keywords pedagogically valuable to non-native language learning and academic exercises. Metadata describing various aspects of resources from which keywords are extracted may be associated with keywords. Metadata describing various aspects of keywords may also be associated with keywords. Extracted keywords may be stored into a keyword store along with any metadata associated with keywords.

Claims (41)

1. A computer-implemented method for extracting keywords from text comprising:

parsing, by a computer, a set of one or more potential keywords from text of a resource containing text;

storing, by the computer, into a keyword store each potential keyword in the set matching a term in a computer file containing a keyword whitelist and each potential keyword in the set matching a collocation in the keyword whitelist;

determining, by the computer, for one or more potential keywords in the set of potential keywords, a word difficulty value associated with each of the one or more potential keywords based on scoring rules determining the word difficulty value;

determining, by the computer, a pedagogical value threshold based upon a proficiency level of a learner and a resource difficulty score of the resource; and

storing, by the computer, into the keyword store each potential keyword having a determined word difficulty value that satisfies the pedagogical value threshold.

2. The method according to claim 1 , further comprising discarding, by the computer, from the set of potential keywords each potential keyword matching a filtered word in a file containing a stop word list.

3. The method according to claim 1 , further comprising discarding, by the computer, from the set of potential keywords each potential keyword having a word difficulty value not satisfying the pedagogical value threshold.

4. The method according to claim 1 , further comprising ranking, by the computer, the potential keywords in the set of potential keywords according to an extraction score determined for each respective potential keyword.

5. The method according to claim 2 , wherein the filtered word in the stop word list is selected from the group consisting of: a proper noun, an ordinal number, a number, a preposition, and a conjunction.

6. The method according to claim 1 , further comprising:

calculating, by the computer, a word difficulty score for a potential keyword; and

associating, by the computer, with the potential keyword metadata containing the word difficulty score of the potential keyword.

7. The method according to claim 1 , further comprising

identifying, by the computer, one or more word attributes associated with a potential keyword; and

associating, by the computer, with the potential keyword metadata indicating each of the one or more word attributes of the potential keyword.

8. The method according to claim 7 , wherein a word attribute in the one or more word attributes is selected from the group consisting of: a word length, a frequency of use of the word in the text, a part-of-speech, a number of syllables, and a word spelling.

9. The method according to claim 7 , further comprising identifying, by the computer, a term frequency-inverse document frequency (TF-IDF) score for the potential keyword, wherein a word attribute in the one or more word attributes is the identified TF-IDF score of the potential keyword.

10. The method according to claim 1 further comprising transmitting, by the computer, each of the potential keywords stored into the keyword store to a computing device of a content curator.

11. A system comprising a processor and non-transitory machine-readable storage containing a keyword extractor module instructing the processor to execute the steps of:

parsing text from a resource into a set of one or more potential keywords;

identifying one or more collocations in the set of potential keywords matching a collocation in a file containing a keyword whitelist;

determining for each potential keyword in the set of one or more potential keywords, a word difficulty value according to one or more scoring rules;

determining a pedagogical value threshold based upon a proficiency level of a learner and a resource difficulty score of the resource; and

storing a set of one or more extracted keywords into a keyword store, wherein the set of extracted keywords comprises each potential keyword having a word difficulty value satisfying the pedagogical value threshold and each identified collocation.

12. The system according to claim 11 , further comprising:

extracting and parsing the potential keywords from text of the document using natural language processing;

identifying one or more word attributes associated with each potential keyword; and

storing the one or more word attributes in the keyword store.

13. The system according to claim 12 , further comprising:

re-scoring each potential keyword according to one or more co-occurrence statistics; and

removing each potential keyword falling below a usage frequency threshold from the set of potential keywords.

14. The system according to claim 12 , further comprising:

tagging each of the potential keywords with one or more metadata tags indicating the word attributes associated with the potential keyword; and

storing the metadata tags in the keyword store.

15. The system according to claim 11 , further comprising:

identifying content of the text that is associated with a potential keyword; and storing the identified content in the keyword store.

16. The system according to claim 11 , wherein the one or more collocations in the keyword whitelist are grouped into one or more difficulty levels.

17. The system according to claim 11 , further comprising determining a word difficulty score for each potential keyword.

18. The system according to claim 11 , further comprising discarding each potential keyword from the set of potential keywords matching a filtered word in a file having a stop word list.

19. The system according to claim 18 , wherein the filtered word in the stop word list is selected from the group consisting of: an ordinal number, a number, a proper noun, an article, a preposition, and a conjunction.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2025
From: ESPRESSO CAPITAL LTD
To: VOXY HOLDINGS LTD.
Reel/Frame 073002/0614 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 21, 2022
From: VOXY, INC.
To: ESPRESSO CAPITAL LTD.
Reel/Frame 059747/0542 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2014
From: NIELSON, KATHARINE; TYSON, NA'IM; BREEN, ANDREW; KIRKHAM, KASEY
To: VOXY, INC.
Reel/Frame 032220/0945 →
Continuity (2)
Provisional Application 61765105 · Feb 15, 2013
Related Publication 20140297266A1 · Oct 2, 2014