IP Library › Granted Patent US 9,678,946
Granted Patent B2
US 9,678,946 · App. 14/793,677 · Granted Jun 13, 2017

Automatic generation of N-grams and concept relations from linguistic input data

Inventors: Fabrice Nauze (Amsterdam, NL); Christian Kissig (Amsterdam, NL); Madalina Zarafin (Bucharest, RO); Maria Begona Villada-Moiron (Almere, NL); Roos Genet (Leiden, NL)
Assignee: ORACLE INTERNATIONAL CORPORATION
G06F17/277G06F17/2735G06F17/2755G06F17/2775G06F17/28G06F17/30734
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,678,946
App. No.
14/793,677
Granted
Jun 13, 2017
Kind
B2
Abstract

A method of automatically generating a lemma dictionary from a web resource may include extracting a plurality of tokens from text-based documents within the web resource, and generating a plurality of N-grams from the plurality of tokens. The method may additionally include receiving one or more filter definitions that identify valid N-grams, and filtering the plurality of N-grams using the one or more filter definitions to generate a lemma dictionary. The method may further include generating an ontology that comprises the lemma dictionary.

Claims (45)

1. A method of automatically generating a lemma dictionary from a web resource, the method comprising:

extracting a plurality of tokens from text-based documents within the web resource;

generating a plurality of N-grams from the plurality of tokens;

receiving one or more filter definitions that identify valid N-grams;

filtering the plurality of N-grams using the one or more filter definitions to generate a lemma dictionary;

generating an ontology that comprises the lemma dictionary;

causing a user interface to display a network of nodes representing the lemma dictionary; and

receiving one or more inputs that assign ontology relationships to the network of nodes.

2. The method of claim 1 wherein extracting the plurality of tokens from the text-based documents comprises identifying and eliminating structural and formatting text that is thereby excluded from the plurality of tokens.

3. The method of claim 1 wherein the web resource comprises a web domain, and wherein the web domain comprises a plurality of HTML webpages.

4. The method of claim 1 wherein generating the plurality of N-grams comprises generating word combinations as they appear in the web resource.

5. The method of claim 1 further comprising:

causing a user interface to be displayed after filtering the plurality of N-grams;

receiving input that eliminates at least one of the N-grams in the lemma dictionary.

6. The method of claim 1 wherein the one or more filter definitions comprises a part-of-speech filter for individual tokens in an N-gram.

7. The method of claim 1 wherein the one or more filter definitions comprises a text pattern.

8. The method of claim 1 wherein the one or more filter definitions comprises a minimum frequency for an N-gram to appear in the web resource.

9. The method of claim 1 wherein the one or more filter definitions comprises a selection of a language.

10. A non-transitory, computer-readable medium comprising instructions which, when executed by one or more processors, causes the one or more processors to perform operations comprising:

extracting a plurality of tokens from text-based documents within the web resource;

generating a plurality of N-grams from the plurality of tokens;

receiving one or more filter definitions that identify valid N-grams;

filtering the plurality of N-grams using the one or more filter definitions to generate a lemma dictionary;

generating an ontology that comprises the lemma dictionary;

causing a user interface to display a network of nodes representing the lemma dictionary; and

receiving one or more inputs that assign ontology relationships to the network of nodes.

11. The non-transitory, computer-readable medium of claim 10 wherein generating the plurality of N-grams comprises generating word combinations as they appear in the web resource.

12. The non-transitory, computer-readable medium of claim 10 wherein the one or more filter definitions comprises a part-of-speech filter for individual tokens in an N-gram.

13. The non-transitory, computer-readable medium of claim 10 wherein the one or more filter definitions comprises a text pattern.

14. The non-transitory, computer-readable medium of claim 10 wherein the one or more filter definitions comprises a minimum frequency for an N-gram to appear in the web resource.

15. The non-transitory, computer-readable medium of claim 10 wherein the one or more filter definitions comprises a selection of a language.

16. A system comprising:

one or more processors; and

one or more memory devices comprising instructions which, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

extracting a plurality of tokens from text-based documents within the web resource;

generating a plurality of N-grams from the plurality of tokens;

receiving one or more filter definitions that identify valid N-grams;

filtering the plurality of N-grams using the one or more filter definitions to generate a lemma dictionary;

generating an ontology that comprises the lemma dictionary;

causing a user interface to display a network of nodes representing the lemma dictionary; and

receiving one or more inputs that assign ontology relationships to the network of nodes.

17. The system of claim 16 wherein generating the plurality of N-grams comprises generating word combinations as they appear in the web resource.

18. The system of claim 16 wherein the one or more filter definitions comprises a part-of-speech filter for individual tokens in an N-gram.

19. The system of claim 16 wherein the one or more filter definitions comprises a minimum frequency for an N-gram to appear in the web resource.

20. The system of claim 16 wherein the one or more filter definitions comprises a selection of a language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2015
From: NAUZE, FABRICE; KISSIG, CHRISTIAN; ZARAFIN, MADALINA; VILLADA-MOIRON, MARIA BEGONA; GENET, ROOS
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 036060/0547 →
Continuity (3)
Provisional Application 62077868 · Nov 10, 2014
Provisional Application 62077887 · Nov 10, 2014
Related Publication 20160132484A1 · May 12, 2016