IP Library Granted Patent US 8,589,387
Granted Patent B1
US 8,589,387 · App. 13/615,811 · Granted Nov 19, 2013

Information extraction from a database

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,589,387
App. No.
13/615,811
Granted
Nov 19, 2013
Kind
B1
Abstract

Techniques for extracting information from a database are provided. A database such as the Web is searched for occurrences of tuples of information. The occurrences of the tuples of information that were found in the database are analyzed to identify a pattern in which the tuples of information were stored. Additional tuples of information can then be extracted from the database utilizing the pattern. This process can be repeated with the additional tuples of information, if desired.

Claims (59)

1. A method performed by one or more devices, the method comprising:

identifying, by the one or more devices, a tuple, from text of a plurality of documents, using a first data pattern;

identifying, by the one or more devices, second data patterns,

the first data pattern being different than each of the second data patterns;

determining, by the one or more devices, a quantity of the second data patterns that match the tuple; and

storing, by the one or more devices, the tuple in a data storage when the quantity of the second data patterns that match the tuple satisfies a particular threshold.

2. A system comprising:

one or more processors; and

one or more memories including a plurality of instructions that, when executed by the one or more processors, cause the one or more processors to:

identify a tuple, from text of a plurality of documents, using a first data pattern;

identify second data patterns,

the first data pattern being different than each of the second data patterns; determine a quantity of the second data patterns that match the tuple; and store the tuple in a data storage when the quantity of the second data patterns that

match the tuple satisfies a particular threshold.

3. A non-transitory computer-readable storage medium comprising:

one or more instructions which, when executed by at least one processor, cause the at least one processor to:

identify a tuple, from text of a plurality of documents, using a first data pattern;

identify second data patterns,

the first data pattern being different than each of the second data patterns; determine a quantity of the second data patterns that match the tuple; and

store the tuple in a data storage when the quantity of the second data patterns that match the tuple satisfies a particular threshold.

4. The method of claim 1 , further comprising:

searching the plurality of documents to identify an occurrence of another tuple in text of the plurality of documents and a context for the other tuple in the text of the plurality of documents;

analyzing the identified occurrence and the context to identify another data pattern that corresponds to the other tuple; and

identifying the tuple using the other data pattern.

5. The method of claim 1 , further comprising:

disregarding the tuple when the quantity of the second data patterns that match the tuple does not satisfy the particular threshold.

6. The method of claim 1 , further comprising:

determining a specificity of the first data pattern based on a quantity of occurrences of the tuple in the text of the plurality of documents.

7. The method of claim 1 , where the first data pattern includes first text and second text, the first text preceding information in the tuple and the second text following the information in the tuple.

8. The system of claim 2 , where the one or more processors are further to:

search the plurality of documents to identify an occurrence of another tuple in text of the plurality of documents and a context for the other tuple in the text of the plurality of documents;

analyze the identified occurrence and the context to identify another data pattern that corresponds to the other tuple; and

identify the tuple using the other data pattern.

9. The system of claim 2 , where the one or more processors are further to:

disregard the tuple when the quantity of the second data patterns that match the tuple does not satisfy the particular threshold.

10. The system of claim 2 , where the one or more processors are further to:

determine a specificity of the first data pattern based on a quantity of occurrences of the tuple in the text of the plurality of documents.

11. The system of claim 2 , where the first data pattern includes first text and second text, the first text preceding information in the tuple and the second text following the information in the tuple.

12. The medium of claim 3 , further comprising:

one or more instructions to search the plurality of documents to identify an occurrence of another tuple in text of the plurality of documents and a context for the other tuple in the text of the plurality of documents;

one or more instructions to analyze the identified occurrence and the context to identify another data pattern that corresponds to the other tuple; and

one or more instructions to identify the tuple using the other data pattern.

13. The medium of claim 3 , further comprising:

one or more instructions to disregard the tuple when the quantity of the second data patterns that match the tuple does not satisfy the particular threshold.

14. The medium of claim 3 , further comprising:

one or more instructions to determine a specificity of the first data pattern based on a quantity of occurrences of the tuple in the text of the plurality of documents.

15. The medium of claim 3 , where the first data pattern includes first text and second text, the first text preceding information in the tuple and the second text following the information in the tuple.

16. The method of claim 4 , where

the other tuple includes a plurality of fields that each represent a character string, and

searching the plurality of documents to identify the occurrence of the other tuple in the text of the plurality of documents includes identifying an occurrence of text, in the plurality of documents, that matches the character strings of the other tuple.

17. The method of claim 4 , where the tuple is different than the other tuple.

18. The system of claim 8 , where

the other tuple includes a plurality of fields that each represent a character string, and

the one or more processors, when searching the plurality of documents to identify the occurrence of the other tuple in the text of the plurality of documents, are further to:

identify an occurrence of text, in the plurality of documents, that matches the character strings of the other tuple.

19. The system of claim 8 , where the tuple is different than the other tuple.

20. The medium of claim 12 , where

the other tuple includes a plurality of fields that each represent a character string, and

one or more instructions to search the plurality of documents to identify the occurrence of the other tuple in the text of the plurality of documents include:

one or more instructions to identify an occurrence of text, in the plurality of documents, that matches the character strings of the other tuple.

Assignments (1)
CHANGE OF NAME Recorded Dec 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044695/0115 →