IP Library Granted Patent US 7,930,322
Granted Patent B2
US 7,930,322 · App. 12/127,017 · Granted Apr 19, 2011

Text based schema discovery and information extraction

Assignee: Microsoft Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,930,322
App. No.
12/127,017
Granted
Apr 19, 2011
Kind
B2
Abstract

Various technologies and techniques are disclosed for text based schema discovery and information extraction. Documents are analyzed to identify sections of the documents and a relationship between the sections. Statistics are stored regarding occurrences of items in the documents. A probabilistic model is generated based on the stored statistics. A database schema is generated with a plurality of tables based upon the probabilistic model. The documents are analyzed against the probabilistic model to determine how the documents map to the tables generated from the database schema. The tables are populated from the documents based on a result of the analysis against the probabilistic model.

Claims (35)

1. A method for creating a database schema from unstructured documents comprising the steps of:

accessing an unstructured document;

extracting information from the unstructured document using text mining, the extracted information comprising terms, phrases and sentiments;

analyzing the extracted information to identify sections of the unstructured document;

storing statistics regarding an occurrence of items in the unstructured document, the items comprising the extracted information and identified sections;

repeating the accessing, extracting, analyzing, and storing steps for a plurality of unstructured documents;

creating a probabilistic model based on the statistics stored for the plurality of unstructured documents;

generating a database schema using the probabilistic model;

receiving user modifications to the probabilistic model;

updating the probabilistic model based upon the user modifications; and

generating a database based on the database schema generated using the probabilistic model.

2. The method of claim 1 , wherein the document is analyzed using machine learning techniques.

3. The method of claim 1 , wherein the document is analyzed using heuristics.

4. The method of claim 1 , wherein the database schema is used to map data from the document to database tables.

5. The method of claim 1 , wherein the items for which statistics are stored include terms.

6. The method of claim 1 , wherein the items for which statistics are stored include fields.

7. The method of claim 1 , wherein the items for which statistics are stored include lists.

8. The method of claim 1 , wherein the items for which statistics are stored include sections.

9. The method of claim 1 , wherein the analyzing step further comprises analyzing the extracted information to identify relationships between different sections of the document.

10. The method of claim 1 , wherein during the extracting step, one or more dictionaries are utilized to aid in the extracting.

11. The method of claim 1 , further comprising the steps of:

using the probabilistic model to generate analytical objects.

12. A computer storage medium having computer-executable instructions for causing a computer to perform steps comprising:

accessing an unstructured document;

extracting information from the unstructured document using text mining, the extracted information comprising terms, phrases and sentiments;

analyzing the extracted information to identify sections of the unstructured document;

storing statistics regarding an occurrence of items in the unstructured document, the items comprising the extracted information and identified sections;

repeating the accessing, extracting, analyzing, and storing steps for a plurality of unstructured documents;

creating a probabilistic model based on the statistics stored for the plurality of unstructured documents;

generating a database schema using the probabilistic model;

receiving user modifications to the probabilistic model;

updating the probabilistic model based upon the user modifications; and

generating a database based on the database schema generated using the probabilistic model.

13. The computer storage medium of claim 12 , wherein the document is analyzed using machine learning techniques.

14. The computer storage medium of claim 12 , wherein the document is analyzed using heuristics.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034564/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2008
From: MACLENNAN, C. JAMES
To: MICROSOFT CORPORATION
Reel/Frame 021326/0055 →
Continuity (1)
Related Publication 20090300043A1 · Dec 3, 2009