IP Library Granted Patent US 7,370,061
Granted Patent B2
US 7,370,061 · App. 11/204,061 · Granted May 6, 2008

Method for querying XML documents using a weighted navigational index

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,370,061
App. No.
11/204,061
Granted
May 6, 2008
Kind
B2
Abstract

A technique for optimizing the archival and management of data stored as XML documents is capable of handling mixed data including highly structured data and unstructured data. The technique maps the structured data to a relational database while storing the unstructured data in its native XML format. The data is updated using a rules database that maps updating rules against attributes and classes of elements within the documents. A document checking/validation engine performs the updates based on rule verification. A search engine searches the documents using both a path index table and a weighted content index.

Claims (32)

1. A method for accessing a collection of mark-up language documents having a document type definition (DTD) defining the documents of the collection, the method comprising the steps of:

mapping nodes of the DTD to create a path index table entry in a path index;

for each document in the collection, performing the steps of:

creating a document object model (DOM) tree for the document, the DOM tree including terminal nodes;

identifying keywords in the terminal nodes; and

for each keyword in each terminal node, creating a weighted content index entry in a weighted content index, the entry including an a-priori probability of a query with that keyword for that terminal node;

performing a query of the collection of documents using the path index and the weighted content index to obtain query results; and

updating the path index and the weighted content index based on the query results.

2. The method of claim 1 , further comprising the step of:

computing the a-priori probability of a query with that keyword for that terminal node by calculating a frequency for the terminal node and keyword.

3. The method of claim 1 , wherein the weighted content index entry is in the form:

F(DocID, NodeID, LevelID, ElementType, KeywdFreq, Probability),

wherein DocID is a document identification, NodeID is a node identification, LevelID is a hierarchy level, ElementType is an element type, KeywdFreq is a frequency of the keyword in the terminal node and Probability is said a-priori probability of a query with that keyword for that terminal node.

4. The method of claim 1 , wherein the step of performing a query of the collection of documents using the path index and the weighted content index to obtain query results further comprises the step of:

referring to the path index to determine the paths to search; and

conducting a constrained search of the determined paths using the weighted content index.

5. A computer program product comprising a computer readable recording medium having recorded thereon a computer program comprising code means for, when executed on a computer, instructing said computer to control steps in a method for accessing a collection of mark-up language documents having a document type definition (DTD) defining documents of the collection, the method comprising the steps of:

mapping nodes of the DTD to create a path index table entry in a path index;

for each document in the collection, performing the steps of:

creating a document object model (DOM) tree for the document, the DOM tree including terminal nodes;

identifying keywords in the terminal nodes; and

for each keyword in each terminal node, creating a weighted content index entry in a weighted content index, the entry including an a-priori probability of a query with that keyword for that terminal node;

performing a query of the collection of documents using the path index and the weighted content index to obtain query results; and

updating the path index and the weighted content index based on the query results.

6. The computer program product of claim 5 , further comprising the step of:

computing the a-priori probability of a query with that keyword for that terminal node by calculating a frequency for the terminal node and keyword.

7. The computer program product of claim 5 , wherein the weighted content index entry is in the form:

F(DocID, NodeID, LevelID, ElementType, KeywdFreq, Probability),

wherein DocID is a document identification, NodeID is a node identification, LevelID is a hierarchy level, Element Type is an element type, KeywdFreq is a frequency of the keyword in the terminal node and Probability is said a-priori probability of a query with that keyword for that terminal node.

8. The computer program product of claim 5 , wherein the step of performing a query of the collection of documents using the path index and the weighted content index to obtain query results further comprises the step of:

referring to the path index to determine the paths to search; and

conducting a constrained search of the determined paths using the weighted content index.

Assignments (3)
MERGER Recorded Apr 5, 2010
From: SIEMENS CORPORATE RESEARCH, INC.
To: SIEMENS CORPORATION
Reel/Frame 024185/0042 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2005
From: HSU, LIANG H.
To: SIEMENS CORPORATE RESEARCH, INC.
Reel/Frame 016679/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2005
From: CHAKRABORTY, AMIT; LIU, ZHIJING; LIU, PEIYA
To: SIEMENS CORPORATE RESEARCH, INC.
Reel/Frame 016679/0392 →