IP Library Granted Patent US 8,572,062
Granted Patent B2
US 8,572,062 · App. 12/643,343 · Granted Oct 29, 2013

Indexing documents using internal index sets

Inventors: Gregory Scott Felderman (Westminster, CO); Brian Keith Hoyt (Thornton, CO); Paula Jean Muir (Boulder, CO)
Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,572,062
App. No.
12/643,343
Granted
Oct 29, 2013
Kind
B2
Abstract

Provided are techniques for retrieving a document that includes for each page an area that is ignored by applications that process the document and that includes a different internal index set associated with each subset of pages of the document, wherein each different internal index set is associated with an area and stores indexes, and wherein each of the indexes consists of a name-value pair. Then, for each page in the document, it is determined whether the page is associated with an internal index set; and, in response to determining that the page is associated with an internal index set, one or more name-value pairs from the internal index set are extracted, wherein each of the one or more name-value pairs provides specific information about the document for use in identifying the document.

Claims (36)

1. A computer-implemented method, comprising:

obtaining, using a processor of a computer, a document with multiple subsets of pages that includes a different internal index set associated with each subset of pages from among the multiple subsets of pages, wherein each different internal index set is located within a first area on a page within the associated subset of pages, is relevant to the page and subsequent pages in the associated subset of pages until one of another internal index set within the document is found and an end of the document is reached, and includes one or more name-value pairs, and wherein the first area is ignored by an application that processes a second area of the document;

extracting the one or more name-value pairs from each different internal index set, wherein each of the one or more name-value pairs provides specific information about the document for use in identifying the document; and

storing the extracted one or more-name value pairs in a table in a database to enable subsequent searching for the document, wherein, for a name-value pair, the name corresponds to a column name in the table, and the value corresponds to a value stored in a row for a column having the name.

2. The computer-implemented method of claim 1 , wherein the extracting further comprises using Application Programming Interfaces (APIs) to extract the one or more name-value pairs.

3. The computer-implemented method of claim 1 , wherein the application is one of a document reader and a document converter.

4. The computer-implemented method of claim 1 , further comprising:

in response to receiving a search request with one or more search keys, comparing the one or more search keys to values in the database; and

in response to determining that one or more of the search keys match at least one value, providing one or more documents that are associated with one or more internal index sets that have the at least one value.

5. The computer-implemented method of claim 1 , further comprising:

in response to receiving a search request, identifying the document based on the search request matching at least one name-value pair stored in the database, wherein the at least one name-value pair is in an internal index set in the document.

6. A system, comprising:

a computer processor; and

storage coupled to the computer processor, wherein the storage stores a program, and wherein the computer processor executes the program to perform operations, wherein the operations comprise:

obtaining a document with multiple subsets of pages that includes a different internal index set associated with each subset of pages from among the multiple subsets of pages, wherein each different internal index is located within a first area on a page within the associated subset of pages, is relevant to the page and subsequent pages in the associated subset of pages until one of another internal index set within the document is found and an end of the document is reached, and includes one or more name-value pairs, and wherein the first area is ignored by an application that processes a second area of the document;

extracting the one or more name-value pairs from each different internal index set, wherein each of the one or more name-value pairs provides specific information about the document for use in identifying the document; and

storing the extracted one or more-name value pairs in a table in a database to enable subsequent searching for the document, wherein, for a name-value pair, the name corresponds to a column name in the table, and the value corresponds to a value stored in a row for a column having the name.

7. The system of claim 6 , wherein the operations for extracting further comprise using Application Programming Interfaces (APIs) to extract the one or more name-value pairs.

8. The system of claim 6 , wherein the application is one of a document reader and a document converter.

9. The system of claim 6 , wherein the operations further comprise:

in response to receiving a search request with one or more search keys, comparing the one or more search keys to values in the database; and

in response to determining that one or more of the search keys match at least one value, providing one or more documents that are associated with one or more internal index sets that have the at least one value.

10. The system of claim 6 , wherein the operations further comprise:

in response to receiving a search request, identifying the document based on the search request matching at least one name-value pair stored in the database, wherein the at least one name-value pair is in an internal index set in the document.

11. A computer program product comprising:

a non-transitory computer readable storage medium including a computer readable program, wherein the computer readable program, when executed by a processor on a computer, causes the computer to:

obtain a document with multiple subsets of pages that includes a different internal index set associated with each subset of pages from among the multiple subsets of pages, wherein each different internal index set is located within an area on a page within the associated subset of pages, is relevant to the page and subsequent pages in the associated subset of pages until one of another internal index set within the document is found and an end of the document is reached, and includes one or more name-value pairs, and wherein the first area is ignored by an application that processes a second area of the document; and

extract the one or more name-value pairs from each different internal index set, wherein each of the one or more name-value pairs provides specific information about the document for use in identifying the document; and

store the extracted one or more-name value pairs in a table in a database to enable subsequent searching for the document, wherein, for a name-value pair, the name corresponds to a column name in the table, and the value corresponds to a value stored in a row for a column having the name.

12. The computer program product of claim 11 , wherein the extracting further comprises using Application Programming Interfaces (APIs) to extract the one or more name-value pairs.

13. The computer program product of claim 11 , wherein the application is one of a document reader and a document converter.

14. The computer program product of claim 11 , wherein the computer readable program, when executed by the processor on the computer causes, the computer to:

in response to receiving a search request with one or more search keys, compare the one or more search keys to values in the database; and

in response to determining that one or more of the search keys match at least one value, provide one or more documents that are associated with one or more internal index sets that have the at least one value.

15. The computer program product of claim 11 , wherein the computer readable program, when executed by the processor on the computer causes, the computer to:

in response to receiving a search request, identify the document based on the search request matching at least one name-value pair stored in the database, wherein the at least one name-value pair is in an internal index set in the document.

Assignments (2)
CONVEYOR IS ASSIGNING UNDIVIDED 50% INTEREST Recorded Jan 11, 2018
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: SERVICENOW, INC.; INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 045060/0977 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2010
From: FELDERMAN, GREGORY S.; HOYT, BRIAN K.; MUIR, PAULA J.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 024226/0006 →
Continuity (1)
Related Publication 20110153640A1 · Jun 23, 2011