IP Library Granted Patent US 8,219,557
Granted Patent B2
US 8,219,557 · App. 12/813,354 · Granted Jul 10, 2012

System for automatically generating queries

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,219,557
App. No.
12/813,354
Granted
Jul 10, 2012
Kind
B2
Abstract

A method, system and article of manufacture therefor, are disclosed for automatically generating a query from document content.

Claims (32)

1. A computer implemented method, comprising:

accessing selected-document content that is selected from a document;

accessing a database of entities, each entity having associated therewith one or more entity-types, each entity type pertaining to one or more themes;

automatically identifying in the selected-document content at least one entity in the database of entities;

accessing a categorization system for categorizing document content having a classification profile defined using an organization of categories that corresponds to categories of document content available through an information retrieval system; the classification profile allowing document content to be assigned to an existing category;

automatically categorizing the selected-document content using the categorization system to assign a subset of one or more categories from the categories of document content available through the information retrieval system;

automatically formulating a query for document content by providing a first element of the query representative of at least one entity identified in the selected-document content, and by providing a second element of the query focusing the query to the subset of categories of document content assigned to the selected-document content;

automatically, after formulating the query, acquiring search results by using the query for querying document content available through the information retrieval system;

wherein at least said identifying, categorizing, formulating and acquiring are performed using one or more processors.

2. The method according to claim 1 , further comprising:

automatically filtering the search results;

wherein said filtering is performed using the one or more processors.

3. The method according to claim 2 , further comprising:

automatically using the search results acquired through the information retrieval system for annotating the selected-document content;

wherein said using is performed using the one or more processors.

4. The method according to claim 3 , further comprising:

making available the selected-document content annotated using the search results:

wherein the selected-document content annotated using the search results is characterized by the theme of at least one entity-type associated with at least one entity identified in the selected-document content.

5. The method according to claim 1 , wherein said identifying identifies in the selected-document content the at least one entity from proper names, times, locations, amounts, citations, and addresses.

6. The method according to claim 1 , wherein said identifying identifies in the selected-document content the at least one entity as a proper name.

7. The method according to claim 1 , wherein said identifying identifies in the selected-document content the at least one entity as a location.

8. The method according to claim 1 , wherein said identifying identifies in the selected-document content the at least one entity using one or a combination of regular expressions, lexicons, keywords, and rules.

9. The method according to claim 1 , wherein entities in the database of entities further comprise lexicons that have associated therewith an entity and an entity-type.

10. The method according to claim 9 , wherein the lexicons in the database of entities have further associated therewith an entity-string.

11. The method according to claim 10 , wherein the lexicons in the database of entities have further associated therewith a part-of-speech tag.

12. The method according to claim 1 , wherein said identifying further comprises tagging the at least one entity in the selected-document content with a part-of-speech tag that identifies grammatical usage of the at least one entity.

13. The method according to claim 1 , wherein said categorizing the selected-document content further comprises assigning the selected-document content a category vocabulary associated with each category in the set of one or more categories.

14. The method according to claim 13 , wherein said formulating further formulates the query for document content by providing a third element specified using the category vocabulary associated with each category in the set of one or more categories.

15. The method according to claim 14 , wherein said formulating further formulates the query for document content by providing a fourth element specified using one or more aspects to augment the at least one entity identified in the selected-document content.

16. The method according to claim 1 , wherein the organization of categories is organized as a hierarchy.

17. The method according to claim 1 , wherein said categorizing further comprises computing a similarity measure between the selected-document content and class profiles of categories in the categorization system.

18. The method according to claim 1 , wherein said formulating further formulates the query for document content by providing a third element specified using one or more aspects to augment the at least one entity identified in the selected-document content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2015
From: XEROX CORPORATION
To: III HOLDINGS 6, LLC
Reel/Frame 036201/0584 →