IP Library Granted Patent US 7,219,098
Granted Patent B2
US 7,219,098 · App. 10/045,064 · Granted May 15, 2007

System and method for processing data in a distributed architecture

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,219,098
App. No.
10/045,064
Granted
May 15, 2007
Kind
B2
Abstract

A system, method, and processor readable medium for processing data in a knowledge management system gathers information content and transmits a work request for the information content gathered. The information content may be registered with a content map and assigned a unique document identifier. A work queue processes the work requests. The processed information may then be transmitted to another work queue for further processing. Further processing may include categorization, full-text indexing, metrics extraction or other process. Control messages may be transmitted to one or more users providing a status of the work request. The information may be analyzed and further indexed. A progress statistics report may be generated for each of the processes performed on the document. The progress statistics may be provided in a record. A shared access to a central data structure representing the metrics history and taxonomy may be provided for all work queues via a CORBA service.

Claims (93)

1. A method for processing data in a distributed architecture, the method comprising the steps of:

receiving a work request that identifies at least one repository for processing, wherein the at least one identified repository is included in a plurality of repositories;

determining a repository type of the at least one repository;

determining a spider type for gathering information content from the at least one identified repository, wherein the spider type is determined based on the repository type;

gathering information content from the at least one repository identified in the work request;

registering the information content;

assigning the information content to at least one document identifier;

transmitting the work request regarding the information content to a first work queue;

processing the work request by generating a meta-document representation of at least a portion of the information content;

transmitting the meta-document representation to a second work queue; and

analyzing the meta-document representation.

2. The method of claim 1 , wherein the meta-document representation comprises extensible markup language (XML) format.

3. The method of claim 1 , wherein the step of analyzing the meta-document representation comprises metrics extraction of the meta-document representation.

4. The method of claim 1 , wherein the step of analyzing the meta-document representation comprises:

indexing the meta-document representation.

5. The method of claim 1 , further comprising the step of:

generating progress statistics regarding the step of analyzing the meta-document representation.

6. The method of claim 5 , further comprising the step of:

transmitting the progress statistics to a third work queue.

7. The method of claim 1 , wherein the first work queue and the second work queue share access to a central data structure.

8. The method of claim 7 , wherein access is shared via a CORBA service.

9. The method of claim 7 , wherein the central data structure represents at least one of a metrics history and taxonomy regarding the information content.

10. A system for processing data in a distributed architecture, the system comprising:

a scheduling module that identifies at least one repository for processing, determines a repository type of the at least one repository, determines a spider type for gathering information content from the at least one identified repository, and generates a work request that identifies the at least one repository for processing,

wherein the spider type is determined based on the repository type, and the repository is included in a plurality of repositories;

an information content gathering module that gathers information content from the repository identified in the work request;

a registering module that registers the information content;

an assigning module that assigns the information content at least one document identifier;

a work request transmitting module that transmits the work request regarding the information content to a first work queue;

a work request processing module that processes the work request by generating a meta-document representation of at least a portion of the information content;

an information content transmitting module that transmits the meta-document representation to a second work queue; and

an information content processing module that analyzes the meta-document representation.

11. The method of claim 1 wherein the step of analyzing the meta-document representation comprises categorizing the meta-document representation.

12. The system of claim 10 , wherein the meta-document representation comprises extensible markup language (XML) format.

13. The system of claim 10 , wherein the information content processing module that analyzes the meta-document representation comprises:

a metrics extraction module that performs metrics extraction on the meta-document representation.

14. The system of claim 10 , wherein the information content processing module that analyzes the meta-document representation comprises:

an indexing module that indexes the meta-document representation.

15. The system of claim 10 , further comprising:

a generating module that generates progress statistics regarding the analyzing of the meta-document representation.

16. The system of claim 15 , further comprising:

a progress statistics transmitting module that transmits the progress statistics to a third work queue.

17. The system of claim 10 , wherein the first work queue and the second work queue share access to a central data structure.

18. The system of claim 17 , wherein access is shared via a CORBA service.

19. The system of claim 17 , wherein the central data structure represents at least one of a metrics history and taxonomy regarding the information content.

20. The system of claim 10 , wherein the information content processing module that analyzes the meta-document representation comprises a categorizing module that categorizes the meta-document representation.

21. A system for processing data in a distributed architecture, the system comprising:

scheduling means for identifying at least one repository for processing, determining a repository type of the at least one repository, determining a spider type for gathering information content from the at least one identified repository, and generating a work request that identifies the at least one repository for processing,

wherein the spider type is determined based on the repository type, and the repository is included in a plurality of repositories;

gathering means for gathering information content from the repository identified in with the work request;

registering means for registering the information content;

assigning means for assigning the information content at least one document identifier;

work request transmitting means for transmitting the work request regarding the information content to a first work queue;

work request processing means for processing the work request by generating a meta-document representation of at least a portion of the information content;

information content transmitting means for transmitting the meta-document representation to a second work queue; and

information content processing means for analyzing the meta-document representation.

22. The system of claim 21 , wherein the meta-document representation comprises extensible markup language (XML) format.

23. The system of claim 21 , wherein the information content processing means for analyzing the meta-document representation comprises:

metrics extraction means for performing metrics extraction on the meta-document representation.

24. The system of claim 21 , wherein the information content processing means for analyzing the meta-document representation comprises:

indexing means for indexing the meta-document representation.

25. The system of claim 21 , further comprising:

progress statistics generating means for generating progress statistics regarding the analyzing of the meta-document representation.

26. The system of claim 25 , further comprising:

progress statistics transmitting means for transmitting the progress statistics to a third work queue.

27. The system of claim 21 , wherein the first work queue and the second work queue share access to a central data structure.

28. The system of claim 27 , wherein access is shared via a CORBA service.

29. The system of claim 27 , wherein the central data structure represents at least one of a metrics history and taxonomy regarding the information content.

30. The system of claim 21 , wherein the information content processing means for analyzing the meta-document representation comprises categorization means for categorizing the meta-document representation.

31. A processor readable medium comprising processor readable code embodied therein for causing a processor to process data in a distributed architecture, the medium comprising:

work request receiving code that causes a processor to receive a work request that identifies at least one repository for processing, wherein the at least one identified repository is included in a plurality of repositories;

repository type determining code that causes a processor to determine a repository type of the at least one repository;

spider type determining code that causes a processor to determine a spider type for gathering information content from the at least one identified repository, wherein the spider type is determined based on the repository type;

information content gathering code that causes a processor to gather information content from the repository identified in the work request;

registering code that causes a processor to register the information content;

assigning code that causes a processor to assign the information content at least one document identifier;

work request transmitting code that causes a processor to transmit the work request regarding the information content to a first work queue;

work request processing code that causes a processor to process the work request by generating a meta-document representation of at least a portion of the information content;

information content transmitting code that causes a processor to transmit the meta-document representation to a second work queue; and

information content processing code that causes a processor to analyze the meta-document representation.

32. The medium of claim 31 , wherein the meta-document representation comprises extensible markup language (XML) format.

33. The medium of claim 31 , wherein the information content processing code that causes a processor to analyze the meta-document representation comprises:

categorizing code that causes a processor to categorize the meta-document representation.

34. The medium of claim 31 , wherein the information content processing code that causes a processor to analyze the meta-document representation comprises:

indexing code that causes a processor to index the meta-document representation.

35. The medium of claim 31 , further comprising:

generating code that causes a processor to generate progress statistics regarding the analyzing of the meta-document representation.

36. The medium of claim 35 , further comprising:

progress statistics transmitting code that causes a processor to transmit the progress statistics to a third work queue.

37. The medium of claim 31 , wherein the first work queue and the second work queue share access to a central data structure.

38. The medium of claim 37 , wherein access is shared via a CORBA service.

39. The medium of claim 37 , wherein the central data structure represents at least one of a metrics history and taxonomy regarding the information content.

40. The medium of claim 31 , wherein the information content processing code that causes a processor to analyze the meta-document representation comprises categorization code that categorizes the meta-document representation.

Assignments (2)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044127/0735 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2011
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: GOOGLE INC.
Reel/Frame 026664/0866 →