IP Library Granted Patent US 8,903,825
Granted Patent B2
US 8,903,825 · App. 13/478,973 · Granted Dec 2, 2014

Semiotic indexing of digital resources

Inventors: Charles T. Parker (East Lansing, MI); George M. Garrity (Okemos, MI)
Assignee: NamesforLife LLC
G06F17/30707G06F17/3071
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,903,825
App. No.
13/478,973
Granted
Dec 2, 2014
Kind
B2
Abstract

A method of classifying a plurality of documents. The method includes steps of providing a first set of classification terms and a second set of classification terms, the second set of classification terms being different from the first set of classification terms; generating a first frequency array of a number of occurrences of each term from the first set of classification terms in each document; generating a second frequency array of a number of occurrences of each term from the second set of classification terms in each document; generating a first similarity matrix from the first frequency array; generating a second similarity matrix from the second frequency array; determining an entrywise combination of the first similarity matrix and the second similarity matrix; and clustering the plurality of documents based on the result of the entrywise combination.

Claims (71)

1. A method of classifying a plurality of documents, comprising:

providing a first set of classification terms and a second set of classification terms, the second set of classification terms being different from the first set of classification terms;

generating a first frequency array of a number of occurrences of each term from the first set of classification terms in each document;

generating a second frequency array of a number of occurrences of each term from the second set of classification terms in each document;

generating a first similarity matrix from the first frequency array;

generating a second similarity matrix from the second frequency array;

determining an entrywise combination of the first similarity matrix and the second similarity matrix; and

clustering the plurality of documents based on the result of the entrywise combination.

2. The method of claim 1 , wherein the first set of classification terms comprises one of an externally managed set of terms and a patent classification code.

3. The method of claim 1 , wherein the first set of classification terms comprises an externally managed set of terms.

4. The method of claim 3 , wherein the externally managed set of terms has been disambiguated.

5. The method of claim 1 , wherein the first set of classification terms comprises an externally managed set of classification terms for Bacteria and Archaea.

6. The method of claim 1 , further comprising reordering metadata associated with the plurality of documents according to the clustering.

7. The method of claim 1 , wherein clustering the documents comprises a hierarchical clustering method selected from the group consisting of: single linkage clustering, complete linkage clustering, group-average clustering, and centroid clustering.

8. The method of claim 1 , wherein clustering the documents comprises a non-hierarchical method selected from the group consisting of: monothetic divisive clustering, minimization of trace clustering, multivariate mixture model clustering, Jardine and Sibsons's K-dend clustering, distribution-based model clustering, density based model clustering, partitioning based clustering, and Bayesian based clustering.

9. The method of claim 1 , further comprising projecting the data as at least one of a heatmap and a hexagonal bin plot.

10. The method of claim 1 , wherein the plurality of documents comprises patent documents.

11. The method of claim 1 , wherein the plurality of documents comprises one of scientific, technical, medical, or legal literature.

12. The method of claim 1 , wherein generating a first similarity matrix comprises generating a first similarity matrix using the Jaccard coefficient.

13. The method of claim 1 , wherein the entrywise combination comprises at least one of multiplication, addition, subtraction, and division of the first similarity matrix and the second similarity matrix.

14. The method of claim 1 , further comprising

providing a third set of classification terms, different from the first and second sets of classification terms;

generating a third frequency array of a number of occurrences of each term from the third set of classification terms in each document; and generating a third similarity matrix from the third frequency array;

wherein determining an entrywise combination of the first similarity matrix and the second similarity matrix further comprises determining an entrywise combination of the first similarity matrix, the second similarity matrix, and the third similarity matrix.

15. The method of claim 14 , wherein the first, second, and third frequency arrays comprise an intersection of documents from the plurality of documents which have at least one term from each of the first, second, and third sets of classification terms.

16. The method of claim 1 , wherein the first frequency array includes only documents having at least one term from the first set of classification terms.

17. The method of claim 16 , wherein the second frequency array includes only documents having at least one term from the first set of classification terms.

18. The method of claim 1 , wherein the second frequency array includes only documents having at least one term from the second set of classification terms.

19. The method of claim 1 , wherein the first frequency array includes only documents having at least two different terms from the first set of classification terms.

20. The method of claim 19 , wherein the second frequency array includes only documents having at least one term from the first set of classification terms.

21. The method of claim 1 , wherein the plurality of documents comprises a digital resource.

22. The method of claim 1 , wherein the first set of classification terms comprises an externally managed set of classification terms for organisms, chemicals, enzymes, genes, proteins, minerals, materials, trademarks, or trade names.

23. The method of claim 1 , wherein the first set of classification terms comprises an externally managed set of classification terms comprising a computable terminology.

24. The method of claim 1 , wherein at least one step is carried out using a microprocessor.

25. A computer-based system for classifying a plurality of documents, the system comprising:

a processor; and

a storage medium operably coupled to the processor, wherein the storage medium includes, program instructions executable by the processor for

providing a first set of classification terms and a second set of classification terms, the second set of classification terms being different from the first set of classification terms;

generating a first frequency array of a number of occurrences of each term from the first set of classification terms in each document;

generating a second frequency array of a number of occurrences of each term from the second set of classification terms in each document;

generating a first similarity matrix from the first frequency array;

generating a second similarity matrix from the second frequency array;

determining an entrywise combination of the first similarity matrix and the second similarity matrix; and

clustering the plurality of documents based on the result of the entrywise combination.

26. The computer-based system of claim 25 , wherein the first set of classification terms comprises one of an externally managed set of terms and a patent classification code.

27. The computer-based system of claim 25 , wherein the first set of classification terms comprises an externally managed set of terms.

28. The computer-based system of claim 27 , wherein the externally managed set of terms has been disambiguated.

29. The computer-based system of claim 25 , wherein the first set of classification terms comprises an externally managed set of classification terms for Bacteria and Archaea.

30. The computer-based system of claim 25 , further comprising reordering metadata associated with the plurality of documents according to the clustering.

31. The computer-based system of claim 25 , wherein clustering the documents comprises a hierarchical clustering method selected from the group consisting of: single linkage clustering, complete linkage clustering, group-average clustering, and centroid clustering.

32. The computer-based system of claim 25 , wherein clustering the documents comprises a nonhierarchical method selected from the group consisting of: monothetic divisive clustering, minimization of trace clustering, multivariate mixture model clustering, Jardine and Sibsons's K-dend clustering, distribution-based model clustering, density based model clustering, partitioning based clustering, and Bayesian based clustering.

33. The computer-based system of claim 25 , further comprising projecting the data as at least one of a heatmap and a hexagonal bin plot.

34. The computer-based system of claim 25 , wherein the plurality of documents comprises patent documents.

35. The computer-based system of claim 25 , wherein the plurality of documents comprises one of scientific, technical, medical, or legal literature.

36. The computer-based system of claim 25 , wherein generating a first similarity matrix comprises generating a first similarity matrix using the Jaccard coefficient.

37. The computer-based system of claim 25 , wherein the entrywise combination comprises at least one of multiplication, addition, subtraction, and division of the first similarity matrix and the second similarity matrix.

38. The computer-based system of claim 25 , wherein the program instructions executable by the processor further comprise instructions for

providing a third set of classification terms, different from the first and second sets of classification terms;

generating a third frequency array of a number of occurrences of each term from the third set of classification terms in each document; and

generating a third similarity matrix from the third frequency array;

wherein determining an entrywise combination of the first similarity matrix and the second similarity matrix further comprises determining an entrywise combination of the first similarity matrix, the second similarity matrix, and the third similarity matrix.

39. The computer-based system of claim 38 , wherein the first, second, and third frequency arrays comprise an intersection of documents from the plurality of documents which have at least one term from each of the first, second, and third sets of classification terms.

40. The computer-based system of claim 25 , wherein the first frequency array includes only documents having at least one term from the first set of classification terms.

41. The computer-based system of claim 40 , wherein the second frequency array includes only documents having at least one term from the first set of classification terms.

42. The computer-based system of claim 25 , wherein the second frequency array includes only documents having at least one term from the second set of classification terms.

43. The computer-based system of claim 25 , wherein the first frequency array includes only documents having at least two different terms from the first set of classification terms.

44. The computer-based system of claim 43 , wherein the second frequency array includes only documents having at least one term from the first set of classification terms.

45. The computer-based system of claim 25 , wherein the plurality of documents comprises a digital resource.

46. The computer-based system of claim 25 , wherein the first set of classification terms comprises an externally managed set of classification terms for organisms, chemicals, enzymes, genes, proteins, minerals, materials, trademarks, or trade names.

47. The computer-based system of claim 25 , wherein the first set of classification terms comprises an externally managed set of classification terms comprising a computable terminology.

48. The computer-based system of claim 25 , wherein at least one step is carried out using a microprocessor.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2023
From: NAMESFORLIFE, LLC
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 063660/0036 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2012
From: PARKER, CHARLES T.; GARRITY, GEORGE M.
To: NAMESFORLIFE LLC
Reel/Frame 028261/0775 →
Continuity (2)
Provisional Application 61489362 · May 24, 2011
Related Publication 20130013603A1 · Jan 10, 2013