IP Library Granted Patent US 7,953,724
Granted Patent B2
US 7,953,724 · App. 11/799,768 · Granted May 31, 2011

Method and system for disambiguating informational objects

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,953,724
App. No.
11/799,768
Granted
May 31, 2011
Kind
B2
Abstract

The present invention provides a Distinct Author Identification System (“DAIS”) for disambiguating data to discern author entities and link or associate authorships with such author entities. The invention provides powerful disambiguation processes applied across one or more databases to yield a disambiguated authority database of authors. An entire database of publications may be processed by the DAIS to group/link authorships and to identify author entities. The author entities may then be matched or associated with actual authors to establish an authority database of authors. After initial evaluation, the DAIS may be used to reevaluate some or all of the database(s) and/or the authority database established by the DAIS may be used to add or update information. DAIS may use “hierarchical clustering” to link authorships and identify authors based on authorship similarity. DAIS evaluates the likelihood that authorships are from the same author.

Claims (59)

1. A computer implemented method comprising:

a. selecting a set of electronic information associated with a set of publications, each publication in the set of publications comprising at least one cited reference and having at least one authorship;

b. disambiguating at least part of the set of electronic information by using a set of at least two cited references associated with a set of at least two publications from the set of publications to determine an authorship similarity, disambiguating including scoring authorship similarity; wherein the disambiguating step includes arriving at a scored authorship similarity attribute;

c. linking authorships based on the determined authorship similarity and clustering two or more linked authorships to form a first cluster and forming a first author entity associated with the first cluster;

d. matching the first author entity with a first actual author, the first cluster of authorships being attributable to the first actual author, and repeating the clustering step to form a plurality of clusters respectively associated with a plurality of unique author entities; and

e. incorporating into an authority database of authors the plurality of unique author entities each associated with a unique actual author and a cluster.

2. The method of claim 1 , wherein incorporating further comprises comprising establishing an authority database of authors comprising the plurality of unique author entities each associated with a unique actual author and a cluster.

3. The method of claim 1 further comprising receiving notice of an erroneous match of an actual author with at least one of an authorship, a cluster, or an author entity, and based on the notice doing one of associating and disassociating the actual author from the at least one of an authorship, a cluster, or an author entity.

4. The method of claim 1 , wherein the disambiguating step includes arriving at a scored authorship similarity attribute and the linking step is based on the scored authorship similarity attribute meeting or exceeding a predetermined degree of similarity.

5. The method of claim 1 , wherein the scored authorship similarity attribute is based at least in part on author name data.

6. The method of claim 5 , wherein the degree of similarity is based at least in part on a commonality of the name data.

7. The method of claim 5 , wherein the degree of similarity is based at least in part on a frequency of occurrence of the name data.

8. The method of claim 1 , wherein the scored authorship similarity attribute is based at least in part on co-authorship data comprising the number of authorships associated with publications, wherein as the number of co-authorships increases, the degree of similarity associated with the co-authorship data decreases.

9. The method of claim 8 , wherein the co-authorship data comprises co-author name data and matching co-author name data among publications increases the scored authorship similarity attribute.

10. The method of claim 1 , wherein the scored authorship similarity attribute is based at least in part on email address data or co-author data contained in the set of at least two publications.

11. The method of claim 1 , wherein the determined authorship similarity is insufficient to form a link in the linking step, and wherein the linking step further comprises processing information derived from the set of electronic information to establish a secondary link between authorships.

12. The method of claim 1 further comprising processing information derived from the set of electronic information to confirm or disassociate links established in the linking step.

13. The method of claim 1 further comprising processing information derived from the set of electronic information to confirm or disassociate clusters established in the clustering step.

14. The method of claim 1 further comprising reevaluating at least a portion of the established authority database of authors based on supplemental information including data representing a threshold number of publications having common author name data.

15. The method of claim 1 , wherein the disambiguating step may further comprise processing at least one of the following elements: email address; co-author data; address data; paper title; cited reference author name; cited by paper; cited by author name; keywords; Publication Discipline Code; and additional author name initial data.

16. The method of claim 1 further comprising:

establishing a communication link with a client;

receiving from the client a query; and

processing the query and presenting the client with disambiguated data.

17. The method of claim 1 , wherein disambiguating at least part of the set of electronic information by using a set of at least two cited references includes at least one of the following: co-citation; bibliographic coupling; and self cite.

18. A computer-based system comprising:

a computer adapted to process a set of electronic information associated with a set of publications, each publication in the set of publications comprising at least one cited reference and having at least one authorship;

software executing on the computer and adapted to disambiguate at least part of the set of electronic information by using a set of at least two cited references associated with a set of at least two publications from the set of publications to determine an authorship similarity;

a database operatively connected to the computer and adapted to receive and store for processing by the computer the set of information;

an authorship similarity routine executing on the computer and adapted to process at least some of the set of electronic information using cited reference data to determine a degree of authorship similarity;

a linking routine executing on the computer and adapted to link authorships based on the degree of authorship similarity;

a clustering routine executing on the computer and adapted to cluster two or more linked authorships to form a first cluster and adapted to form a first author entity associated with the first cluster, and wherein the database comprises an authority database of authors comprised of a plurality of distinct actual authors matched respectively with a plurality of unique author entities.

19. The computer-based system of claim 18 wherein the clustering routine is further adapted to match the first author entity with a first actual author, the first cluster of authorships being attributable to the first actual author.

20. The computer-based system of claim 18 , wherein a plurality of clusters are respectively associated with a plurality of unique author entities.

21. The computer-based system of claim 18 , wherein the system receives electronic notice of an erroneous match of an actual author with at least one of an authorship, a cluster, or an author entity, and the system having a attribution routine adapted to do one of associate or disassociate the actual author from the at least one of an authorship, a cluster, or an author entity based on the notice.

22. The computer-based system of claim 18 wherein the degree of authorship similarity is based at least in part on author name data.

23. The computer-based system of claim 22 wherein the degree of authorship similarity is based at least in part on a commonality of the author name data.

24. The computer-based system of claim 22 wherein the degree of authorship similarity is based at least in part on a frequency of occurrence of the name data.

25. The computer-based system of claim 18 wherein the degree of authorship similarity is based at least in part on co-authorship data comprising the number of authorships associated with publications, as the number of co-authorships increases, the degree of similarity associated with the co-authorship component decreases.

26. The computer-based system of claim 18 wherein the degree of authorship similarity is based at least in part on co-authorship data comprising co-author name data, whereby publications having matching co-author name data results in a higher degree of authorship similarity.

27. The computer-based system of claim 18 , wherein the degree of authorship similarity is insufficient to form a link, the system further comprising an alternate linking routine adapted to process information derived from the set of electronic information to establish a secondary link between authorships.

28. The computer-based system of claim 18 wherein the linking routine is further adapted to process information derived from the set of electronic information to confirm or disassociate links.

29. The computer-based system of claim 18 wherein the clustering routine is further adapted to process information derived from the set of electronic information to confirm or disassociate linked authorships from clusters.

30. The computer-based system of claim 18 further comprising a reevaluation routine executing on the computer and adapted to process at least a portion of the authority database of authors based on supplemental information.

31. The computer-based system of claim 30 wherein the supplemental information includes data representing a threshold number of publications having common author name data, the system determining whether to execute the reevaluation routine being based at least in part on the threshold number.

32. The computer-based system of claim 18 , wherein the degree of authorship similarity is based at least in part on: email address; address; co-author name; cited reference paper; cited reference author name; cited by paper; cited by author name; keywords; Publication Discipline Code; and additional author name initial.

33. The computer-based system of claim 18 , wherein a client-based computer is in communication with the database and is adapted to query against the authority database of authors, whereby the query is processed and the client is presented with disambiguated data.

34. The computer-based system of claim 33 , wherein the client-based computer, in conjunction with a research productivity software, accesses and queries the database and publications databases to develop bibliographic data records.

35. The computer-based system of claim 18 further comprising:

establishing a communication link with a client;

receiving from the client a query; and

processing the query and presenting the client with disambiguated data.

36. The computer-based system of claim 18 , wherein the authorship similarity routine is further adapted to disambiguate the at least some of the set of electronic information by using at least one of the following: co-citation; bibliographic coupling; and self cite.

37. A computer implemented method for maintaining an authority database of authors used in searching at least one publications database for publications of interest, the method comprising:

a. receiving publications, each publication containing at least one cited reference and having at least one authorship; and

b. disambiguating the received publications by comparing the at least one cited references with data associated with the authority database of authors to determine an authorship similarity between publication authorships, disambiguating including scoring authorship similarity; wherein the disambiguating step includes arriving at a scored authorship similarity attribute;

c. linking authorships based on the determined authorship similarity and clustering two or more linked authorships to form a first cluster and forming a first author entity associated with the first cluster;

d. matching the first author entity with a first actual author, the first cluster of authorships being attributable to the first actual author, and repeating the clustering step to form a plurality of clusters respectively associated with a plurality of unique author entities; and

e. incorporating into the authority database of authors the plurality of unique author entities each associated with a unique actual author and a cluster.

Assignments (14)
SECURITY INTEREST Recorded Dec 3, 2021
From: DECISION RESOURCES, INC.; DR/DECISION RESOURCES, LLC; CPA GLOBAL (FIP) LLC; CPA GLOBAL PATENT RESEARCH LLC; INNOGRAPHY, INC.; CAMELOT UK BIDCO LIMITED
To: WILMINGTON TRUST, NATIONAL ASSOCATION
Reel/Frame 058907/0091 →
SECURITY INTEREST Recorded Dec 17, 2019
From: CAMELOT UK BIDCO LIMITED
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 051323/0875 →
SECURITY INTEREST Recorded Dec 17, 2019
From: CAMELOT UK BIDCO LIMITED
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 051323/0972 →
SECURITY INTEREST Recorded Nov 1, 2019
From: CAMELOT UK BIDCO LIMITED
To: BANK OF AMERICA, N.A.
Reel/Frame 050906/0284 →
SECURITY INTEREST Recorded Nov 1, 2019
From: CAMELOT UK BIDCO LIMITED
To: WILMINGTON TRUST, N.A. AS COLLATERAL AGENT
Reel/Frame 050906/0553 →
RELEASE OF SECURITY INTEREST Recorded Nov 1, 2019
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: CAMELOT UK BIDCO LIMITED
Reel/Frame 050911/0796 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2016
From: THOMSON REUTERS GLOBAL RESOURCES
To: CAMELOT UK BIDCO LIMITED
Reel/Frame 040206/0448 →
SECURITY INTEREST Recorded Oct 3, 2016
From: CAMELOT UK BIDCO LIMITED
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 040205/0156 →
MERGER AND CHANGE OF NAME Recorded Jul 1, 2015
From: THOMSON REUTERS (SCIENTIFIC) INC.; THOMSON REUTERS (SCIENTIFIC) LLC.
To: THOMSON REUTERS (SCIENTIFIC) LLC.
Reel/Frame 035991/0204 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2014
From: THOMSON REUTERS (SCIENTIFIC) LLC
To: THOMSON REUTERS GLOBAL RESOURCES
Reel/Frame 034365/0854 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2009
From: THOMSON REUTERS GLOBAL RESOURCES
To: THOMSON REUTERS (SCIENTIFIC) INC.
Reel/Frame 023452/0838 →
CHANGE OF NAME Recorded Oct 6, 2008
From: THOMSON GLOBAL RESOURCES
To: THOMSON REUTERS GLOBAL RESOURCES
Reel/Frame 021630/0917 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2007
From: THOMSON CORPORATION
To: THOMSON GLOBAL RESOURCES
Reel/Frame 020110/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2007
From: GRIFFITH, ROBERT ALBERT
To: THOMSON CORPORATION
Reel/Frame 019322/0553 →