IP Library Granted Patent US 10,372,717
Granted Patent B2
US 10,372,717 · App. 14/922,585 · Granted Aug 6, 2019

Systems and methods for identifying documents based on citation history

Inventors: Paul Zhang (Centerville, OH); Harry R. Silver (Shaker Heights, OH)
Assignee: LexisNexis, a division of Reed Elsevier Inc.
G06F16/24578G06F16/93G06F16/9535
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,372,717
App. No.
14/922,585
Granted
Aug 6, 2019
Kind
B2
Abstract

Systems, methods, and computer-executable instructions for identifying a document are described. A method includes receiving a query from a graphical user interface having one or more concepts, normalizing a set of terms or concepts in the query to create a normalized query, comparing the normalized query to a set of document centric concept profiles associated with a set of documents in a corpus, where each document centric concept includes a plurality of concepts and at least one reference value for each concept, where the reference value is calculated by tabulating the number of times a document associated with one of the document centric concept profiles is cited by a citing instance for the concept, and surfacing a document from the corpus with the highest reference value for the concept.

Claims (46)

1. A method to identify a document within a corpus of documents generated as a result of a bibliometric process for the purposes of determining whether the document is a most frequently cited document for a reason for citation, the method comprising automatically:

receiving a query from a graphical user interface comprising one or more concepts;

normalizing a set of terms or concepts in the query to create a normalized query;

comparing the normalized query to a set of document centric concept profiles, wherein each document centric concept profile is associated with a document in the corpus, and wherein each document centric concept profile comprises:

a plurality of key concepts identified by:

establishing a reason for citation associated with each citing instance that has cited the document associated with that document centric concept profile, wherein each reason for citation is based on a text area within each respective citing instance; and

comparing each reason for citation to a normalized key concept list to identify at least one key concept; and

a reference value for each key concept, wherein the reference value is calculated by tabulating a number of times the document associated with that document centric concept profile has been cited by citing instances for each key concept; and

surfacing a particular document from the corpus, the particular document having a highest reference value for a key concept corresponding to the normalized query, wherein the particular document is the most frequently cited document for the reason for citation.

2. The method of claim 1 , wherein the set of terms or concepts from the query are not present in the particular document surfaced from the corpus.

3. The method of claim 1 , wherein one or more document centric concept profiles of the set of document centric concept profiles comprises a reference value assigned to a cluster of concepts.

4. The method of claim 1 , wherein the surfacing occurs for the particular document from the corpus based on a subset of the set of document centric concept profiles.

5. The method of claim 4 , wherein the corpus comprises documents chosen from a set of case opinions, statutes, and regulations.

6. The method of claim 5 , wherein the particular document from the surfacing, when scored using Term-Frequency-Inverse-Document-Frequency techniques, has a lower score and a higher reference value than a second document.

7. The method of claim 1 , wherein at least one document in the corpus is associated with more than one document centric concept profile.

8. The method of claim 1 , wherein surfacing the particular document from the corpus comprises generating a user interface component that includes the particular document at the top of a list of search results.

9. A non-transitory computer-readable memory comprising computer-executable instructions for execution by a computer machine to identify a document within a corpus of documents generated as a result of a bibliometric process for the purposes of determining whether the document is a most frequently cited document for a reason for citation, the computer-executable instructions, when executed, cause the computer machine to:

receive a query including at least one concept;

compare the at least one concept of the query to each document centric profile of a set of document centric concept profiles contained in a corpus located in a computerized database, wherein each document centric concept profile comprises:

a plurality of key concepts identified by:

establishing a reason for citation associated with each citing instance that has cited the document associated with that document centric concept profile, wherein each reason for citation is based on a text area within each respective citing instance; and

comparing each reason for citation to a normalized key concept list to identify at least one key concept; and

a reference value for each key concept wherein the reference value is calculated by tabulating a number of times the document associated with that document centric concept profile has been cited by citing instances for each key concept;

surface a set of documents, wherein each document of the set of documents is associated with a document centric concept profiles having a key concept that matches the at least one concept of the query; and

rank the set of documents by their respective reference value associated with the key concept that matches the at least one concept of the query, wherein the highest-ranked document of the set of documents represents the most frequently cited document for the reason for citation.

10. The non-transitory computer-readable memory of claim 9 , wherein the computer machine is chosen from the group consisting of: a mobile device, a desktop, and a laptop.

11. The non-transitory computer-readable memory of claim 9 , wherein the query includes multiple concepts.

12. The non-transitory computer-readable memory of claim 11 , wherein the multiple concepts form a concept cluster.

13. The non-transitory computer-readable memory of claim 12 , wherein the computer-executable instructions are further configured to compare the concept cluster to each document centric concept profile of the set of document centric concept profiles.

14. A system to identify a document within a corpus of documents generated as a result of a bibliometric process for the purposes of determining whether the document is a most frequently cited document for a reason for citation, the system comprising:

a processing device; and

a non-transitory, processor-readable storage medium, the non-transitory processor-readable storage medium comprising one or more programming instructions that, when executed, cause the processing device to:

receive a query from a graphical user interface comprising one or more concepts;

normalize a set of terms or concepts in the query to create a normalized query;

compare the normalized query to a set of document centric concept profiles, wherein each document centric concept profile is associated with a document in the corpus, and wherein each document centric concept profile comprises:

a plurality of key concepts identified by:

establishing a reason for citation associated with each citing instance that has cited the document associated with that document centric concept profile, wherein each reason for citation is based on a text area within each respective citing instance; and

comparing each reason for citation to a normalized key concept list to identify at least one key concept; and

a reference value for each key concept, wherein the reference value is calculated by tabulating a number of times the document associated with that document centric concept profile has been cited by citing instances for each key concept; and

surface a particular document from the corpus, the particular document having a highest reference value for the key concept corresponding to the normalized query, wherein the particular document is the most frequently cited document for the reason for citation.

15. The system of claim 14 , wherein the set of terms or concepts from the query are not present in the particular document surfaced from the corpus.

16. The system of claim 14 , wherein one or more document centric concept profiles of the set of document centric concept profiles comprises a reference value assigned to a cluster of concepts.

17. The system of claim 14 , wherein the one or more programming instructions that, when executed, cause the processing device to surface the particular document further cause the processing device to surface the particular document from the corpus based on a subset of the set of document centric concept profiles.

18. The system of claim 17 , wherein the corpus comprises documents chosen from a set of case opinions, statutes, and regulations.

19. The system of claim 18 , wherein the particular document from the surfacing, when scored using Term-Frequency-Inverse-Document-Frequency techniques, has a lower score and a higher reference value than a second document.

20. The system of claim 14 , wherein at least one document in the corpus is associated with more than one document centric concept profile.

Assignments (2)
CHANGE OF NAME Recorded Dec 3, 2019
From: LEXISNEXIS; REED ELSEVIER INC.
To: RELX INC.
Reel/Frame 051198/0325 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2019
From: ZHANG, PAUL; SILVER, HARRY R.
To: LEXISNEXIS, A DIVISION OF REED ELSEVIER INC.
Reel/Frame 049548/0033 →
Continuity (2)
Continuation 13755874 · Jan 31, 2013
Related Publication 20160098407A1 · Apr 7, 2016