IP Library Granted Patent US 9,305,082
Granted Patent B2
US 9,305,082 · App. 13/249,760 · Granted Apr 5, 2016

Systems, methods, and interfaces for analyzing conceptually-related portions of text

Inventors: Dietmar H. Dorr (Eagan, MN); Masoud Makrehchi (Waterloo, CA); Carol Steele (Apple Valley, MN)
Assignee: Thomson Reuters Global Resources
G06F17/30705G06F17/277G06F17/2735G06Q10/10G06Q50/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,305,082
App. No.
13/249,760
Granted
Apr 5, 2016
Kind
B2
Abstract

A method includes analyzing a cluster of conceptually-related portions of text to develop a model and calculating a novelty measurement between a first identified conceptually-related portion of text and the model. The method further includes transmitting a second identified conceptually-related portion of text and a score associated with the novelty measurement from a server to an access device via a signal. Another method includes determining at least two corpora of conceptually-related portions of text. The method also includes calculating a common neighbors similarity measurement between the at least two corpora of conceptually-related portions of text and if the common neighbors similarity measurement exceeds a threshold, merging the at least two corpora of conceptually-related portions of text into a cluster or if the common neighbors similarity measurement does not exceed a threshold, maintaining a non-merge of the at least two corpora of conceptually-related portions of text.

Claims (51)

1. A method comprising:

analyzing a first cluster of conceptually-related portions of text to identify a probability for each of the one or more portions of texts within the first cluster of conceptually-related portions of text, wherein the first cluster of conceptually-related portions of text comprises one or more contracts and each of the one or more contracts comprises one or more contract clauses and wherein the probability is calculated based on the number of occurrences of a given contract clause divided by the total number of contracts in the first cluster of conceptually-related portions of text;

developing a model based on the one or more probabilities corresponding to the one or more portions of texts within the first cluster of conceptually-related portions of text;

calculating a novelty measurement between a first identified conceptually-related portion of text and the model; and

transmitting a second identified conceptually-related portion of text and a score associated with the novelty measurement.

2. The method of claim 1 , wherein the cluster of conceptually-related portions of text comprises contracts, the model comprises a model contract, the second identified conceptually-related portion of text comprises one or more contracts.

3. The method of claim 1 , wherein the second identified conceptually-related portion of text each comprise one or more contracts.

4. A method comprising:

analyzing a first cluster of conceptually-related portions of text to identify a probability for each of the one or more portions of texts within the first cluster of conceptually-related portions of text, wherein the first cluster of conceptually-related portions of text comprises one or more sentences and wherein the probability is calculated based on the number of first time occurrences of a word in each of the one or more sentences divided by the total number of sentences in the first cluster of conceptually-related portions of text;

developing a model based on the one or more probabilities corresponding to the one or more portions of texts within the first cluster of conceptually-related portions of text;

calculating a novelty measurement between a first identified conceptually-related portion of text and the model; and

transmitting a second identified conceptually-related portion of text and a score associated with the novelty measurement.

5. The method of claim 4 wherein the cluster of conceptually-related portions of text comprises contract clauses, the model comprises a model contract clause, and the second identified conceptually-related portion of text comprises one or more contract clauses.

6. The method of claim 4 wherein the first identified conceptually-related portion of text and the second identified conceptually-related portion of text are each contract clauses.

7. The method of claim 4 further comprising receiving a first identified conceptually-related portion of text from a third party.

8. The method of claim 4 , wherein the first identified conceptually-related portions of a text is a paragraph.

9. The method of claim 4 , wherein the first cluster of conceptually-related portions of text is generated by:

determining at least two corpora of conceptually-related portions of text;

calculating a common neighbors similarity measurement between the at least two corpora of conceptually-related portions of text; and

merging the at least two corpora of conceptually-related portions of text into the first cluster of conceptually-related portions of text where the common neighbors similarity measurement exceeds a threshold.

10. The method of claim 9 wherein the at least two corpora of conceptually-related portions of text comprises contracts, the cluster of conceptually-related portions of text comprises contracts and the model comprises a model contract.

11. The method of claim 9 wherein the at least two corpora of conceptually-related portions of text comprises contract clauses, the cluster of conceptually-related portions of text comprises contract clauses and the model comprises a model contract clause.

12. A system comprising:

a processor;

a memory coupled to the processor;

a processing program stored in the memory for execution by the processor, the processing program comprising:

an analysis module, the analysis module configured to:

analyze a first cluster of conceptually-related portions of text to identify a probability for each of the one or more portions of texts within the first cluster of conceptually-related portions of text, wherein the first cluster of conceptually-related portions of text comprises one or more contracts and each of the one or more contracts comprises one or more contract clauses and wherein the probability is calculated based on the number of occurrences of a given contract clause divided by the total number of contracts in the first cluster of conceptually-related portions of text; and

develop a model based on the one or more probabilities corresponding to the one or more portions of texts within the first cluster of conceptually-related portions of text;

a novelty module, the novelty module configured to calculate a novelty measurement between a first identified conceptually-related portion of text and the model; and

a transmission module, the transmission module configured to transmit a second identified conceptually-related portion of text and a score associated with the novelty measurement.

13. The system of claim 12 , wherein the cluster of conceptually-related portions of text comprises contracts, the model comprises a model contract, the second identified conceptually-related portion of text comprises one or more contracts.

14. The system of claim 12 , wherein the second identified conceptually-related portion of text each comprise one or more contracts.

15. A system comprising:

a processor;

a memory coupled to the processor;

a processing program stored in the memory for execution by the processor, the processing program comprising:

an analysis module, the analysis module configured to:

analyze a first cluster of conceptually-related portions of text to identify a probability for each of the one or more portions of texts within the first cluster of conceptually-related portions of text, wherein the first cluster of conceptually-related portions of text comprises one or more sentences and wherein the probability is calculated based on the number of first time occurrences of a word in each of the one or more sentences divided by the total number of sentences in the first cluster of conceptually-related portions of text; and

develop a model based on the one or more probabilities corresponding to the one or more portions of texts within the first cluster of conceptually-related portions of text;

a novelty module, the novelty module configured to calculate a novelty measurement between a first identified conceptually-related portion of text and the model; and

a transmission module, the transmission module configured to transmit a second identified conceptually-related portion of text and a score associated with the novelty measurement.

16. The system of claim 15 further comprising a common neighbors module, the common neighbors module configured to:

determine at least two corpora of conceptually-related portions of text;

calculate a common neighbors similarity measurement between the at least two corpora of conceptually-related portions of text; and

merge the at least two corpora of conceptually-related portions of text into the first cluster of conceptually-related portions of text where the common neighbors similarity measurement exceeds a threshold.

17. The system of claim 16 wherein the at least two corpora of conceptually-related portions of text comprises contract clauses, the cluster of conceptually-related portions of text comprises contract clauses and the model comprises a model contract clause.

18. The system of claim 15 wherein the cluster of conceptually-related portions of text comprises contract clauses, the model comprises a model contract clause, and the second identified conceptually-related portion of text comprises a contract clause.

19. The system of claim 15 wherein the first identified conceptually-related portion of text and the second identified conceptually-related portion of text are each contract clauses.

20. The system of claim 15 wherein the novelty module is further configured to receive the first identified conceptually-related portion of text from a third party.

21. The system of claim 15 , wherein the first identified conceptually-related portions of a text is a paragraph.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2020
From: THOMSON REUTERS GLOBAL RESOURCES UNLIMITED COMPANY
To: THOMSON REUTERS ENTERPRISE CENTRE GMBH
Reel/Frame 052035/0862 →
CHANGE OF NAME Recorded Nov 30, 2017
From: THOMSON REUTERS GLOBAL RESOURCES
To: THOMSON REUTERS GLOBAL RESOURCES UNLIMITED COMPANY
Reel/Frame 044264/0277 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2012
From: WEST SERVICES INC.
To: THOMSON REUTERS GLOBAL RESOURCES
Reel/Frame 028257/0443 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2012
From: THOMSON REUTERS HOLDINGS INC.
To: THOMSON REUTERS GLOBAL RESOURCES
Reel/Frame 028257/0522 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2012
From: STEELE, CAROL
To: WEST SERVICES INC.
Reel/Frame 028219/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2012
From: DORR, DIETMAR H.; MAKREHCHI, MASOUD
To: THOMSON REUTERS HOLDINGS INC.
Reel/Frame 028219/0675 →
Continuity (1)
Related Publication 20130086470A1 · Apr 4, 2013