IP Library Granted Patent US 9,690,849
Granted Patent B2
US 9,690,849 · App. 14/201,134 · Granted Jun 27, 2017

Systems and methods for determining atypical language

Inventors: Sameena Shah (White Plains, NY); Dietmar Dorr (Sunnyvale, CA); Khalid Al-Kofahi (Rosemount, MN); Jacob Sisk (Brooklyn, NY)
Assignee: Thomson Reuters Global Resources Unlimited Company
G06F17/30705G06F17/277G06F17/2735G06F17/30699G06Q10/10G06Q50/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,690,849
App. No.
14/201,134
Granted
Jun 27, 2017
Kind
B2
Abstract

A method includes analyzing a cluster of conceptually-related portions of text to develop a model and calculating a novelty measurement between a first identified conceptually-related portion of text and the model. The method further includes transmitting a second identified conceptually-related portion of text and a score associated with the novelty measurement from a server to an access device via a signal. Another method includes determining at least two corpora of conceptually-related portions of text. The method also includes calculating a common neighbors similarity measurement between the at least two corpora of conceptually-related portions of text and if the common neighbors similarity measurement exceeds a threshold, merging the at least two corpora of conceptually-related portions of text into a cluster or if the common neighbors similarity measurement does not exceed a threshold, maintaining a non-merge of the at least two corpora of conceptually-related portions of text.

Claims (24)

1. A computer-implemented method comprising:

analyzing a first cluster of conceptually-related portions of text to identify a probability for each of the one or more portions of texts within the first cluster of conceptually-related portions of text, wherein the first cluster of conceptually-related portions of text comprises one or more financial documents and each of the one or more financial documents comprises one or more financial document sections and each of the one or more financial document sections comprises one or more sentences, and wherein the probability is calculated based on the number of occurrences of a given token of a given sentence of a given financial document of the first cluster of conceptually-related portions of text;

developing a model based on the one or more probabilities corresponding to the one or more portions of texts within the first cluster of conceptually-related portions of text;

calculating an abnormality score for each of the one or more sentences of the one or more financial document sections of a first identified conceptually-related portion of text as compared to the model; and

transmitting a second identified conceptually-related portion of text based upon the abnormality score satisfying a threshold.

2. The method of claim 1 , wherein the first cluster of conceptually-related portions of text is generated by aggregating one or more financial documents according to an assigned key value.

3. The method of claim 2 , wherein the assigned key value is a time period of a given financial document.

4. The method of claim 2 , wherein the assigned key value is a sector for a given financial document.

5. The method of claim 2 , wherein the assigned key value is a market cap for a given financial document.

6. The method of claim 1 , wherein the second identified conceptually-related portion of text comprises one or more sentences identified as atypical.

7. The method of claim 1 , wherein the second identified conceptually-related portion of text comprises one or more sentences identified as typical.

8. A system comprising:

a processor;

a memory coupled to the processor; and

a processing program stored in the memory for execution by the processor, the processing program comprising:

an analysis module, the analysis module configured to analyze a cluster of conceptually-related portions of text to identify a probability for each of the one or more portions of texts within the first cluster of conceptually-related portions of text, wherein the first cluster of conceptually-related portions of text comprises one or more financial documents and each of the one or more financial documents comprises one or more financial document sections' and each of the one or more financial document sections comprises one or more sentences, and wherein the probability is calculated based on the number of occurrences of a given token of a given sentence of a given financial document of the first cluster of conceptually-related portions of text and to develop a model based on the one or more probabilities corresponding to the one or more portions of texts within the first cluster of conceptually-related portions of text;

a novelty module, the novelty module configured to calculate an abnormality score for each of the one or more sentences of the one or more financial document sections of a first identified conceptually-related portion of text as compared to the model; and

a transmission module, the transmission module configured to transmit a second identified conceptually-related portion of text based upon the abnormality score satisfying a threshold.

9. The system of claim 8 , wherein the first cluster of conceptually-related portions of text is generated by aggregating one or more financial documents according to an assigned key value.

10. The system of claim 9 , wherein the assigned key value is a time period of a given financial document.

11. The system of claim 9 , wherein the assigned key value is a sector for a given financial document.

12. The system of claim 9 , wherein the assigned key value is a market cap for a given financial document.

13. The system of claim 8 , wherein the second identified conceptually-related portion of text comprises one or more sentences identified as atypical.

14. The system of claim 8 , wherein the second identified conceptually-related portion of text comprises one or more sentences identified as typical.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 5, 2020
From: THOMSON REUTERS GLOBAL RESOURCES UNLIMITED COMPANY
To: THOMSON REUTERS ENTERPRISE CENTRE GMBH
Reel/Frame 052029/0558 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2017
From: THOMSON REUTERS GLOBAL RESOURCES
To: THOMSON REUTERS GLOBAL RESOURCES UNLIMITED COMPANY
Reel/Frame 042553/0942 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2015
From: THOMSON REUTERS HOLDINGS, INC.
To: THOMSON REUTERS GLOBAL RESOURCES
Reel/Frame 036133/0461 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2015
From: SHAH, SAMEENA; AL-KOFAHI, KHALID
To: THOMSON REUTERS HOLDINGS, INC.
Reel/Frame 036066/0777 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2015
From: SISK, JACOB; DORR, DIETMAR
To: THOMSON REUTERS GLOBAL RESOURCE
Reel/Frame 036066/0811 →
Continuity (3)
Continuation In Part 13249760 · Sep 30, 2011
Provisional Application 61774863 · Mar 8, 2013
Related Publication 20140344279A1 · Nov 20, 2014