IP Library Granted Patent US 10,885,124
Granted Patent B2
US 10,885,124 · App. 15/451,069 · Granted Jan 5, 2021

Domain-specific negative media search techniques

Inventors: Gary Shiffman (Arlington, VA); Jeffrey Borowitz (Arlington, VA)
Assignee: Giant Oak, Inc.
G06F16/951G06F16/334G06F16/338G06F16/3344G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,885,124
App. No.
15/451,069
Granted
Jan 5, 2021
Kind
B2
Abstract

In some implementations, systems and methods that are capable of customizing negative media searches using domain-specific search indexes are described. Data indicating a search query associated with a negative media search for an entity and a corpus of documents to be searched are obtained. Content from a particular collection of documents from among the corpus of documents is obtained and processed. Multiple scores for the entity are computed based on processing the content obtained from the collection of documents. The multiple scores are aggregated to compute a priority indicator that represents a likelihood that the collection of documents includes content that is descriptive of derogatory information.

Claims (87)

1. A method performed by one or more computers, the method comprising:

receiving data indicating (i) a search query associated with a negative media search for an entity, and (ii) a corpus of documents to be searched using the search query, the corpus of documents including documents that are predetermined to satisfy one or more criteria associated with the negative media search for the entity;

obtaining content from a particular collection of documents from among the corpus of documents that are determined to be responsive to the search query;

processing the content obtained from the particular collection of documents;

based on processing the content obtained from the particular collection of documents:

computing concept scores, wherein each concept score included in the concept scores represents a likelihood that the entity is associated with a predetermined negative attribute of a corresponding reference entity group of a plurality of reference entity groups predetermined to be associated with derogatory information, and

computing relevancy scores, wherein each relevancy score of the relevancy scores represents a likelihood that a corresponding document included in the particular collection of documents includes content that is descriptive of the predetermined negative attributes;

determining a number of reference entity groups having corresponding concept scores satisfying a first threshold associated with the predetermined negative attributes;

determining a number of documents having corresponding relevancy scores satisfying a second threshold associated with the predetermined negative attributes;

aggregating the concept scores and the relevancy scores to compute a priority indicator, wherein:

the priority indicator is computed based at least on (i) the number of reference entity groups having corresponding concept scores satisfying the first threshold, and (ii) the number of documents having corresponding relevancy scores satisfying the second threshold, and

the priority indicator represents a likelihood that the particular collection of documents includes content descriptive of the derogatory information; and

enabling a user to perceive a representation of the priority indicator.

2. The method of claim 1 , wherein processing the content obtained from the particular collection of documents comprises:

computing, for each document included within the particular collection of documents, a reliability score representing a likelihood that a particular document is associated with the entity;

determining that the reliability scores for one or more documents does not satisfy a predetermined threshold; and

removing the one or more documents from the particular collection of documents based on determining that the reliability scores for one or more documents does not satisfy the predetermined threshold.

3. The method of claim 2 , wherein computing the reliability score comprises:

obtaining one or more text fragments from a particular document;

determining a respective topic associated with each of the one or more text fragments; and

determining a likelihood that at least one of the topics are associated with the entity.

4. The method of claim 2 , further comprising:

computing a commonality score that represents a probability that a text fragment corresponding to a name of the entity will be included in a particular document included within the particular collection of documents; and

wherein the reliability scores that are computed for each document included within the particular collection of documents is computed based at least on the commonality score.

5. The method of claim 1 , wherein the concept scores are computed based at least on determining that the entity is included in a list of sanctioned entities.

6. The method of claim 1 , further comprising:

computing additional scores for the entity based on processing the content obtained from the particular collection of documents, the additional scores indicating respective likelihoods that documents within the particular collection of documents are associated with the entity; and

wherein computing the priority indicator comprises aggregating the concept scores, the relevancy scores, and the additional scores.

7. The method of claim 1 , further comprising:

obtaining data indicating a set of documents that are (i) associated with the plurality of reference entity groups and (ii) manually identified as being associated with derogatory information;

processing the set of documents to identify a set of negative attributes for the plurality of reference entity groups; and

designating the set of negative attributes as the predetermined negative attributes.

8. The method of claim 1 , wherein:

the one or more criteria associated with the negative media search for the entity comprises a criterion representing a risk-associated behavior; and

the particular collection of documents comprises documents that are (i) obtained from a plurality of distinct data sources and (ii) include historical information indicating occurrence of the risk-associated behavior.

9. A system comprising:

one or more computers; and

one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving data indicating (i) a search query associated with a negative media search for an entity, and (ii) a corpus of documents to be searched using the search query, the corpus of documents including documents that are predetermined to satisfy one or more criteria associated with the negative media search for the entity;

obtaining content from a particular collection of documents from among the corpus of documents that are determined to be responsive to the search query;

processing the content obtained from the particular collection of documents;

based on processing the content obtained from the particular collection of documents:

computing concept scores, wherein each concept score included in the concept scores represents a likelihood that the entity is associated with a predetermined negative attribute of a corresponding reference entity group of a plurality of reference entity groups predetermined to be associated with derogatory information, and

computing relevancy scores, wherein each relevancy score of the relevancy scores represents a likelihood that a corresponding document included in the particular collection of documents includes content that is descriptive of the predetermined negative attributes;

determining a number of reference entity groups having corresponding concept scores satisfying a first threshold associated with the predetermined negative attributes;

determining a number of documents having corresponding relevancy scores satisfying a second threshold associated with the predetermined negative attributes;

aggregating the concept scores and the relevancy scores to compute a priority indicator, wherein:

the priority indicator is computed based at least on (i) the number of reference entity groups having corresponding concept scores satisfying the first threshold, and (ii) the number of documents having corresponding relevancy scores satisfying the second threshold, and

the priority indicator represents a likelihood that the particular collection of documents includes content descriptive of the derogatory information; and

enabling a user to perceive a representation of the priority indicator.

10. The system of claim 9 , wherein processing the content obtained from the particular collection of documents comprises:

computing, for each document included within the particular collection of documents, a reliability score representing a likelihood that a particular document is associated with the entity;

determining that the reliability scores for one or more documents does not satisfy a predetermined threshold; and

removing the one or more documents from the particular collection of documents based on determining that the reliability scores for one or more documents does not satisfy the predetermined threshold.

11. The system of claim 10 , wherein computing the reliability score comprises:

obtaining one or more text fragments from a particular document;

determining a respective topic associated with each of the one or more text fragments; and

determining a likelihood that at least one of the topics are associated with the entity.

12. The system of claim 10 , further comprising:

computing a commonality score that represents a probability that a text fragment corresponding to a name of the entity will be included in a particular document included within the particular collection of documents; and

wherein the reliability scores that are computed for each document included within the particular collection of documents is computed based at least on the computed commonality score.

13. The system of claim 9 , wherein the concept scores are computed based at least on determining that the entity is included in a list of sanctioned entities.

14. A non-transitory computer-readable storage device storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising:

receiving data indicating (i) a search query associated with a negative media search for an entity, and (ii) a corpus of documents to be searched using the search query, the corpus of documents including documents that are predetermined to satisfy one or more criteria associated with the negative media search for the entity;

obtaining content from a particular collection of documents from among the corpus of documents that are determined to be responsive to the search query;

processing the content obtained from the particular collection of documents;

based on processing the content obtained from the particular collection of documents:

computing concept scores, wherein each concept score included in the concept scores represents a likelihood that the entity is associated with a predetermined negative attribute of a corresponding reference entity group of a plurality of reference entity groups predetermined to be associated with derogatory information, and

computing relevancy scores, wherein each relevancy score of the relevancy scores represents a likelihood that a corresponding document included in the particular collection of documents includes content that is descriptive of the predetermined negative attributes;

determining a number of reference entity groups having corresponding concept scores satisfying a first threshold associated with the predetermined negative attributes;

determining a number of documents having corresponding relevancy scores satisfying a second threshold associated with the predetermined negative attributes;

aggregating the concept scores and the relevancy scores to compute a priority indicator, wherein:

the priority indicator is computed based at least on (i) the number of reference entity groups having corresponding concept scores satisfying the first threshold, and (ii) the number of documents having corresponding relevancy scores satisfying the second threshold, and

the priority indicator represents a likelihood that the particular collection of documents includes content descriptive of the derogatory information; and

enabling a user to perceive a representation of the priority indicator.

15. The storage device of claim 14 , wherein processing the content obtained from the particular collection of documents comprises:

computing, for each document included within the particular collection of documents, a reliability score representing a likelihood that a particular document is associated with the entity;

determining that the reliability scores for one or more documents does not satisfy a predetermined threshold; and

removing the one or more documents from the particular collection of documents based on determining that the reliability scores for one or more documents does not satisfy the predetermined threshold.

16. The storage device of claim 15 , wherein computing the reliability scores comprises:

obtaining one or more text fragments from a particular document;

determining a respective topic associated with each of the one or more text fragments; and

determining a likelihood that at least one of the topics are associated with the entity.

17. The storage device of claim 15 , wherein the operations further comprise:

computing a commonality score that represents a probability that a text fragment corresponding to a name of the entity will be included in a particular document included within the particular collection of documents; and

wherein the reliability scores that are computed for each document included within the particular collection of documents is computed based at least on the commonality score.

18. The storage device of claim 14 , wherein the concept scores are computed based at least on determining that the entity is included in a list of sanctioned entities.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUS RECORDATION OF ASSIGNMENT AGAINST U.S. APPLICATION NO. 17/365,807; ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF THE ENTIRE INTEREST PREVIOUSLY RECORDED ON REEL 68527 FRAME 467. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 15, 2025
From: GIANT OAK, INC.
To: FMR LLC
Reel/Frame 073136/0965 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2024
From: GIANT OAK, INC.
To: FMR LLC
Reel/Frame 068527/0467 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2017
From: SHIFFMAN, GARY; BOROWITZ, JEFFREY
To: GIANT OAK, INC.
Reel/Frame 041489/0987 →
Continuity (2)
Provisional Application 62304108 · Mar 4, 2016
Related Publication 20170255700A1 · Sep 7, 2017
Cited By (1)
US 12,314,330