IP Library › Granted Patent US 11,163,840
Granted Patent B2
US 11,163,840 · App. 15/988,526 · Granted Nov 2, 2021

Systems and methods for intelligent content filtering and persistence

Inventors: Martin Brousseau (Mont-Saint-Hilaire, CA); Steve Pettigrew (Montreal, CA)
Assignee: OPEN TEXT SA ULC
G06F16/9535G06F16/9558G06F40/295G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,163,840
App. No.
15/988,526
Granted
Nov 2, 2021
Kind
B2
Abstract

A source content processor receives content from a crawler and calls a text mining engine. The text mining engine mines the content and provides metadata about the content. The source content processor applies a source content filtering rule to the content utilizing the metadata from the text mining engine. The source content filtering rule is previously built based on at least one of a named entity, a category, or a sentiment. The source content processor determines whether to persist the content according to a result from applying the source content filtering rule to the content and either stores the content in a data store or deletes the contents from the data ingestion pipeline such that the content is not persisted anywhere. Embodiments disclosed herein can significantly reduce the amount of irrelevant content through the data ingestion pipeline, prior to data persistence.

Claims (48)

1. A method, comprising:

receiving, by a source content processor, content from a crawler, the source content processor working in conjunction with the crawler and a data ingestion pipeline running on a server machine, the crawler communicatively connected to a data store through the data ingestion pipeline, the server machine operating in an enterprise computing environment;

prior to persisting the content from the crawler, calling, by the source content processor, a text mining engine with the content from the crawler;

receiving, by the source content processor from the text mining engine, metadata that describes the content from the crawler;

applying, by the source content processor, a source content filtering rule to the content from the crawler utilizing the metadata that describes the content from the crawler, wherein the source content filtering rule is previously built based on at least one of a named entity, a category, and a sentiment;

determining, by the source content processor, whether to persist the content from the crawler according to a result from the applying; and

responsive to a determination by the source content processor to persist the content from the crawler, storing the content in the data store.

2. The method according to claim 1 , further comprising:

accessing a source content filtering rules database; and

retrieving the source content filtering rule from the source content filtering rules database based on a type of the metadata.

3. The method according to claim 1 , wherein the metadata from the text mining engine comprise named entities, categories, and sentiments.

4. The method according to claim 1 , wherein responsive to a determination by the source content processor not to persist the content from the crawler, the source content processor is operable to push the content to a file, generate a link to the file, and store the link.

5. The method according to claim 1 , wherein the data store comprises a relational database management system, a data store, or a content repository.

6. The method according to claim 1 , wherein responsive to a determination by the source content processor not to persist the content from the crawler, the source content processor deletes the content from the data ingestion pipeline such that the content is not persisted anywhere in the enterprise computing environment.

7. A system, comprising:

a processor;

a non-transitory computer-readable medium; and

stored instructions translatable by the processor to implement a source content filter for:

receiving content from a crawler, the source content processor working in conjunction with the crawler and a data ingestion pipeline running on the system, the crawler communicatively connected to a data store through the data ingestion pipeline;

prior to persisting the content from the crawler, calling a text mining engine with the content from the crawler;

receiving, from the text mining engine, metadata that describes the content from the crawler;

applying a source content filtering rule to the content from the crawler utilizing the metadata that describes the content from the crawler, wherein the source content filtering rule is previously built based on at least one of a named entity, a category, and a sentiment;

determining whether to persist the content from the crawler according to a result from the applying; and

responsive to a determination to persist the content from the crawler, storing the content in the data store.

8. The system of claim 7 , wherein the stored instructions are further translatable by the processor to perform:

accessing a source content filtering rules database; and

retrieving the source content filtering rule from the source content filtering rules database based on a type of the metadata.

9. The system of claim 7 , wherein the metadata from the text mining engine comprise named entities, categories, and sentiments.

10. The system of claim 7 , wherein the stored instructions are further translatable by the processor to perform:

responsive to a determination not to persist the content from the crawler, pushing the content to a file, generating a link to the file, and storing the link.

11. The system of claim 7 , wherein the data store comprises a relational database management system, a data store, or a content repository.

12. The system of claim 7 , wherein the stored instructions are further translatable by the processor to perform:

responsive to a determination not to persist the content from the crawler, deleting the content from the data ingestion pipeline such that the content is not persisted anywhere on the system.

13. A computer program product comprising a non-transitory computer-readable medium storing instructions translatable by a processor to implement a source content filter for:

receiving content from a crawler, the source content processor working in conjunction with the crawler and a data ingestion pipeline running on a server machine, the crawler communicatively connected to a data store through the data ingestion pipeline, the server machine operating in an enterprise computing environment;

prior to persisting the content from the crawler, calling a text mining engine with the content from the crawler;

receiving, from the text mining engine, metadata that describes the content from the crawler;

applying a source content filtering rule to the content from the crawler utilizing the metadata that describes the content from the crawler, wherein the source content filtering rule is previously built based on at least one of a named entity, a category, and a sentiment;

determining whether to persist the content from the crawler according to a result from the applying; and

responsive to a determination to persist the content from the crawler, storing the content in the data store.

14. The computer program product of claim 13 , wherein the instructions are further translatable by the processor to perform:

accessing a source content filtering rules database; and

retrieving the source content filtering rule from the source content filtering rules database based on a type of the metadata.

15. The computer program product of claim 13 , wherein the metadata from the text mining engine comprise named entities, categories, and sentiments.

16. The computer program product of claim 13 , wherein the instructions are further translatable by the processor to perform:

responsive to a determination not to persist the content from the crawler, pushing the content to a file, generating a link to the file, and storing the link.

17. The computer program product of claim 13 , wherein the instructions are further translatable by the processor to perform:

responsive to a determination not to persist the content from the crawler, deleting the content from the data ingestion pipeline such that the content is not persisted anywhere in the enterprise computing environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2018
From: BROUSSEAU, MARTIN; PETTIGREW, STEVE
To: OPEN TEXT SA ULC
Reel/Frame 045896/0312 →
Continuity (1)
Related Publication 20190362024A1 · Nov 28, 2019
Cited By (1)
US 12,306,884