IP Library Granted Patent US 10,621,182
Granted Patent B2
US 10,621,182 · App. 14/844,623 · Granted Apr 14, 2020

System and process for analyzing, qualifying and ingesting sources of unstructured data via empirical attribution

Inventors: Anthony J. Scriffignano (West Caldwell, NJ); Yiem Sunbhanich (Lewis Center, OH); Robin Fry Davies (Plymouth, MI); Warwick Matthews (Glen Iris, AU)
Assignee: THE DUN & BRADSTREET CORPORATION
G06F16/24578G06F17/2211
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,621,182
App. No.
14/844,623
Granted
Apr 14, 2020
Kind
B2
Abstract

There is provided a method that includes (a) receiving data from a data source, (b) attributing the data source in accordance with rules, thus yielding an attribute, (c) analyzing the data to identify a confounding characteristic in the data, (d) calculating a qualitative measure of the attribute, thus yielding a weighted attribute, (e) calculating a qualitative measure of the confounding characteristic, thus yielding a weighted confounding characteristic, (f) analyzing the weighted attribute and the weighted confounding characteristic, to produce a disposition, (g) filtering the data in accordance with the disposition, thus yielding extracted data, and (h) transmitting the extracted data to a downstream process. There is also provided a system that executes the method, and a storage device that contains instructions for controlling a processor to perform the method.

Claims (74)

1. A method comprising:

(a) receiving data from a first data source;

(b) performing a source analysis process that includes:

attributing said first data source at a context level, a source file level, and a content level, in accordance with rules, thus yielding an attribute of said first data source;

analyzing said data to identify a characteristic of said data that confounds a meaning of said data, thus yielding a confounding characteristic;

calculating a qualitative measure of said attribute of said first data source, thus yielding a weighted attribute of said first data source;

calculating a qualitative measure of said confounding characteristic, thus yielding a weighted confounding characteristic; and

analyzing said weighted attribute of said first data source and said weighted confounding characteristic, to produce a disposition that includes disposition instructions;

(c) processing said data in accordance with said disposition instructions, thus yielding extracted data;

(d) transmitting said extracted data to a downstream process;

(e) determining, based on said weighted attribute of said first data source and said weighted confounding characteristic, to use said first data source as a seed of an automated data source discovery process;

(f) executing said automated data discovery process using search terms based on content of said first data source to discover a second data source; and

(g) performing said source analysis process on data from said second data source.

2. The method of claim 1 , further comprising:

generating feedback based on said disposition; and

improving said method based on said feedback.

3. The method of claim 1 , wherein said analyzing is performed in a dimension selected from the group consisting of entity extraction, semantic disambiguation, sentiment analysis, language extraction, linguistic transformation, and basic metadata.

4. The method of claim 1 , wherein said confounding characteristic is selected from the group consisting of sarcasm, neologism, grammar variation, inappropriately phrased text, punctuation, polylingual data, spelling, obfuscation, encryption, context, and use of a combination of media.

5. The method of claim 1 , wherein said disposition is selected from the group consisting of:

(a) set a rule that files analogous to said first data source be taken in and stored in toto,

(b) divide up files from said first data source and take in and store only parts that meet certain criteria,

(c) take in and store an entire file from said first data source, but flag data with a source-specific quality level indicator,

(d) set a rule that files from said first data source always be rejected, and

(e) tentatively take in and store files from said first data source, but hold them pending additional corroboration.

6. A system comprising:

a processor; and

a memory that contains instructions that are readable by said processor to cause said processor to:

(a) receive data from a first data source;

(b) perform a source analysis process in which said instructions cause said processor to:

attribute said first data source at a context level, a source file level, and a content level, in accordance with rules, thus yielding an attribute of said first data source;

analyze said data to identify a characteristic of said data that confounds a meaning of said data, thus yielding a confounding characteristic;

calculate a qualitative measure of said attribute of said first data source, thus yielding a weighted attribute of said first data source;

calculate a qualitative measure of said confounding characteristic, thus yielding a weighted confounding characteristic; and

analyze said weighted attribute of said first data source and said weighted confounding characteristic, to produce a disposition that includes disposition instructions;

(c) process said data in accordance with said disposition instructions, thus yielding extracted data;

(d) transmit said extracted data to a downstream process;

(e) determine, based on said weighted attribute of said first data source and said weighted confounding characteristic, to use said first data source as a seed of an automated data source discovery process;

(f) execute said automated data discovery process using search terms based on content of said first data source to discover a second data source; and

(g) perform said source analysis process on data from said second data source.

7. The system of claim 6 , wherein said instructions also cause said processor to:

generate feedback based on said disposition; and

improve said method based on said feedback.

8. The system of claim 6 , wherein said instructions that cause said processor to analyze said data cause said processor to analyze said data in a dimension selected from the group consisting of entity extraction, semantic disambiguation, sentiment analysis, language extraction, linguistic transformation, and basic metadata.

9. The system of claim 6 , wherein said confounding characteristic is selected from the group consisting of sarcasm, neologism, grammar variation, inappropriately phrased text, punctuation, polylingual data, spelling, obfuscation, encryption, context, and use of a combination of media.

10. The system of claim 6 , wherein said disposition is selected from the group consisting of:

(a) set a rule that files analogous to said first data source be taken in and stored in toto,

(b) divide up files from said first data source and take in and store only parts that meet certain criteria,

(c) take in and store an entire file from said first data source, but flag data with a source-specific quality level indicator,

(d) set a rule that files from said first data source always be rejected, and

(e) tentatively take in and store files from said first data source, but hold them pending additional corroboration.

11. A non-transitory computer readable storage medium, comprising instructions that are readable by a processor to cause said processor to:

(a) receive data from a first data source;

(b) perform a source analysis process in which said instructions cause said processor to:

attribute said first data source at a context level, a source file level, and a content level, in accordance with rules, thus yielding an attribute of said first data source;

analyze said data to identify a characteristic of said data that confounds a meaning of said data, thus yielding a confounding characteristic;

calculate a qualitative measure of said attribute of said first data source, thus yielding a weighted attribute of said first data source;

calculate a qualitative measure of said confounding characteristic, thus yielding a weighted confounding characteristic; and

analyze said weighted attribute of said first data source and said weighted confounding characteristic, to produce a disposition that includes disposition instructions;

(c) process said data in accordance with said disposition instructions, thus yielding extracted data;

(d) transmit said extracted data to a downstream process;

(e) determine, based on said weighted attribute of said first data source and said weighted confounding characteristic, to use said first data source as a seed of an automated data source discovery process;

(f) execute said automated data discovery process using search terms based on content of said first data source to discover a second data source; and

(g) perform said source analysis process on data from said second data source.

12. The non-transitory computer readable storage medium of claim 11 , wherein said instructions also cause said processor to:

generate feedback based on said disposition; and

improve said method based on said feedback.

13. The non-transitory computer readable storage medium of claim 11 , wherein said instructions that cause said processor to analyze said data cause said processor to analyze said data in a dimension selected from the group consisting of entity extraction, semantic disambiguation, sentiment analysis, language extraction, linguistic transformation, and basic metadata.

14. The non-transitory computer readable storage medium of claim 11 , wherein said confounding characteristic is selected from the group consisting of sarcasm, neologism, grammar variation, inappropriately phrased text, punctuation, polylingual data, spelling, obfuscation, encryption, context, and use of a combination of media.

15. The non-transitory computer readable storage medium of claim 11 , wherein said disposition is selected from the group consisting of:

(a) set a rule that files analogous to said first data source be taken in and stored in toto,

(b) divide up files from said first data source and take in and store only parts that meet certain criteria,

(c) take in and store an entire file from said first data source, but flag data with a source-specific quality level indicator,

(d) set a rule that files from said first data source always be rejected, and

(e) tentatively take in and store files from said first data source, but hold them pending additional corroboration.

Assignments (6)
RELEASE OF SECURITY INTEREST Recorded Aug 27, 2025
From: BANK OF AMERICA, N.A. AS AGENT
To: THE DUN & BRADSTREET CORPORATION; DUN & BRADSTREET EMERGING BUSINESSES CORP.; DUN & BRADSTREET, INC.; HOOVER’S, INC.; LATTICE ENGINES, INC.
Reel/Frame 072591/0843 →
SECURITY INTEREST Recorded Aug 27, 2025
From: DUN & BRADSTREET EMERGING BUSINESSES CORP.; DUN & BRADSTREET, INC.; LATTICE ENGINES, INC.; THE DUN AND BRADSTREET CORPORATION
To: ARES CAPITAL CORPORATION, AS COLLATERAL AGENT
Reel/Frame 072643/0196 →
INTELLECTUAL PROPERTY RELEASE AND TERMINATION Recorded Jan 18, 2022
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: THE DUN & BRADSTREET CORPORATION; DUN & BRADSTREET EMERGING BUSINESSES CORP.; DUN & BRADSTREET, INC.; HOOVER'S, INC.
Reel/Frame 058757/0232 →
PATENT SECURITY AGREEMENT Recorded Feb 12, 2019
From: THE DUN & BRADSTREET CORPORATION; DUN & BRADSTREET EMERGING BUSINESSES CORP.; DUN & BRADSTREET, INC.; HOOVER'S, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 048306/0375 →
PATENT SECURITY AGREEMENT Recorded Feb 12, 2019
From: THE DUN & BRADSTREET CORPORATION; DUN & BRADSTREET EMERGING BUSINESSES CORP.; DUN & BRADSTREET, INC.; HOOVER'S INC.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 048306/0412 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2015
From: SCRIFFIGNANO, ANTHONY J.; SUNBHANICH, YIEM; DAVIES, ROBIN FRY; MATTHEWS, WARWICK
To: THE DUN & BRADSTREET CORPORATION
Reel/Frame 036773/0272 →