IP Library › Granted Patent US 12,596,692
Granted Patent B2
US 12,596,692 · App. 18/019,736 · Granted Apr 7, 2026

Source scoring for entity representation systems

Inventors: Dwayne Collins (Conway, AR); Pavan Roy Marupally (Conway, AR)
Assignee: LiveRamp, Inc.
G06F16/215G06F16/288G06F16/9024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,692
App. No.
18/019,736
Granted
Apr 7, 2026
Kind
B2
Abstract

A method for computationally scoring the trustworthiness of source data begins with a subset of the raw source data. Fields and related fields are identified, profiled, and results aggregated. In addition, the subset of raw source data is input to an entity resolution system, with results summarized and aggregated. The output of these two streams of processes are used to compute a scorecard. In addition, a sandbox of the entity resolution system's data graph is constructed, the sandbox is modified with the source data, and the difference between the baseline sandbox and modified sandbox is computed. Changes to entities are computed for all entities as well as for most sought-after entities, and the results are summarized and aggregated into the overall scorecard.

Claims (22)

1 . A computerized method for performing source valuation within a parallel processing environment for a data source considered for inclusion in an entity resolution data graph, the method comprising the steps of:

using a plurality of processors in a parallel processing computer system to:

construct a first sample of the data source, wherein the first sample comprises a first subset of records in the data source associated with a dense representation of a geographically localized area;

construct a second sample of the data source, wherein the second sample comprises a second subset of records randomly selected from the data source, excluding the first subset of records;

automatically partition the first sample and second sample into a plurality of partitions using a partitioning algorithm that distributes data across multiple processors based on processing capacity, wherein each partition corresponds to an available processor within the parallel processing environment;

automatically identify corresponding relationships between sets of fields in the data source and nodes in the entity resolution data graph using automated field mapping algorithms that compare data structures and analyze entity relationship patterns;

execute in parallel across the plurality of processors a profiling process that analyzes one or more fields and the identified sets of fields to automatically generate sets of distributions, sets of counts, and sets of examples into a machine readable profiling scorecard that characterizes entity resolution quality metrics;

automatically aggregate the sets of distributions, set of counts, and set of examples into a machine-readable profiling scorecard that integrates entity resolution confidence scores with data quality assessments;

automatically compare the set of distributions, set of counts, and set of examples to at least one authoritative source distribution using statistical comparison algorithms to produce a quality score for determining data source trustworthiness, wherein the quality score combines data profiling results with entity matching validation;

pass the first sample and second sample into a match service, wherein the match service uses an identity graph comprising unique person identifiers, address identifiers, and household identifiers, and executes cascade level matching at a plurality of cascade levels;

return from the match service, for each record in the first sample and second sample, person identifiers for each of the plurality of cascade levels;

construct a key tuple for each record, wherein the key tuple comprises a plurality of sub-tuples each corresponding to a cascade level and containing person identifiers returned from the match service for that cascade level, wherein a first positive integer value indicates a known person in the entity resolution data graph, wherein different positive integer values for different cascade levels within a single record indicate that persons matched at different cascade levels may be the same person and represent a potential consolidation, and wherein a negative value indicates that a touchpoint is not in the entity resolution data graph and represents new information;

aggregate the key tuples into a distribution document whose values are counts of records sharing the same key; and

determine, from the distribution document, a level of trustworthiness of the data source and an expected impact on the entity resolution data graph if the data source is added, wherein the expected impact indicates at least one of: new touchpoints to be added, new relationships for known touchpoints, or potential consolidation of persons.

2 . The method of claim 1 , wherein the first subset of records comprises about two percent of a total number of records in the data source.

3 . The method of claim 1 , wherein the second subset of records comprises about three percent of a total number of records in the data source.

4 . The method of claim 1 , wherein the sets of fields comprise postal state and phone area code.

5 . The method of claim 1 , wherein the sets of fields comprise gender and age.

6 . The method of claim 1 , wherein the authoritative source distribution comprises US Census data.

7 . The method of claim 1 , wherein the authoritative source distribution comprises US Postal Service data.

8 . The method of claim 1 , wherein the plurality of matched cascade levels comprises a postal address only level.

9 . The method of claim 1 , wherein the plurality of matched cascade levels comprises a name plus telephone number level.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2023
From: COLLINS, DWAYNE; MARUPALLY, PAVAN ROY
To: LIVERAMP, INC.
Reel/Frame 062589/0361 →
Continuity (2)
Provisional Application 63063791 · Aug 10, 2020
Related Publication 20240256502A1 · Aug 1, 2024
References Cited (23)
US 9251470B2 · Hua et al. · 2016 [cited by applicant]
US 9275114B2 · Milton et al. · 2016 [cited by applicant]
US 9311301B1 · Balluru et al. · 2016 [cited by applicant]
US 9348891B2 · Srivastava et al. · 2016 [cited by applicant]
US 9659052B1 · Glennon · 2017 [cited by examiner]
US 9794358B1 · Jurgens · 2017 [cited by applicant]
US 10078679B1 · Shefferman et al. · 2018 [cited by applicant]
US 10339147B1 · Barmes · 2019 [cited by examiner]
US 12316610B1 · Muth · 2025 [cited by examiner]
US 20120278297A1 · Yin et al. · 2012 [cited by applicant]
US 20160371435A1 · Pauletto · 2016 [cited by examiner]
US 20170154058A1 · May · 2017 [cited by examiner]
US 20180203920A1 · Chen · 2018 [cited by examiner]
US 20190303494A1 · Meyer · 2019 [cited by examiner]
US 20190361853A1 · Lutsaievska · 2019 [cited by examiner]
US 20190392075A1 · Han · 2019 [cited by examiner]
US 20210042333A1 · Edwards · 2021 [cited by examiner]
US 20220164374A1 · Herrera · 2022 [cited by examiner]
Yin, Xiaoxin et al., “Semi-Supervised Truth Discovery,” WWW 2011-Session:Trust and Diversity (2011). [cited by applicant]
Deng, Yang et al., “MedTruth: A Semi-Supervised Approach to Discovering Knowledge Condition Information from Multi-Source Medical Data,” arXiv:1809.10404v2 (Aug. 19, 2019). [cited by applicant]
Fang, Xiu Susie et al., “Value Veracity Estimation for Multi-Truth Objects via a Graph-Based Approach,” IW3C2 Companion (Apr. 3-7, 2017). [cited by applicant]
Li, Yaliang et al., “A Survey on Truth Discovery,” arXiv:1505.02463v2 (Nov. 4, 2015). [cited by applicant]
Pasternack, Jeff et al., Knowing What to Believe (When You Already Know Something), Procs. of the 23rd Int'l Conf. on Computational Linguistics (coling 2010) (Aug. 2010). [cited by applicant]