IP Library › Granted Patent US 11,468,129
Granted Patent B2
US 11,468,129 · App. 16/444,444 · Granted Oct 11, 2022

Automatic false positive estimation for website matching

Inventors: Avishay Meron (Tel Aviv, IL); Tomer Handelman (Tel Aviv, IL); Shay Elbaz (Tel Aviv, IL); Shuly Lev-Yehudi (Tel Aviv, IL)
Assignee: PAYPAL, INC.
G06F16/951G06F16/958G06N7/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,468,129
App. No.
16/444,444
Granted
Oct 11, 2022
Kind
B2
Abstract

A system and method for generating automatic false positive estimations for website matching is described. Several sets of assets and Uniform Resource Locators (URLs) are aggregated. Each of the several sets of assets is searched across webpage content corresponding to the several URLs to determine matches between the sets of assets and webpage content. One or more false positive estimations is determined, where each of the one or more false positive estimations corresponds to the one or more matches. A combined score is generated based on the one or more false positive estimations.

Claims (37)

1. A system for performing an automatic false positive estimation for website matching comprising:

a non-transitory memory storing instructions; and

one or more hardware processors coupled to the non-transitory memory and configured to read the instructions from the non-transitory memory to cause the system to perform operations comprising:

aggregating a plurality of sets of assets and a plurality of uniform resource locators (URLs), wherein each set of assets and each URL are associated and correspond to one of a plurality of customer identifiers;

for each asset:

searching across webpage content corresponding to each of the plurality of URLs for one or more matches between the asset and the webpage content; and

calculating a false positive estimation based on a ratio of a number of times that the asset matched to webpage content of one URL to a number of times that the asset matched to webpage content of two or more URLs and subtracting the ratio from one;

for each set of assets:

generating a combined score based on the false positive estimations calculated for each asset in the set, wherein the combined score comprises a confidence score that an associated URL belongs to the corresponding customer identifier; and

determining at least one URL with a corresponding generated combined score exceeding a predetermined threshold, the at least one URL belonging to one of the plurality of customer identifiers; and

crawling webpages corresponding to the at least one URL to extract information associated with the one of the plurality of customer identifiers.

2. The system of claim 1 , wherein each set of the plurality of sets of assets comprises at least one of a name, a phone number, an address, or an email address.

3. The system of claim 1 , wherein the one or more matches are determined in a fuzzy manner.

4. The system of claim 3 , wherein a Levenshtein distance is used to determine the one or more matches in the fuzzy manner.

5. The system of claim 1 , wherein the false positive estimation is based on a top n results of the search across webpage content.

6. The system of claim 1 , wherein the combined score is based on a top n results of the search across webpage content.

7. A method for performing an automatic false positive estimation for website matching comprising:

aggregating a plurality of sets of assets and a plurality of uniform resource locators (URLs), wherein each set of assets and each URL are associated and correspond to one of a plurality of customer identifiers;

for each asset:

searching across webpage content corresponding to each of the plurality of URLs for one or more matches between the asset and the webpage content; and

calculating a false positive estimation by determining a ratio of a number of times that the asset matched to webpage content of one URL in relation to a number of times that the asset matched to webpage content of two or more URLs and subtracting the ratio from one;

for each set of assets:

generating a combined score based on the false positive estimations calculated for each asset in the set, wherein the combined score comprises a confidence score that an associated URL belongs to the corresponding customer identifier; and

determining at least one associated URL with a corresponding combined score exceeding a predetermined threshold, the at least one URL belonging to one of the plurality of customer identifiers; and

crawling webpages corresponding to the at least one URL to extract information associated with the one of the plurality of customer identifiers.

8. The method of claim 7 , wherein each set of the plurality of sets of assets comprises at least one of a name, a phone number, an address, or an email address.

9. The method of claim 7 , wherein the one or more matches are determined in a fuzzy manner.

10. The method of claim 9 , wherein a Levenshtein distance is used to determine the one or more matches in the fuzzy manner.

11. A non-transitory machine-readable medium having stored thereon machine-readable instructions executable to cause performance of operations comprising:

aggregating a plurality of sets of assets and a plurality of uniform resource locators (URLs), wherein each set of assets and each URL are associated and correspond to one of a plurality of customer identifiers;

for each asset:

searching across webpage content corresponding to each of the plurality of URLs for one or more matches between the asset and the webpage content; and

calculating a false positive estimation based on a ratio of a number of times that the asset matched to webpage content of one URL to a number of times that the asset matched to webpage content of two or more URLs, and subtracting the ratio from one;

for each set of assets, determining at least one URL with a corresponding generated combined score exceeding a predetermined threshold, the at least one URL belonging to one of the plurality of customer identifiers; and

crawling webpages corresponding to the at least one URL to extract information associated with the one of the plurality of customer identifiers.

12. The non-transitory machine-readable medium of claim 11 , wherein each set of the plurality of sets of assets comprises at least one of a name, a phone number, an address, or an email address.

13. The non-transitory machine-readable medium of claim 11 , wherein the one or more matches are determined in a fuzzy manner, and wherein a Levenshtein distance is used to determine the one or more matches in the fuzzy manner.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2019
From: MERON, AVISHAY; HANDELMAN, TOMER; ELBAZ, SHAY; LEV-YEHUDI, SHULY
To: PAYPAL, INC.
Reel/Frame 049504/0695 →
Continuity (1)
Related Publication 20200401633A1 · Dec 24, 2020