IP Library › Granted Patent US 10,785,188
Granted Patent B2
US 10,785,188 · App. 15/986,585 · Granted Sep 22, 2020

Domain name processing systems and methods

Inventors: Harold Nguyen (Burlingame, CA); Ali Mesdaq (San Jose, CA); Kevin Dedon (Austin, TX); Michael Fox (Lago Vista, TX); Gaurav Dalal (Fremont, CA)
Assignee: Proofpoint, Inc.
H04L61/2046G06F16/9535H04L61/1511
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,785,188
App. No.
15/986,585
Granted
Sep 22, 2020
Kind
B2
Abstract

Disclosed is a domain filter capable of determining an n-gram distance between a seed domain and each of a plurality of candidate domains. The domain filter loads a seed domain n-gram for the seed domain and a candidate domain n-gram for each candidate domain in memory, compares the seed domain n-gram and the candidate domain n-gram to identify any identical grams, removes any identical grams from the seed domain n-gram, and determines how many grams are left in the seed domain n-gram, representing the n-gram distance between the seed domain and the candidate domain. The domain filter then compares n-gram distances thus determined with a predetermined threshold, eliminates any candidate domain having an n-gram distance from the seed domain that exceeds the predetermined threshold, and provides remaining candidate domains to a downstream computing facility such as a user interface or an analytical module operating in an enterprise computing environment.

Claims (69)

1. A method, comprising:

determining an n-gram distance between a seed domain and each candidate domain of a plurality of candidate domains, the determining performed by a domain filter running on a computing device having a computer memory, the determining comprising:

loading a seed domain n-gram for the seed domain and a candidate domain n-gram for the each candidate domain in the computer memory;

comparing the seed domain n-gram and the candidate domain n-gram to identify any identical grams in the seed domain n-gram and the candidate domain n-gram;

removing any identical grams from the seed domain n-gram in the computer memory; and

counting a number of grams left in the seed domain n-gram in the computer memory after the removing, the number representing the n-gram distance between the seed domain and the each candidate domain;

comparing n-gram distances determined by the domain filter with a predetermined threshold;

eliminating, from the plurality of candidate domains, any candidate domain having an n-gram distance from the seed domain that exceeds the predetermined threshold; and

providing candidate domains left from the eliminating to a downstream computing facility.

2. The method according to claim 1 , further comprising:

performing the determining for each of a plurality of seed domains.

3. The method according to claim 1 , further comprising:

accessing a seed domain database;

retrieving a seed domain name for the seed domain from the seed domain database, the seed domain name comprising a character string;

generating an original seed domain n-gram from the character string; and

storing the original seed domain n-gram in the computer memory of the computing device, wherein the loading comprises making a copy of the original seed domain n-gram and storing the copy in the computer memory.

4. The method according to claim 3 , wherein the computing device comprises a mobile device, a laptop computer, or a tablet computer, wherein the seed domain database resides on a server machine operating in an enterprise computing environment and wherein the retrieving comprises obtaining the seed domain name for the seed domain from the seed domain database over a secure network connection.

5. The method according to claim 1 , further comprising:

accessing an Internet domain database;

retrieving candidate domain names for the plurality of candidate domains from the Internet domain database, each candidate domain name comprising a character string; and

generating the candidate domain n-gram from the character string.

6. The method according to claim 1 , wherein the downstream computing facility comprises a user interface, an edit distance analyzer, or an analytical module running on a computer operating in an enterprise computing environment.

7. The method according to claim 1 , wherein the computing device comprises a mobile device, a laptop computer, or a tablet computer.

8. An apparatus, comprising:

a processor;

a computer memory; and

stored instructions translatable by the processor to perform:

determining an n-gram distance between a seed domain and each candidate domain of a plurality of candidate domains, the determining comprising:

loading a seed domain n-gram for the seed domain and a candidate domain n-gram for the each candidate domain in the computer memory;

comparing the seed domain n-gram and the candidate domain n-gram to identify any identical grams in the seed domain n-gram and the candidate domain n-gram;

removing any identical grams from the seed domain n-gram in the computer memory; and

counting a number of grams left in the seed domain n-gram in the computer memory after the removing, the number representing the n-gram distance between the seed domain and the each candidate domain;

comparing n-gram distances from the determining with a predetermined threshold;

eliminating, from the plurality of candidate domains, any candidate domain having an n-gram distance from the seed domain that exceeds the predetermined threshold; and

providing candidate domains left from the eliminating to a downstream computing facility.

9. The apparatus of claim 8 , wherein the stored instructions are further translatable by the processor to perform the determining for each of a plurality of seed domains.

10. The apparatus of claim 8 , wherein the stored instructions are further translatable by the processor to perform:

accessing a seed domain database;

retrieving a seed domain name for the seed domain from the seed domain database, the seed domain name comprising a character string;

generating an original seed domain n-gram from the character string; and

storing the original seed domain n-gram in the computer memory of the computing device, wherein the loading comprises making a copy of the original seed domain n-gram and storing the copy in the computer memory.

11. The apparatus of claim 10 , wherein the apparatus comprises a mobile device, a laptop computer, or a tablet computer, wherein the seed domain database resides on a server machine operating in an enterprise computing environment and wherein the retrieving comprises obtaining the seed domain name for the seed domain from the seed domain database over a secure network connection.

12. The apparatus of claim 8 , wherein the stored instructions are further translatable by the processor to perform:

accessing an Internet domain database;

retrieving candidate domain names for the plurality of candidate domains from the Internet domain database, each candidate domain name comprising a character string; and

generating the candidate domain n-gram from the character string.

13. The apparatus of claim 8 , wherein the downstream computing facility comprises a user interface, an edit distance analyzer, or an analytical module running on a computer operating in an enterprise computing environment.

14. The apparatus of claim 10 , wherein the apparatus comprises a mobile device, a laptop computer, or a tablet computer.

15. A computer program product comprising a non-transitory computer-readable medium storing instructions translatable by a computing device having a computer memory to perform:

determining an n-gram distance between a seed domain and each candidate domain of a plurality of candidate domains, the determining comprising:

loading a seed domain n-gram for the seed domain and a candidate domain n-gram for the each candidate domain in the computer memory;

comparing the seed domain n-gram and the candidate domain n-gram to identify any identical grams in the seed domain n-gram and the candidate domain n-gram;

removing any identical grams from the seed domain n-gram in the computer memory; and

counting a number of grams left in the seed domain n-gram in the computer memory after the removing, the number representing the n-gram distance between the seed domain and the each candidate domain;

comparing n-gram distances from the determining with a predetermined threshold;

eliminating, from the plurality of candidate domains, any candidate domain having an n-gram distance from the seed domain that exceeds the predetermined threshold; and

providing candidate domains left from the eliminating to a downstream computing facility.

16. The computer program product of claim 15 , wherein the instructions are further translatable by the computing device to perform the determining for each of a plurality of seed domains.

17. The computer program product of claim 15 , wherein the instructions are further translatable by the computing device to perform:

accessing a seed domain database;

retrieving a seed domain name for the seed domain from the seed domain database, the seed domain name comprising a character string;

generating an original seed domain n-gram from the character string; and

storing the original seed domain n-gram in the computer memory of the computing device, wherein the loading comprises making a copy of the original seed domain n-gram and storing the copy in the computer memory.

18. The computer program product of claim 15 , wherein the computing device comprises a mobile device, a laptop computer, or a tablet computer, wherein the seed domain database resides on a server machine operating in an enterprise computing environment and wherein the retrieving comprises obtaining the seed domain name for the seed domain from the seed domain database over a secure network connection.

19. The computer program product of claim 15 , wherein the instructions are further translatable by the computing device to perform:

accessing an Internet domain database;

retrieving candidate domain names for the plurality of candidate domains from the Internet domain database, each candidate domain name comprising a character string; and

generating the candidate domain n-gram from the character string.

20. The computer program product of claim 15 , wherein the downstream computing facility comprises a user interface, an edit distance analyzer, or an analytical module running on a computer operating in an enterprise computing environment.

Assignments (5)
SECOND LIEN INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 8, 2025
From: PROOFPOINT, INC.
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 073889/0677 →
RELEASE OF SECOND LIEN SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded Mar 21, 2024
From: GOLDMAN SACHS BANK USA, AS AGENT
To: PROOFPOINT, INC.
Reel/Frame 066865/0648 →
FIRST LIEN INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Aug 31, 2021
From: PROOFPOINT, INC.
To: GOLDMAN SACHS BANK USA, AS COLLATERAL AGENT
Reel/Frame 057389/0615 →
SECOND LIEN INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Aug 31, 2021
From: PROOFPOINT, INC.
To: GOLDMAN SACHS BANK USA, AS COLLATERAL AGENT
Reel/Frame 057389/0642 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2018
From: NGUYEN, HAROLD; MESDAQ, ALI; DEDON, KEVIN; FOX, MICHAEL; DALAL, GAURAV
To: PROOFPOINT, INC.
Reel/Frame 045876/0429 →
Continuity (1)
Related Publication 20190364011A1 · Nov 28, 2019
Cited By (1)
US 12,323,460