IP Library Granted Patent US 11,171,916
Granted Patent B2
US 11,171,916 · App. 16/866,297 · Granted Nov 9, 2021

Domain name classification systems and methods

Inventors: Sharon Huffner (Munich, DE); Ali Mesdaq (San Jose, CA)
Assignee: Proofpoint, Inc.
H04L61/3005H04L61/1511H04L63/1483
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,171,916
App. No.
16/866,297
Granted
Nov 9, 2021
Kind
B2
Abstract

Disclosed is a domain engineering analysis solution that determines relevance of a domain name to a brand name in which a domain name, brand name, and identification of a substring of the domain name may be provided to or obtained by a computer embodying a domain engineering analyzer. A list of features may be determined. The list of features may include a lexicon, or a set of key-value pairs that encode information about terms included as substrings in the domain name. Determining the features may include obtaining a language model for each term, analyzing a cluster of language models closest to the obtained language model, and determining and scoring a relevance of each term to the brand name. The determined relevance and score of each term may be provided to a client. This relevance analysis can be dynamically applied in an online process or proactively applied in an offline process.

Claims (74)

1. A computer-implemented method for domain name classification, the method comprising:

obtaining, by a pre-processing engine executing on a processor, input domain names from multiple data sources, the input domain names associated with a brand or entity;

extracting, by the pre-processing engine utilizing a lexicon for the brand or entity, a set of words of interest to the brand or entity from the input domain names;

determining, by the pre-processing engine, a word vector for each respective word of the set of words extracted from the input domain names;

analyzing, by the pre-processing engine, the word vector for each respective word of the set of words extracted from the input domain names;

determining, by the pre-processing engine based on the analyzing, a list of true positives;

determining, by the pre-processing engine from the input domain names, a list of domain names containing a substring that is an exact match of the brand or entity;

determining, by an analyzer executing on a processor utilizing the list of true positives and the list of domain names, candidate words included as substrings in a candidate domain;

obtaining, by the analyzer, a word vector for each respective candidate word of the candidate words included as substrings in the candidate domain;

determining, by the analyzer for each respective candidate word of the candidate words included as substrings in the candidate domain, neighbors of the word vector;

determining, by the analyzer, a percentage of the neighbors that belong to the brand or entity;

determining, by the analyzer for each respective candidate word of the candidate words included as substrings in the candidate domain, a score based at least on the percentage of the neighbors that belong to the brand or entity; and

classifying, by a classifier, the candidate domain based at least on the score.

2. The computer-implemented method according to claim 1 , further comprising:

determining, by the analyzer, a percentage of the neighbors that are not of interest to the brand or entity, wherein the score is determined further based on the percentage of the neighbors that are not of interest to the brand or entity.

3. The computer-implemented method according to claim 1 , further comprising:

obtaining entries pertaining to the brand or the entity from disparate sources on the Internet; and

generating the lexicon for the brand or entity from the entries.

4. The computer-implemented method according to claim 1 , further comprising:

generating key-value pairs based on the candidate words included as substrings in the candidate domain.

5. The computer-implemented method according to claim 4 , further comprising:

providing the key-value pairs and relevance scores corresponding to the candidate words to a user device.

6. The computer-implemented method according to claim 1 , wherein the input domain names are registered domain names of the brand or entity.

7. The computer-implemented method according to claim 1 , wherein the input domain names are obtained from the multiple data sources by request of the pre-processing engine, scheduled for periodic download, or on demand.

8. A system for domain name classification, the system comprising:

a processor;

a non-transitory computer-readable medium; and

stored instructions translatable by the processor for:

obtaining input domain names from multiple data sources, the input domain names associated with a brand or entity;

extracting, utilizing a lexicon for the brand or entity, a set of words of interest to the brand or entity from the input domain names;

determining a word vector for each respective word of the set of words extracted from the input domain names;

analyzing the word vector for each respective word of the set of words extracted from the input domain names;

determining, based on the analyzing, a list of true positives;

determining, from the input domain names, a list of domain names containing a substring that is an exact match of the brand or entity;

determining, utilizing the list of true positives and the list of domain names, candidate words included as substrings in a candidate domain;

obtaining a word vector for each respective candidate word of the candidate words included as substrings in the candidate domain;

determining, for each respective candidate word of the candidate words included as substrings in the candidate domain, neighbors of the word vector;

determining a percentage of the neighbors that belong to the brand or entity;

determining, for each respective candidate word of the candidate words included as substrings in the candidate domain, a score based at least on the percentage of the neighbors that belong to the brand or entity; and

classifying the candidate domain based at least on the score.

9. The system of claim 8 , wherein the stored instructions are further translatable by the processor for:

determining a percentage of the neighbors that are not of interest to the brand or entity, wherein the score is determined further based on the percentage of the neighbors that are not of interest to the brand or entity.

10. The system of claim 8 , wherein the stored instructions are further translatable by the processor for:

obtaining entries pertaining to the brand or the entity from disparate sources on the Internet; and

generating the lexicon for the brand or entity from the entries.

11. The system of claim 8 , wherein the stored instructions are further translatable by the processor for:

generating key-value pairs based on the candidate words included as substrings in the candidate domain.

12. The system of claim 11 , wherein the stored instructions are further translatable by the processor for:

providing the key-value pairs and relevance scores corresponding to the candidate words to a user device.

13. The system of claim 8 , wherein the input domain names are registered domain names of the brand or entity.

14. The system of claim 8 , wherein the input domain names are obtained from the multiple data sources by request, scheduled for periodic download, or on demand.

15. A computer program product for domain name classification, the computer program product comprising a non-transitory computer-readable medium storing instructions translatable by a processor for:

obtaining input domain names from multiple data sources, the input domain names associated with a brand or entity;

extracting, utilizing a lexicon for the brand or entity, a set of words of interest to the brand or entity from the input domain names;

determining a word vector for each respective word of the set of words extracted from the input domain names;

analyzing the word vector for each respective word of the set of words extracted from the input domain names;

determining, based on the analyzing, a list of true positives;

determining, from the input domain names, a list of domain names containing a substring that is an exact match of the brand or entity;

determining, utilizing the list of true positives and the list of domain names, candidate words included as substrings in a candidate domain;

obtaining a word vector for each respective candidate word of the candidate words included as substrings in the candidate domain;

determining, for each respective candidate word of the candidate words included as substrings in the candidate domain, neighbors of the word vector;

determining a percentage of the neighbors that belong to the brand or entity;

determining, for each respective candidate word of the candidate words included as substrings in the candidate domain, a score based at least on the percentage of the neighbors that belong to the brand or entity; and

classifying the candidate domain based at least on the score.

16. The computer program product of claim 15 , wherein the instructions are further translatable by the processor for:

determining a percentage of the neighbors that are not of interest to the brand or entity, wherein the score is determined further based on the percentage of the neighbors that are not of interest to the brand or entity.

17. The computer program product of claim 15 , wherein the instructions are further translatable by the processor for:

obtaining entries pertaining to the brand or the entity from disparate sources on the Internet; and

generating the lexicon for the brand or entity from the entries.

18. The computer program product of claim 15 , wherein the instructions are further translatable by the processor for:

generating key-value pairs based on the candidate words included as substrings in the candidate domain.

19. The computer program product of claim 18 , wherein the instructions are further translatable by the processor for:

providing the key-value pairs and relevance scores corresponding to the candidate words to a user device.

20. The computer program product of claim 15 , wherein the input domain names are obtained from the multiple data sources by request, scheduled for periodic download, or on demand.

Assignments (5)
SECOND LIEN INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 8, 2025
From: PROOFPOINT, INC.
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 073889/0677 →
RELEASE OF SECOND LIEN SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded Mar 21, 2024
From: GOLDMAN SACHS BANK USA, AS AGENT
To: PROOFPOINT, INC.
Reel/Frame 066865/0648 →
FIRST LIEN INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Aug 31, 2021
From: PROOFPOINT, INC.
To: GOLDMAN SACHS BANK USA, AS COLLATERAL AGENT
Reel/Frame 057389/0615 →
SECOND LIEN INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Aug 31, 2021
From: PROOFPOINT, INC.
To: GOLDMAN SACHS BANK USA, AS COLLATERAL AGENT
Reel/Frame 057389/0642 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2020
From: HÜFFNER, SHARON; MESDAQ, ALI
To: PROOFPOINT, INC.
Reel/Frame 052939/0084 →
Continuity (2)
Continuation 15687660 · Aug 28, 2017
Related Publication 20200267119A1 · Aug 20, 2020
Cited By (2)
US 12,335,306 US 12,568,116