IP Library Patent Application 16709637
Patent Application
App. No. 16/709,637

CLASSIFYING DOMAIN NAMES BASED ON CHARACTER EMBEDDING AND DEEP LEARNING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/709,637
Filed
Dec 10, 2019
Art Unit
2657
USPC
706/20
Abstract

An apparatus may include a processor that may be caused to access a plurality of known domain names. The processor may be caused to determine a character embedding based on the plurality of known domain names. The character embedding may map each character of a known domain name to a respective vector. The processor may be caused to input the character embedding to a deep learning layer of a neural network. The processor may be caused to access a target domain name to be classified. The processor may be caused to classify the target domain name based on an output of the deep learning layer.

Claims (49)

1 . An apparatus comprising:

a processor; and

a non-transitory machine-readable storage medium on which is stored instructions that when executed by the processor, cause the processor to:

access a plurality of known domain names;

determine a character embedding based on the plurality of known domain names, the character embedding mapping each character of a known domain name to a respective vector;

input the character embedding to a deep learning layer of a neural network;

access a target domain name to be classified; and

classify the target domain name based on an output of the deep learning layer.

2 . The apparatus of claim 1 , wherein to determine the character embedding, the processor is further caused to:

for each character of the known domain name, identify N continuous characters that neighbor the character in the known domain name, wherein N represents a number of continuous characters.

3 . The apparatus of claim 2 , wherein the processor is further caused to:

determine similarities among the N continuous characters with other continuous characters in the plurality of known domain names that neighbor other characters in the plurality of known domain names.

4 . The apparatus of claim 3 , wherein to determine the similarities, the processor is further caused to:

for each character, determine similarities among the N continuous characters that precede the character and the other continuous characters that precede the other characters.

5 . The apparatus of claim 3 , wherein to determine the similarities, the processor is further caused to:

for each character, determine similarities among the N continuous characters that follow the character and the other continuous characters that follow the other characters.

6 . The apparatus of claim 1 , wherein the deep learning layer comprises a Long Short-Term Memory (LSTM) layer.

7 . The apparatus of claim 1 , wherein the processor is further caused to:

provide the output of the deep learning layer to a classifier layer that classifies the target domain name.

8 . The apparatus of claim 7 , wherein to classify the target domain name, the processor is further caused to:

determine, based on an output of the classifier layer, whether or not the target domain name is associated with a malicious class of domain names.

9 . The apparatus of claim 7 , wherein the classifier layer comprises a softmax layer that determines a first probability that the target domain name is a malicious domain name, a second probability that the target domain name is a non-algorithmically-generated benign domain name, and a third probability that the target domain name is an algorithmically-generated benign domain name.

10 . The apparatus of claim 9 , wherein to access the plurality of known domain names, the processor is caused to:

access a first plurality of malicious domain names;

access a second plurality of non-algorithmically-generated benign domain names; and

access a third plurality of algorithmically-generated benign domain names.

11 . The apparatus of claim 1 , wherein the deep learning layer is trained without manual feature generation.

12 . A method, comprising:

learning, by a processor, a character embedding from a plurality of known domain names;

providing, by the processor, the character embedding as an input to a Long Short-Term Memory (LSTM) layer;

accessing, by the processor, a target domain name to be classified; and

classifying, by the processor, the target domain name via a fully connected softmax layer.

13 . The method of claim 12 , wherein learning the character embedding comprises determining the character embedding in a reverse direction.

14 . The method of claim 12 , wherein learning the character embedding comprises determining the character embedding in a forward direction.

15 . The method of claim 12 , wherein classifying the target domain name comprises:

providing an output of the LSTM to a softmax layer that classifies the target domain name into one or more of a plurality of classes.

16 . The method of claim 15 , wherein the plurality of classes comprises a malicious domain name class, a non-algorithmically-generated benign domain name class, and an algorithmically-generated benign domain name class.

17 . A non-transitory machine-readable storage medium on which is stored machine-readable instructions that when executed by a processor, cause the processor to:

access a plurality of known domain names;

determine a character embedding based on the plurality of known domain names, the character embedding mapping each character of a known domain name to a respective vector;

input the character embedding to a deep learning layer of a neural network;

access a target domain name to be classified; and

provide an output of the deep learning layer to a classifier layer that classifies the target domain name based on the output.

18 . The non-transitory machine-readable storage medium of claim 17 , wherein to determine the character embedding, the machine-readable instructions further cause the processor to:

determine the character embedding in a reverse direction.

19 . The non-transitory machine-readable storage medium of claim 17 , wherein to determine the character embedding, the machine-readable instructions further cause the processor to:

determine the character embedding in a forward direction.

20 . The non-transitory machine-readable storage medium of claim 17 , wherein the classifier layer comprises a softmax layer, and wherein the machine-readable instructions further cause the processor to:

classify, based on an output of the softmax layer, the target domain name into one or more of at least: a malicious domain name class, a non-algorithmically-generated benign domain name class, or an algorithmically-generated benign domain name class.

Assignments (6)
RELEASE OF SECURITY INTEREST REEL/FRAME 052294/0522 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC; MICRO FOCUS SOFTWARE INC. (F/K/A NOVELL, INC.); NETIQ CORPORATION
Reel/Frame 062624/0449 →
RELEASE OF SECURITY INTEREST REEL/FRAME 052295/0041 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC; MICRO FOCUS SOFTWARE INC. (F/K/A NOVELL, INC.); NETIQ CORPORATION
Reel/Frame 062625/0754 →
SECURITY AGREEMENT Recorded Apr 2, 2020
From: MICRO FOCUS LLC; BORLAND SOFTWARE CORPORATION; MICRO FOCUS SOFTWARE INC.; NETIQ CORPORATION; MICRO FOCUS (US), INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 052294/0522 →
SECURITY AGREEMENT Recorded Apr 2, 2020
From: MICRO FOCUS LLC; BORLAND SOFTWARE CORPORATION; MICRO FOCUS SOFTWARE INC.; NETIQ CORPORATION; MICRO FOCUS (US), INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 052295/0041 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CORRESPONDENT NAME: MANNAVA & KANG, P.C., ADDRESS: 3201 JERMANTOWN ROAD, SUITE 525, FAIRFAX, VIRGINIA 22030 PREVIOUSLY RECORDED ON REEL 051237 FRAME 0499. ASSIGNOR(S) HEREBY CONFIRMS THE CORRESPONDENT NAME SHOULD BE: MICRO FOCUS LLC, ADDRESS: 500 WESTOVER DR. #12603, SANFORD, NORTH CAROLINA 27330. Recorded Jan 23, 2020
From: MANADHATA, PRATYUSA K.; ARLITT, MARTIN
To: MICRO FOCUS LLC
Reel/Frame 051737/0348 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2019
From: MANADHATA, PRATYUSA K.; ARLITT, MARTIN
To: MICRO FOCUS LLC
Reel/Frame 051237/0499 →