IP Library › Granted Patent US 12,412,076
Granted Patent B1
US 12,412,076 · App. 18/961,535 · Granted Sep 9, 2025

Traffic anomaly detection method and system based on improved BERT integrating contrastive learning

Inventors: Feiran Huang (Guangzhou, CN); Jinming Zhong (Guangzhou, CN); Zhibo Zhou (Guangzhou, CN); Jian Weng (Guangzhou, CN)
Assignee: JINAN UNIVERSITY
G06N3/045G06N3/084G06N3/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,076
App. No.
18/961,535
Granted
Sep 9, 2025
Kind
B1
Abstract

A traffic anomaly detection method and system based on improved BERT integrating contrastive learning is provided, the method includes: obtaining traffic data and preprocessing the data; building an improved BERT model including an embedding layer and 12 Transformer encoder networks; performing weight sharing operation among a first 6 Transformer encoder networks and a last 6 Transformer encoder networks, respectively; building a classification network; building a total loss function based on a cross-entropy loss and a contrastive loss; performing unsupervised pre-training on the improved Bidirectional Encoder Representations from Transformers, BERT model; fine-tune training the improved BERT model; updating model parameters through backpropagation to obtain a trained improved BERT model; feed test traffic data into the trained improved BERT model; and obtain a traffic detection outcome. The present disclosure enhances the generalization ability of the model while maintaining stability and accuracy.

Claims (159)

1. A traffic anomaly detection method based on improved Bidirectional Encoder Representations from Transformers, BERT integrating contrastive learning, comprising the following steps:

obtaining traffic data and preprocessing the traffic data to obtain a byte sequence;

building an improved BERT model comprising an embedding layer and 12 Transformer encoder networks, wherein the byte sequence is fed into the embedding layer, and is taken to be a text segment, and a byte is taken to be a word; extracting a token embedding, a segment embedding and a position embedding as representations for the word; summing up the token embedding, the segment embedding and the position embedding to obtain a corresponding vector; performing weight sharing operation among a first 6 Transformer encoder networks and a last 6 Transformer encoder networks, respectively; and outputting, by the improved BERT model and for the byte, a vector representation comprising traffic contextual feature information;

building a classification network;

building a total loss function based on a cross-entropy loss and a contrastive loss, and it comprises:

Loss

=

λ

·

CELoss

+

(

1

-

λ

)

·

SCLoss

CELoss

=

∑

(

y

,

t

)

-

[

(

t

⁢

log

⁢

y

)

+

(

1

-

t

)

⁢

log

⁢

(

1

-

y

)

]

SCLoss

=

∑

i

≠

j

-

y

·

log

⁡

(

σ

⁢

(

S

⁢

(

x

i

,

x

j

)

/

τ

)

)

-

(

1

-

y

)

·

log

⁡

(

1

-

σ

⁢

(

S

⁢

(

x

i

,

x

j

)

/

τ

)

)

wherein the CELoss denotes the cross-entropy loss, the SCLoss denotes contrastive loss, the λ is used to control weights between the cross-entropy loss and the contrastive loss, the t denotes a real label of a network stream, the y denotes a probability result calculated by a softmax layer, the σ denotes a sigmoid function, the i and the j denote different network traffic samples from a same batch, the x i and x j each denotes a vector output by a 10 th layer encoder of the improved BERT model, the S denotes a function for measuring a similarity between the x i and x j , the y indicates whether the x i and x j belong to a same class, and τ is a hyperparameter;

applying masking to the byte sequence and feeding a masked byte sequence into the improved BERT model for unsupervised pre-training; calculating a probability distribution of a byte at a masked position using forward computation; and calculating a difference between a predicted probability distribution and a real label based on the cross-entropy loss function;

passing an output vector from a 12 th layer Transformer encoder network of the improved BERT model to the classification network; outputting, by the classification network, a probability distribution corresponding to a number of class labels; and calculating the contrastive loss using an output vector of the 10 th layer encoder network of the improved BERT model;

updating a model parameter through backpropagation to obtain a trained improved BERT model;

obtaining test traffic data; feeding the test traffic data into the trained improved BERT model; and obtaining a traffic detection outcome.

2. The method according to claim 1 , wherein the preprocessing the traffic data comprises data splitting, vocabulary expanding, data cleaning, and unifying data length,

wherein the data splitting comprises splitting an original traffic data set into data stream sets using a network session as a splitting criterion, wherein a data stream is a sequence of multiple data packets, wherein a data packet in the data stream comprises a five-tuple: {source IP, destination IP, source port, destination port, network protocol}, and is arranged according to a temporal order, and the data packet is a sequence of multiple bytes,

the vocabulary expanding comprises adding new words to a BERT vocabulary,

the data cleaning comprises removing information that does not meet a specified condition, and

the unifying data length comprises a step of unifying data stream length and a step of unifying data packet length.

3. The method according to claim 1 , wherein the classification network comprises several fully connected layers and the softmax layer,

the vector representation outputted from the improved BERT model is fed into the classification network and through the several fully connected layers to generate a numerical distribution list for traffic classification,

the softmax layer performs softmax calculation on the numerical distribution list to convert the list into a probability distribution for traffic classification.

4. The method according to claim 1 , wherein following the trained improved BERT model is obtained, the method further comprises: assessing a performance of the BERT model according to accuracy, recall, precision, and F1 score, denoted by:

Accuracy

=

TP

+

TN

TP

+

TN

+

FP

+

FN

Recall

=

TP

TP

+

FN

Precision

=

TP

TP

+

FP

F

⁢

1

⁢

_score

=

2

×

Precison

×

Recall

Precison

+

Recall

wherein the Accuracy denotes the accuracy, the Recall denotes the recall, the Precision denotes the precision, the F1_score denotes the F1 score, the TP indicates the BERT model has correctly predicted an actual anomaly traffic to be anomalous, the TN indicates the BERT model has correctly predicted an actual normal traffic to be normal, the FP indicates the BERT model has erroneously predicted an actual normal traffic to be anomalous, the FN indicates the BERT model has erroneously predicted an actual anomaly traffic to be normal.

Priority Claims (1)
CN 202410258967.3 · Mar 7, 2024 · national
References Cited (16)
US 11843624B1 · Estep · 2023 [cited by applicant]
US 20190102678A1 · Chang · 2019 [cited by examiner]
US 20220129621A1 · Guda · 2022 [cited by applicant]
US 20240273374A1 · Lee · 2024 [cited by examiner]
CN 114781392A · 2022 [cited by applicant]
CN 114861601B · 2022 [cited by applicant]
CN 116257698A · 2023 [cited by applicant]
CN 116541838A · 2023 [cited by applicant]
CN 116595407A · 2023 [cited by applicant]
CN 116910341A · 2023 [cited by applicant]
CN 117082004A · 2023 [cited by applicant]
WO 2024000944A1 · 2024 [cited by applicant]
Hang, Zijun, et al. “Flow-MAE: Leveraging Masked AutoEncoder for Accurate, Efficient and Robust Malicious Traffic Classification.” Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and D… [cited by examiner]
He, Ju, et al. “Transfg: A transformer architecture for fine-grained recognition.” Proceedings of the AAAI conference on artificial intelligence. vol. 36. No. 1. 2022. (Year: 2022). [cited by examiner]
He Shasha “Research on Encrypted Traffic Classification Based on CNN-Transformer and Mutal Contrastive Learning” CMFD, Information Technology, Sep. 15, 2023 (Sep. 15, 2023), pp. | 139-89. [cited by applicant]
Zijun Hang “Flow-MAE: Leveraging Masked AutoEncoder for Accurate, Efficient and Robust Malicious Traffic Classification” RAID '23: Proceedings of the 26th International Symposium on Research in Attacks, Oct. 16, 2023 (O… [cited by applicant]