IP Library › Granted Patent US 12,242,939
Granted Patent B2
US 12,242,939 · App. 18/686,563 · Granted Mar 4, 2025

Method, system, and computer program product for synthetic oversampling for boosting supervised anomaly detection

Inventors: Kwei-Herng Lai (Houston, TX); Lan Wang (Sunnyvale, CA); Huiyuan Chen (San Jose, CA); Mangesh Bendre (Sunnyvale, CA); Mahashweta Das (Campbell, CA); Hao Yang (San Jose, CA)
Assignee: Visa International Service Association
G06N20/00G06F18/24147
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,939
App. No.
18/686,563
Granted
Mar 4, 2025
Kind
B2
Abstract

Methods, systems, and computer program products may formulate an iterative data mix up problem into a Markov decision process (MDP) with a tailored reward signal to guide a learning process. To solve the MDP, a deep deterministic actor-critic framework may be modified to adapt a discrete-continuous decision space for training a data augmentation policy.

Claims (649)

1. A method, comprising:

obtaining, with at least one processor, a training dataset X train including a plurality of source samples including a plurality of labeled normal samples and a plurality of labeled anomaly samples;

executing, with the at least one processor, a training episode by:

(i) initializing a timestamp t;

(ii) receiving, from an actor network π of an actor critic framework including the actor network π and a critic network Q, an action vector a t for the timestamp t, wherein the actor network π is configured to generate the action vector a t based on a state s t , wherein the state s t is determined based on a current pair of source samples of the plurality of source samples, and wherein the action vector a t includes a size of a nearest neighborhood k, a composition ratio α, a number of oversampling n, and a termination probability ∈;

(iii) combining the current pair of source samples according to the composition ratio α and the number of oversampling n to generate a labeled synthetic sample x syn associated with a label y syn ;

(iv) training, using the labeled synthetic sample x syn and the label y syn , a machine learning classifier ϕ;

(v) obtaining, based on the size of a nearest neighborhood k, source samples in the k-nearest neighborhood of the labeled synthetic sample x syn ;

(vi) generating, with the machine learning classifier ϕ, for the source samples in the k-nearest neighborhood of the labeled synthetic sample x syn and a subset of the plurality of source samples of the training dataset X train in a validation dataset X val , a plurality of classifier outputs;

(vii) selecting, from the source samples in the k-nearest neighborhood of the labeled synthetic sample x syn , a next pair of source samples;

(viii) storing, in a memory buffer, the state s t , the action vector a t , a next state s t+1 , and a reward r t , wherein the next state s t+1 is determined based on the next pair of source samples, and wherein the reward r t is determined based on the plurality of classifier outputs;

(ix) determining whether the termination probability ∈ satisfies a termination threshold;

(x) in response to determining that the termination probability ∈ fails to satisfy the termination threshold, incrementing the timestamp t, for a number of training steps S:

training the critic network Q according to a critic loss function that depends on the state s t , the action vector a t , and the reward r t ; and

training the actor network π according to an actor loss function that depends on an output of the critic network, and

after training the actor network π and the critic network Q for the number of training steps S, returning to step (ii) with the next pair of source samples as the current pair of source samples;

(xi) in response to determining that the termination probability ∈ satisfies the termination threshold, determining whether the number of training episodes executed satisfies a threshold number of training episodes;

(xii) in response to determining that the number of training episodes executed fails to satisfy the threshold number of training episodes, return to step (i) to execute a next training episode; and

(xiii) in response to determining that the number of training episodes executed satisfies the threshold number of training episodes, provide the machine learning classifier ϕ, wherein the plurality of source samples is associated with a plurality of transactions in a transaction processing network, wherein the plurality of labeled normal samples is associated with a plurality of non-fraudulent transactions of the plurality of transactions, and wherein the plurality of labeled anomaly samples is associated with a plurality of fraudulent transactions of the plurality of transactions;

receiving, with the at least one processor, transaction data associated with a transaction currently being processed in the transaction processing network:

processing, with the at least one processor, using the trained machine learning classifier ϕ, the transaction data to classify the transaction as a fraudulent or non-fraudulent transaction; and

in response to classifying the transaction as a fraudulent transaction, denying, with the at least one processor, authorization of the transaction in the transaction processing network.

2. The method of claim 1 , wherein the current pair of source samples are combined according to the composition ratio α to generate the labeled synthetic sample x syn according to the following Equations:

x syn =α*x 0 +(1−α)* x 1

y

syn

=

{

y

0

,

α

≥

0.5

y

1

,

otherwise

where x 0 is a first sample of the current pair of samples, x i is a second sample of the current pair of samples, y syn is a hard label for the labeled synthetic sample x syn , y 0 is a first hard label value, and y 1 is a second hard label value.

3. The method of claim 1 , wherein the reward r t is determined according to the following Equations:

Δ

⁢

ℳ

⁡

(

ϕ

t

)

=

ℳ

⁡

(

ϕ

t

(

𝒳

val

)

,

y

val

)

-

∑

i

=

t

-

m

t

-

1

⁢

ℳ

⁡

(

ϕ

i

(

𝒳

val

)

,

y

val

)

m

-

1

𝒞

⁡

(

ϕ

t

⁢

❘

"\[LeftBracketingBar]"

s

t

,

a

t

)

=

1

k

⁢

∑

i

=

0

k

P

⁡

(

y

i

=

0

⁢

❘

"\[LeftBracketingBar]"

x

i

,

ϕ

t

)

⁢

P

⁡

(

y

i

=

1

⁢

❘

"\[LeftBracketingBar]"

x

i

,

ϕ

t

)

where M is an evaluation metric, ΔM(ϕ t ) measures a performance improvement of the trained classifier ϕ t , X val is the validation data set, y val is a label set for the training data set, where

∑

i

=

t

-

m

t

-

1

⁢

ℳ

⁡

(

ϕ

i

(

𝒳

val

)

,

y

val

)

m

-

1

is a baseline for the timestamp t, m is a hyperparameter to define a buffer size for forming the baseline, C(οt|st, at) evaluates a model confidence of the trained classifier ϕ t , P is a model exploration function, k is the size of the nearest neighborhood specified by the action vector a t , x i is a k-nearest neighborhood of the labeled synthetic sample x syn in timestamp t, and y 1 is a label for x 1 .

4. The method of claim 1 , wherein the actor loss function is defined according to the following Equation:

L

π

(

θ

1

)

=

-

1

N

⁢

∑

i

=

1

N

Q

⁡

(

s

i

,

π

⁡

(

s

i

)

⁢

❘

"\[LeftBracketingBar]"

θ

2

)

where N is a number of transitions, π(s i |θ 2 ) is a projected action for a state s i , and Q(s i , π(s i )|θ 2 ) is an output of the critic network for the projected action π(s i |θ 2 ) and state s i , and

wherein the critic loss function is defined according to the following Equation:

L Q (θ 2 )=[ Q ( s t ,a t )− b t ] 2

where b t =R(s t , a t )+γQ(s t+1 , π(s t+1 |θ 1 )|θ 2 ), π(s t+1 |θ 1 ) is an action specified by the actor network, and γ is a decade factor.

5. The method of claim 1 , further comprising:

before executing the training episode:

training, with the at least one processor, using the training dataset X train the machine learning classifier ϕ, and

pre-computing, with the at least one processor, each k-nearest neighborhood for each source sample of the plurality of source samples in the training dataset X train .

6. A system, comprising:

at least one processor programmed and/or configured to:

obtain a training dataset X train including a plurality of source samples including a plurality of labeled normal samples and a plurality of labeled anomaly samples;

execute a training episode by:

(i) initializing a timestamp t;

(ii) receiving, from an actor network π of an actor critic framework including the actor network π and a critic network Q, an action vector a t for the timestamp t, wherein the actor network π is configured to generate the action vector a t based on a state s t , wherein the state s t is determined based on a current pair of source samples of the plurality of source samples, and wherein the action vector a t includes a size of a nearest neighborhood k, a composition ratio α, a number of oversampling n, and a termination probability ∈;

(iii) combining the current pair of source samples according to the composition ratio α and the number of oversampling n to generate a labeled synthetic sample x syn associated with a label y syn ;

(iv) training, using the labeled synthetic sample x syn and the label y syn , a machine learning classifier ϕ;

(v) obtaining, based on the size of a nearest neighborhood k, source samples in the k-nearest neighborhood of the labeled synthetic sample x syn ;

(vi) generating, with the machine learning classifier ϕ, for the source samples in the k-nearest neighborhood of the labeled synthetic sample x syn and a subset of the plurality of source samples of the training dataset X train in a validation dataset X val , a plurality of classifier outputs;

(vii) selecting, from the source samples in the k-nearest neighborhood of the labeled synthetic sample x syn , a next pair of source samples;

(viii) storing, in a memory buffer, the state s t , the action vector at, a next state s t +1, and a reward r t , wherein the next state s t+1 is determined based on the next pair of source samples, and wherein the reward r t is determined based on the plurality of classifier outputs;

(ix) determining whether the termination probability ∈ satisfies a termination threshold;

(x) in response to determining that the termination probability ∈ fails to satisfy the termination threshold, incrementing the timestamp t, for a number of training steps S:

training the critic network Q according to a critic loss function that depends on the state s t , the action vector a t , and the reward r t ; and

training the actor network π according to an actor loss function that depends on an output of the critic network, and

after training the actor network π and the critic network Q for the number of training steps S, returning to step (ii) with the next pair of source samples as the current pair of source samples;

(xi) in response to determining that the termination probability E satisfies the termination threshold, determining whether the number of training episodes executed satisfies a threshold number of training episodes;

(xii) in response to determining that the number of training episodes executed fails to satisfy the threshold number of training episodes, return to step (i) to execute a next training episode; and

(xiii) in response to determining that the number of training episodes executed satisfies the threshold number of training episodes, provide the machine learning classifier ϕ, wherein the plurality of source samples is associated with a plurality of transactions in a transaction processing network, wherein the plurality of labeled normal samples is associated with a plurality of non-fraudulent transactions of the plurality of transactions, and wherein the plurality of labeled anomaly samples is associated with a plurality of fraudulent transactions of the plurality of transactions;

receive transaction data associated with a transaction currently being processed in the transaction processing network;

process, using the trained machine learning classifier ϕ, the transaction data to classify the transaction as a fraudulent or non-fraudulent transaction; and

in response to classifying the transaction as a fraudulent transaction, deny authorization of the transaction in the transaction processing network.

7. The system of claim 6 , wherein the current pair of source samples are combined according to the composition ratio α to generate the labeled synthetic sample x syn according to the following Equations:

x syn =α*x 0 +(1−α)* x 1

y

syn

=

{

y

0

,

α

≥

0.5

y

1

,

otherwise

where x 0 is a first sample of the current pair of samples, x i is a second sample of the current pair of samples, y syn is a hard label for the labeled synthetic sample x syn , y 0 is a first hard label value, and y 1 is a second hard label value.

8. The system of claim 6 , wherein the reward re is determined according to the following Equations:

Δ

⁢

ℳ

⁡

(

ϕ

t

)

=

ℳ

⁡

(

ϕ

t

(

𝒳

val

)

,

y

val

)

-

∑

i

=

t

-

m

t

-

1

⁢

ℳ

⁡

(

ϕ

i

(

𝒳

val

)

,

y

val

)

m

-

1

𝒞

⁡

(

ϕ

t

⁢

❘

"\[LeftBracketingBar]"

s

t

,

a

t

)

=

1

k

⁢

∑

i

=

0

k

P

⁡

(

y

i

=

0

⁢

❘

"\[LeftBracketingBar]"

x

i

,

ϕ

t

)

⁢

P

⁡

(

y

i

=

1

⁢

❘

"\[LeftBracketingBar]"

x

i

,

ϕ

t

)

where M is an evaluation metric, ΔM(ϕ t ) measures a performance improvement of the trained classifier ϕ t , X val is the validation data set, y val is a label set for the training data set, where

∑

i

=

t

-

m

t

-

1

⁢

ℳ

⁡

(

ϕ

i

(

𝒳

val

)

,

y

val

)

m

-

1

is a baseline for the timestamp t, m is a hyperparameter to define a buffer size for forming the baseline, C(ϕ t |st, at) evaluates a model confidence of the trained classifier ϕ t , P is a model exploration function, k is the size of the nearest neighborhood specified by the action vector a t , x i is a k-nearest neighborhood of the labeled synthetic sample x syn in timestamp t, and y 1 is a label for x 1 .

9. The system of claim 6 , wherein the actor loss function is defined according to the following Equation:

L

π

(

θ

1

)

=

-

1

N

⁢

∑

i

=

1

N

Q

⁡

(

s

i

,

π

⁡

(

s

i

)

⁢

❘

"\[LeftBracketingBar]"

θ

2

)

where N is a number of transitions, π(s i |θ 2 ) is a projected action for a state s i , and Q(s i , π(s i )|θ 2 ) is an output of the critic network for the projected action π(s i |θ 2 ) and state s i , and

wherein the critic loss function is defined according to the following Equation:

L Q (θ 2 )=[ Q ( s t ,a t )− b t ] 2

where b t =R(s t , a t )+γQ(s t+1 , π(s t+1 |θ 1 )|θ 2 ), π(s t+1 |θ 1 ) is an action specified by the actor network, and γ is a decade factor.

10. The system of claim 6 , wherein the at least one processor is further programmed and/or configured to:

before executing the training episode:

train, using the training dataset X train , the machine learning classifier ϕ; and

pre-compute each k-nearest neighborhood for each source sample of the plurality of source samples in the training dataset X train .

11. A computer program product including a non-transitory computer readable medium including program instructions which, when executed by at least one processor, cause the at least one processor to:

obtain a training dataset X train including a plurality of source samples including a plurality of labeled normal samples and a plurality of labeled anomaly samples; and

execute a training episode by:

(i) initializing a timestamp t;

(ii) receiving, from an actor network π of an actor critic framework including the actor network π and a critic network Q, an action vector a t for the timestamp t, wherein the actor network π is configured to generate the action vector a t based on a state s t , wherein the state s t is determined based on a current pair of source samples of the plurality of source samples, and wherein the action vector a t includes a size of a nearest neighborhood k, a composition ratio α, a number of oversampling n, and a termination probability ∈;

(iii) combining the current pair of source samples according to the composition ratio α and the number of oversampling n to generate a labeled synthetic sample x syn associated with a label y syn ;

(iv) training, using the labeled synthetic sample x syn and the label y syn , a machine learning classifier ϕ;

(v) obtaining, based on the size of a nearest neighborhood k, source samples in the k-nearest neighborhood of the labeled synthetic sample x syn ;

(vi) generating, with the machine learning classifier ϕ, for the source samples in the k-nearest neighborhood of the labeled synthetic sample x syn and a subset of the plurality of source samples of the training dataset X train in a validation dataset X val , a plurality of classifier outputs;

(vii) selecting, from the source samples in the k-nearest neighborhood of the labeled synthetic sample x syn , a next pair of source samples;

(viii) storing, in a memory buffer, the state s t , the action vector a t , a next state s t +1, and a reward r t , wherein the next state s t+1 is determined based on the next pair of source samples, and wherein the reward r t is determined based on the plurality of classifier outputs;

(ix) determining whether the termination probability ∈ satisfies a termination threshold;

(x) in response to determining that the termination probability ∈ fails to satisfy the termination threshold, incrementing the timestamp t, for a number of training steps S:

training the critic network Q according to a critic loss function that depends on the state s t , the action vector a t , and the reward r t ; and

training the actor network π according to an actor loss function that depends on an output of the critic network, and

after training the actor network π and the critic network Q for the number of training steps S, returning to step (ii) with the next pair of source samples as the current pair of source samples;

(xi) in response to determining that the termination probability ∈ satisfies the termination threshold, determining whether the number of training episodes executed satisfies a threshold number of training episodes;

(xii) in response to determining that the number of training episodes executed fails to satisfy the threshold number of training episodes, return to step (i) to execute a next training episode; and

(xiii) in response to determining that the number of training episodes executed satisfies the threshold number of training episodes, provide the machine learning classifier ϕ, wherein the plurality of source samples is associated with a plurality of transactions in a transaction processing network, wherein the plurality of labeled normal samples is associated with a plurality of non-fraudulent transactions of the plurality of transactions, and wherein the plurality of labeled anomaly samples is associated with a plurality of fraudulent transactions of the plurality of transactions;

receive transaction data associated with a transaction currently being processed in the transaction processing network;

process, using the trained machine learning classifier ϕ, the transaction data to classify the transaction as a fraudulent or non-fraudulent transaction; and

in response to classifying the transaction as a fraudulent transaction, deny authorization of the transaction in the transaction processing network.

12. The computer program product of claim 11 , wherein the current pair of source samples are combined according to the composition ratio α to generate the labeled synthetic sample x syn according to the following Equations:

x syn =α*x 0 +(1−α)* x 0

y

syn

=

{

y

0

,

α

≥

0.5

y

1

,

otherwise

where x 0 is a first sample of the current pair of samples, x i is a second sample of the current pair of samples, y syn is a hard label for the labeled synthetic sample x syn , y 0 is a first hard label value, and y 1 is a second hard label value.

13. The computer program product of claim 11 , wherein the reward re is determined according to the following Equations:

Δ

⁢

ℳ

⁡

(

ϕ

t

)

=

ℳ

⁡

(

ϕ

t

(

𝒳

val

)

,

y

val

)

-

∑

i

=

t

-

m

t

-

1

⁢

ℳ

⁡

(

ϕ

i

(

𝒳

val

)

,

y

val

)

m

-

1

𝒞

⁡

(

ϕ

t

⁢

❘

"\[LeftBracketingBar]"

s

t

,

a

t

)

=

1

k

⁢

∑

i

=

0

k

P

⁡

(

y

i

=

0

⁢

❘

"\[LeftBracketingBar]"

x

i

,

ϕ

t

)

⁢

P

⁡

(

y

i

=

1

⁢

❘

"\[LeftBracketingBar]"

x

i

,

ϕ

t

)

where M is an evaluation metric, ΔM(ϕ t ) measures a performance improvement of the trained classifier ϕ t , X val is the validation data set, y val is a label set for the training data set, where

∑

i

=

t

-

m

t

-

1

⁢

ℳ

⁡

(

ϕ

i

(

𝒳

val

)

,

y

val

)

m

-

1

is a baseline for the timestamp t, m is a hyperparameter to define a buffer size for forming the baseline, C(ϕ t |st, at) evaluates a model confidence of the trained classifier ϕ t , P is a model exploration function, k is the size of the nearest neighborhood specified by the action vector a t , x i is a k-nearest neighborhood of the labeled synthetic sample x syn in timestamp t, and y 1 is a label for x 1 .

14. The computer program product of claim 11 , wherein the actor loss function is defined according to the following Equation:

L

π

(

θ

1

)

=

-

1

N

⁢

∑

i

=

1

N

Q

⁡

(

s

i

,

π

⁡

(

s

i

)

⁢

❘

"\[LeftBracketingBar]"

θ

2

)

where N is a number of transitions, π(s i |θ 2 ) is a projected action for a state s i , and Q(s i , π(s i )|θ 2 ) is an output of the critic network for the projected action π(s i |θ 2 ) and state s i , and

wherein the critic loss function is defined according to the following Equation:

L Q (θ 2 )=[ Q ( s t ,a t )− b t ] 2

where b t =R(s t , a t )+γQ(s t+1 , π(s t+1 |θ 1 )|θ 2 ), π(s t+1 |θ 1 ) is an action specified by the actor network, and γ is a decade factor.

15. The computer program product of claim 11 , wherein the program instructions, when executed by at least one processor, further cause the at least one processor to:

before executing the training episode:

train, using the training dataset X train , the machine learning classifier ϕ; and

pre-compute, each k-nearest neighborhood for each source sample of the plurality of source samples in the training dataset X train .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2024
From: LAI, KWEI-HERNG; WANG, LAN; CHEN, HUIYUAN; BENDRE, MANGESH; DAS, MAHASHWETA; YANG, HAO
To: VISA INTERNATIONAL SERVICE ASSOCIATION
Reel/Frame 066565/0755 →
Continuity (2)
Provisional Application 63397719 · Aug 12, 2022
Related Publication 20240281718A1 · Aug 22, 2024
References Cited (44)
US 20160110584A1 · Remiszewski et al. · 2016 [cited by applicant]
US 20190227536A1 · Cella et al. · 2019 [cited by applicant]
US 20190259041A1 · Jackson · 2019 [cited by applicant]
US 20200314127A1 · Wilson et al. · 2020 [cited by applicant]
US 20220391299A1 · Bharadwaj · 2022 [cited by examiner]
US 20230118240A1 · Wong · 2023 [cited by examiner]
Ma, “AESMOTE: Adversarial Reinforcement Learning With SMOTE for Anomaly Detection”, IEEE Transactions on Network Science and Engineering, vol. 8, No. 2, Apr.-Jun. 2021. (Year: 2021). [cited by examiner]
Harjai, “Detecting Fraudulent Insurance Claims Using Random Forests and Synthetic Minority Oversampling Technique”, 2019 4th International Conference on Information Systems and Computer Networks (ISCON) GLA University, … [cited by examiner]
Al-Hashedi et al., “Financial fraud detection applying data mining techniques: A comprehensive review from 2009 to 2019.” Computer Science Review 40 (2021): 100402. [cited by applicant]
Bunkhumpornpat et al., “DBSMOTE: density-based synthetic minority over-sampling technique.” Applied Intelligence 36 (2012): 664-684. [cited by applicant]
Cao et al., “Learning to transfer examples for partial domain adaptation.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019. [cited by applicant]
Chapelle et al., “Semi-supervised learning (chapelle, o et al., eds.; 2006)[book reviews].” IEEE Transactions on Neural Networks 20.3 (2009): 542-542. [cited by applicant]
Chawla et al., “SMOTE: synthetic minority over-sampling technique.” Journal of artificial intelligence research 16 (2002): 321-357. [cited by applicant]
Feng et al., “A survey of data augmentation approaches for NLP.” arXiv preprint arXiv:2105.03075 (2021). [cited by applicant]
Haarnoja et al., “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.” International conference on machine learning. PMLR, 2018. [cited by applicant]
Han et al., “Borderline-SMOTE: a new over-sampling method in imbalanced data sets learning.” International conference on intelligent computing. Berlin, Heidelberg: Springer Berlin Heidelberg, 2005. [cited by applicant]
He et al., “ADASYN: Adaptive synthetic sampling approach for imbalanced learning.” 2008 IEEE international joint conference on neural networks (IEEE world congress on computational intelligence). Ieee, 2008. [cited by applicant]
Hsu et al., “Multiple time-series convolutional neural network for fault detection and diagnosis and empirical study in semiconductor manufacturing.” Journal of Intelligent Manufacturing 32 (2021): 823-836. [cited by applicant]
Ienco et al., “A semisupervised approach to the detection and characterization of outliers in categorical data.” IEEE transactions on neural networks and learning systems 28.5 (2016): 1017-1029. [cited by applicant]
Kabra et al., “Mixboost: Synthetic oversampling with boosted mixup for handling extreme imbalance.” arXiv preprint arXiv:2009.01571 (2020). [cited by applicant]
Kim et al., “AI-IDS: Application of deep learning to real-time Web intrusion detection.” IEEE Access 8 (2020): 70245-70261. [cited by applicant]
Kovacs, “Smote-variants: A python implementation of 85 minority oversampling techniques.” Neurocomputing 366 (2019): 352-354. [cited by applicant]
Li et al., “Autobalance: Optimized loss functions for imbalanced data.” Advances in Neural Information Processing Systems 34 (2021): 3163-3177. [cited by applicant]
Lillicrap et al., “Continuous control with deep reinforcement learning.” arXiv preprint arXiv: 1509.02971 (2015). [cited by applicant]
Liu et al., “The influence of class imbalance on cost-sensitive learning: An empirical study.” sixth international conference on data mining (ICDM'06). IEEE, 2006. [cited by applicant]
Maceachern et al., “Configurable fpga-based outlier detection for time series data.” 2020 IEEE 63rd International Midwest Symposium on Circuits and Systems (MWSCAS). IEEE, 2020. [cited by applicant]
Nguyen et al., “Deep neural networks are easily fooled: High confidence predictions for unrecognizable images.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2015. [cited by applicant]
Ouali et al., “An overview of deep semi-supervised learning.” arXiv preprint arXiv:2006.05278 (2020). [cited by applicant]
Pang et al., “Deep anomaly detection with deviation networks.” Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. 2019. [cited by applicant]
Pang et al., “Deep learning for anomaly detection: A review.” ACM computing surveys (CSUR) 54.2 (2021): 1-38. [cited by applicant]
Pang et al., “Deep weakly-supervised anomaly detection.” Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2023. [cited by applicant]
Pang et al., “Toward deep supervised anomaly detection: Reinforcement learning from partially labeled anomaly data.” Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining. 2021. [cited by applicant]
Perini et al., “Quantifying the confidence of anomaly detectors in their example-wise predictions.” Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Cham: Springer International Publis… [cited by applicant]
Rayana, “Odds library [http://odds. cs. stonybrook. edu]. stony brook, ny: Stony brook university.” Department of Computer Science (2016). [cited by applicant]
Ruff et al., “Deep one-class classification.” International conference on machine learning. PMLR, 2018. [cited by applicant]
Shorten et al., “A survey on image data augmentation for deep learning.” Journal of big data 6.1 (2019): 1-48. [cited by applicant]
Sutton et al., Reinforcement learning: An introduction. MIT press, 2018. [cited by applicant]
Vanschoren et al., “OpenML: networked science in machine learning.” ACM SIGKDD Explorations Newsletter 15.2 (2014): 49-60. [cited by applicant]
Wen et al., “Time series data augmentation for deep learning: A survey.” arXiv preprint arXiv:2002.12478 (2020). [cited by applicant]
Wu et al., “Rlad: Time series anomaly detection through reinforcement learning and active learning.” arXiv preprint arXiv:2104.00543 (2021). [cited by applicant]
Zha et al., “Meta-AAD: Active anomaly detection with deep reinforcement learning.” 2020 IEEE International Conference on Data Mining (ICDM). IEEE, 2020. [cited by applicant]
Zhang et al., “mixup: Beyond empirical risk minimization.” arXiv preprint arXiv:1710.09412 (2017). [cited by applicant]
Zhao et al., “Xgbod: improving supervised outlier detection with unsupervised representation learning.” 2018 International Joint Conference on Neural Networks (IJCNN). IEEE, 2018. [cited by applicant]
Zhu, “Semi-supervised learning literature survey.” (2005). [cited by applicant]