IP Library Granted Patent US 12,340,308
Granted Patent B2
US 12,340,308 · App. 17/695,325 · Granted Jun 24, 2025

Edge-side federated learning for anomaly detection

Inventors: Dongjin Song (Princeton, NJ); Yuncong Chen (Plainsboro, NJ); Cristian Lumezanu (Princeton Junction, NJ); Takehiko Mizoguchi (Princeton, NJ); Haifeng Chen (West Windsor, NJ); Wei Zhu (Rochester, NY)
Assignee: NEC Corporation
G06N3/08G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,308
App. No.
17/695,325
Granted
Jun 24, 2025
Kind
B2
Abstract

Methods and systems for training a neural network include collecting model exemplar information from edge devices, each model exemplar having been trained using information local to the respective edge devices. The collected model exemplar information is aggregated together using federated averaging. Global model exemplars are trained using federated constrained clustering. The trained global exemplars are transmitted to respective edge devices.

Claims (164)

1. A method for training a neural network, comprising:

training an edge model exemplar at an edge device, using an initialized global model exemplar, based on information collected at the edge device, including optimizing an objective function:

min

θ

,

C

-

1

n

i

=

1

n

K

L

(

p

i

q

i

)

-

α

T

log

(

1

n

i

=

1

n

q

i

)

+

1

/

n

i

=

1

n

M

(

X

i

)

where n is a number of multivariate time series segments X i , θ is a set of parameters for the neural network to be learned, C is a set of edge model exemplars, KL(·) is a Kullback-Leibler divergence, p i is a target cluster membership vector for an i th locally gathered information, q i is a cluster membership vector for the i th locally gathered information, α is a prior distribution over the edge model exemplars, and M(X i ) is a term that preserves local similarity of an original feature space;

transmitting the edge model exemplar to a server without transmitting the information collected at the edge device;

receiving an updated global model exemplar that is based on the edge model exemplar and at least one other model exemplar from another edge device; and

retraining the edge model exemplar using the updated global model exemplar.

2. The method of claim 1 , wherein the updated global model exemplar is a federated average of the edge model exemplar and the at least one other model exemplar.

3. The method of claim 2 , wherein the federated average is an element-wise average of exemplars.

4. The method of claim 1 , further comprising repeating the transmitting, receiving, and retraining based on additional information collected at the edge device.

5. The method of claim 1 , wherein the edge model exemplar includes the neural network, which includes a bidirectional long-short term memory layer.

6. The method of claim 1 , further comprising determining an anomaly score using the retrained edge model exemplar based on the information gathered at the edge device.

7. The method of claim 6 , wherein determining the anomaly score is based on a similarity between new information and existing exemplars.

8. The method of claim 1 , wherein the retrained edge model exemplar recognizes operating conditions from cyber-physical systems associated with a plurality of edge devices.

9. A system for training a neural network, comprising: a hardware processor; and a memory that stores a computer program, which, when executed by the hardware processor, causes the hardware processor to:

train an edge model exemplar at an edge device, using an initialized global model exemplar, based on information collected at the edge device, including optimization of an objective function:

min

θ

,

C

-

1

n

i

=

1

n

K

L

(

p

i

q

i

)

-

α

T

log

(

1

n

i

=

1

n

q

i

)

+

1

/

n

i

=

1

n

M

(

X

i

)

where n is a number of multivariate time series segments X i , θ is a set of parameters for the neural network to be learned, C is a set of edge model exemplars, KL(·) is a Kullback-Leibler divergence, p i is a target cluster membership vector for an i th locally gathered information, q i is a cluster membership vector for the i th locally gathered information, α is a prior distribution over the edge model exemplars, and M(X i ) is a term that preserves local similarity of an original feature space;

transmit the edge model exemplar to a server without transmission of the information collected at the edge device;

receive an updated global model exemplar that is based on the edge model exemplar and at least one other model exemplar from another edge device; and

retrain the edge model exemplar using the updated global model exemplar.

10. The system of claim 9 , wherein the updated global model exemplar is a federated average of the edge model exemplar and the at least one other model exemplar.

11. The system of claim 10 , wherein the federated average is an element-wise average of exemplars.

12. The system of claim 9 , wherein the computer program further causes the hardware processor to repeat the transmission, receipt, and retraining based on additional information collected at the edge device.

13. The system of claim 9 , wherein the edge model exemplar is a includes the neural network, which includes a bidirectional long-short term memory layer.

14. The system of claim 9 , wherein the computer program causes the hardware processor to determine an anomaly score using the retrained edge model exemplar based on the information gathered at the edge device.

15. The system of claim 14 , wherein the anomaly score is based on a similarity between new information and existing exemplars.

16. The system of claim 9 , wherein the retrained edge model exemplar recognizes operating conditions from cyber-physical systems associated with a plurality of edge devices.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 071095/0825 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2022
From: SONG, DONGJIN; CHEN, YUNCONG; LUMEZANU, CRISTIAN; MIZOGUCHI, TAKEHIKO; CHEN, HAIFENG; ZHU, WEI
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 059270/0897 →
Continuity (6)
Continuation 17395118 · Aug 5, 2021
Provisional Application 63291560 · Dec 20, 2021
Provisional Application 63075450 · Sep 8, 2020
Provisional Application 63070437 · Aug 26, 2020
Provisional Application 63062031 · Aug 6, 2020
Related Publication 20220215256A1 · Jul 7, 2022
References Cited (23)
US 11593485B1 · Briliauskas · 2023 [cited by examiner]
US 11763197B2 · McMahan · 2023 [cited by examiner]
US 20120284791A1 · Miller · 2012 [cited by examiner]
US 20180032915A1 · Nagaraju · 2018 [cited by examiner]
US 20190050515A1 · Su · 2019 [cited by examiner]
US 20200027033A1 · Garg · 2020 [cited by examiner]
US 20200028862A1 · Lin · 2020 [cited by examiner]
US 20200364579A1 · Misu · 2020 [cited by examiner]
Clark, “A Comparison of Rule and Exemplar-Based Learning Systems”, in P. B. Brazdil et al. (eds.), Machine Learning, Meta-Reasoning and Logics, © Kluwer Academic Publishers 1990, pp. 159-186. (Year: 1990). [cited by examiner]
Wang et al., “Adaptive Federated Learning in Resource Constrained Edge Computing Systems”, arXiv ID: 1804.05271v3, Feb. 17, 2019, pp. 1-20. (Year: 2019). [cited by examiner]
Corinzia et al., “Variational FederatedMulti-Task Learning”, arXiv ID: 1906.06268, Jun. 14, 2019, pp. 1-10. (Year: 2019). [cited by examiner]
Nguyen et al., “DÏoT: A Federated Self-learning Anomaly Detection System for IoT”, 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS), Jul. 2019, pp. 756-767. (Year: 2019). [cited by examiner]
Kwak, et al., “DeepHealth: Deep Learning for Health Informatics reviews, challenges, and opportunities on medical imaging, electronic health records, genomics, sensing, and online communication health”, arXiv ID: 1909.0… [cited by examiner]
Lim et al., “Federated Learning in Mobile Edge Networks: A Comprehensive Survey”, arXiv ID: 1909.11875, Sep. 26, 2019. (Year: 2019). [cited by examiner]
Savazzi et al., “Federated Learning with Cooperating Devices: A Consensus Approach for Massive IoT Networks”, arXiv ID: 1912.13163, Dec. 27, 2019, pp. 1-16. (Year: 2019). [cited by examiner]
Wang et al., “Federated Learning with Matched Averaging”, arXiv ID: 2002.06440; pub. date on Feb. 15, 2020, pp. 1-16. (Year: 2020). [cited by examiner]
Liu et al., “Privacy-preserving Traffic Flow Prediction: A Federated Learning Approach”, arXiv ID: 2003.08725, Mar. 19, 2020, pp. 1-12. (Year: 2020). [cited by examiner]
Duan et al., “Self-Balancing Federated Learning With Global Imbalanced Data in Mobile Systems”, IEEE Transactions on Parallel and Distributed Systems, vol. 32, No. 1, date of publication Jul. 15, 2020, pp. 59-71. (Year:… [cited by examiner]
Liu et al., “Deep Anomaly Detection for Time-series Data in Industrial IoT: A Communication-Efficient On-device Federated Learning Approach”, arXiv ID: 2007.09712; pub. date on Jul. 19, 2020, pp. 1-11. (Year: 2020). [cited by examiner]
Sun et al., “Intrusion Detection with Segmented Federated Learning for Large-Scale Multiple LANs”, 2020 International Joint Conference on Neural Networks (IJCNN), Jul. 19-24, 2020, pp. 1-8. (Year: 2020). [cited by examiner]
Mcmahan, H. Brendan, et al. “Communication-Efficient Learning of Deep Networks from Decentralized Data”, InArtificial intelligence and statistics, PMLR. Apr. 10, 2017, pp. 1-10. [cited by applicant]
Pang, Guansong, et al. “Deep Learning for Anomaly Detection: A Review”, ACM Computing Surveys (CSUR). Mar. 5, 2021, pp. 1-38. [cited by applicant]
Wang, Hongyi, et al. “Federated Learning with Matched Averaging”, arXiv preprint arXiv:2002.06440. Feb. 15, 2020, pp. 1-16. [cited by applicant]