IP Library › Granted Patent US 12,572,778
Granted Patent B2
US 12,572,778 · App. 16/889,652 · Granted Mar 10, 2026

Technique to perform neural network architecture search with federated learning

Inventors: Holger Reinhard Roth (Rockville, MD); Dong Yang (Pocatello, ID); Wenqi Li (London, GB); Andriy Myronenko (San Francisco, CA); Wentao Zhu (Bethesda, MD); Ziyue Xu (Reston, VA); Xiaosong Wang (Rockville, MD); Daguang Xu (Potomac, MD)
Assignee: NVIDIA Corporation
G06N3/045G06N3/08G06T7/10G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,778
App. No.
16/889,652
Granted
Mar 10, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to select a neural network architecture from a plurality of neural networks in a federated learning (FL) setting. In at least one embodiment, a neural network is trained by combining training results from different FL computing systems, where each of the different FL computing systems, for example, trains different portions of the neural network.

Claims (102)

1 . One or more processors, comprising:

circuitry to:

select different portions of two or more second neural networks from a plurality of neural networks based, at least in part, on training data used to train the selected different portions;

use the selected different portions to train correspondingly different portions of a first neural network from the plurality of neural networks by:

passing a data point to the first neural network;

reconstructing the data point using a generator network to generate a reconstructed data point;

comparing the reconstructed data point with the data point; and

updating one or more weights of the first neural network.

2 . The one or more processors of claim 1 , wherein the selected different portions of the two or more second neural networks are further selected based, at least in part, on information to be inferenced using the first neural network.

3 . The one or more processors of claim 1 , wherein the circuitry is further to:

receive training results comprising model parameters from each of the two or more second neural networks; and

train the first neural network using the model parameters.

4 . The one or more processors of claim 1 , wherein the first neural network is a supernetwork.

5 . The one or more processors of claim 1 , wherein the circuitry is further to use results from the comparing to indicate which operation to select from each layer of the first neural network to determine a second neural network for the data point.

6 . The one or more processors of claim 5 , wherein operations that are performed on the data point at each layer are averaged prior to being fed to a subsequent layer to perform additional operations on the data point.

7 . The one or more processors of claim 5 , wherein the circuitry is further to:

pass a second data point, using at least one computer system, to the first neural network for inferencing;

use the generator network to reconstruct the second data point to generate a reconstructed second data point;

perform a comparison of the reconstructed second data point with the second data point; and

use results from the comparison to indicate which operation to select from each layer of the first neural network to select another second neural network of the two or more second neural networks for the second data point, wherein the another second neural network is either the same second neural network or a different second neural network.

8 . The one or more processors of claim 7 , wherein the second data point is of the same type of input as the data point.

9 . The one or more processors of claim 1 , wherein the comparing is performed by using a local validation set.

10 . The one or more processors of claim 9 , wherein a loss is determined based on the local validation set.

11 . A system, comprising:

one or more computers having one or more processors to:

select different portions of two or more second neural networks from a plurality of neural networks based, at least in part, on training data used to train the selected different portions;

use the selected different portions to train correspondingly different portions of a first neural network from the plurality of neural networks by:

passing a data point to the first neural network;

reconstructing the data point using a generator network to generate a reconstructed data point;

comparing the reconstructed data point with the data point; and

updating one or more weights of the first neural network.

12 . The system of claim 11 , wherein training the correspondingly different portions of the first neural network further comprises the one or more computers having one or more processors to select a second neural network from the two or more second neural networks to train.

13 . The system of claim 11 , wherein a second neural network of the two or more second neural networks comprise an optimal path that is determined based, at least in part, on information to be inferenced using at least one of the two or more second neural networks.

14 . The system of claim 11 , wherein the one or more processors are further to:

receive model weights from multiple different computer systems; and

aggregate the model weights and use information from the aggregated model weights to train the correspondingly different portions of the first neural network.

15 . The system of claim 14 , wherein the one or more processors are further to use results from the comparing to indicate which operation to select from each layer of the first neural network to select a second neural network for the data point.

16 . The system of claim 15 , further comprising the one or more processors are to generate a weighted average from operations that are to be performed on the data point at each layer before feeding the data point to a subsequent layer to perform additional operations on the data point.

17 . The system of claim 15 , wherein the comparing is performed by using a local validation set.

18 . The system of claim 17 , wherein a loss is determined based on the local validation set.

19 . The system of claim 11 , wherein the one or more computers having one or more processors further cause the two or more second neural networks to be selected as different second neural networks for different data points of the same type.

20 . A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to:

select different portions of two or more second neural networks from a plurality of neural networks based, at least in part, on training data used to train the selected different portions;

use the selected different portions to train correspondingly different portions of a first neural network from the plurality of neural networks by:

passing a data point to the first neural network;

reconstructing the data point using a generator network to generate a reconstructed data point;

comparing the reconstructed data point with the data point; and

updating one or more weights of the first neural network.

21 . The machine-readable medium of claim 20 ,

wherein the set of instructions, which if performed by the one or more processors, further cause the one or more processors to:

receive model parameters from each of the two or more second neural networks based at least in part on each of the two or more second neural networks training a portion of the first neural network using local data for each of the two or more second neural networks; and

aggregate the model parameters and use information from the aggregated model parameters to generate the first neural network.

22 . The machine-readable medium of claim 21 , wherein the set of instructions, which if performed by the one or more processors, further cause the one or more processors to use results from the comparing to indicate which operation to select from each layer of the first neural network to select a second neural network for the data point.

23 . The machine-readable medium of claim 22 , wherein each correspondingly different portion of the first neural network comprises the selected second neural network.

24 . The machine-readable medium of claim 23 , wherein the second neural network comprises an optimal path that is determined based, at least in part, on information to be inferenced using the first neural network.

25 . The machine-readable medium of claim 24 , wherein information to be inferenced comprises local data accessible to each individual computer system of different computer systems and inaccessible to other computer systems.

26 . The machine-readable medium of claim 20 , wherein the set of instructions, which if performed by the one or more processors, further cause the one or more processors to train the first neural network to perform medical image segmentation.

27 . The machine-readable medium of claim 20 , wherein the first neural network is a convolutional neural network.

28 . One or more processors, comprising:

circuitry to:

use a first neural network to infer information, wherein correspondingly different portions of the first neural network are trained using:

selected different portions of two or more second neural networks from a plurality of neural networks based, at least in part, on training data used to train the selected different portions, by:

passing a data point to the first neural network;

reconstructing the data point using a generator network to generate a reconstructed data point;

comparing the reconstructed data point with the data point; and

updating one or more weights of the first neural network.

29 . The one or more processors of claim 28 , wherein the first neural network is a supernetwork.

30 . The one or more processors of claim 29 , wherein different portions of the supernetwork are trained at different computer systems by selecting a second neural network of the two or more second neural networks from the supernetwork to train, wherein the second neural network is selected by each of the different computer systems based, at least in part, on information, inaccessible to other computer systems, to be inferenced using the first neural network.

31 . The one or more processors of claim 30 , wherein the circuitry is further to:

receive training results comprising model parameters from each of the different computer systems, wherein each of the different computer systems trains a portion of the first neural network using data not shared with other computer systems; and

train the supernetwork using the model parameters.

32 . The one or more processors of claim 31 , wherein the circuitry is further to use results from the comparing to indicate which operation to select from each layer of the first neural network to select a second neural network for the data point.

33 . The one or more processors of claim 32 , wherein the circuitry is further to generate results from the comparing between the reconstructed data point and the data point by using cross-entropy in a reconstruction loss function.

34 . The one or more processors of claim 33 , wherein the one or more first neural networks is trained to perform image segmentation.

35 . The one or more processors of claim 28 , wherein the one or more first neural networks comprises a convolutional neural network.

36 . A system, comprising:

one or more computers having one or more processors to:

use a first neural network to infer information, wherein correspondingly different portions of the first neural network are trained using:

selected different portions of two or more second neural networks from a plurality of neural networks based, at least in part, on training data used to train the selected different portions, by:

passing a data point to the first neural network;

reconstructing the data point using a generator network to generate a reconstructed data point;

comparing the reconstructed data point with the data point; and

updating one or more weights of the first neural network.

37 . The system of claim 36 , wherein training the correspondingly different portions of the first neural network further comprise the one or more computers having one or more processors to select a path from the two or more second neural networks, wherein the path is selected based, at least in part, on information to be inferenced using the first neural network.

38 . The system of claim 37 , wherein the one or more processors are further to:

receive model weights from each of the two or more second neural networks; and

aggregate the model weights and use information from the aggregated model weights to train the first neural network.

39 . The system of claim 38 , wherein the one or more processors are further to use results from the comparing to indicate which operation to select from each layer of the first neural network to construct the path for the data point.

40 . The system of claim 39 , wherein data point used by each of the different computer systems is inaccessible to other computer systems.

41 . The system of claim 36 , wherein the correspondingly different portions of the first neural network is trained to adapt to domains for different computer systems.

42 . A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to:

use a first neural network to infer information, wherein correspondingly different portions of the first neural network are trained using:

selected different portions of two or more second neural networks from a plurality of neural networks based, at least in part, on training data used to train the selected different portions, by:

passing a data point to the first neural network;

reconstructing the data point using a generator network to generate a reconstructed data point;

comparing the reconstructed data point with the data point; and

updating one or more weights of the first neural network.

43 . The machine-readable medium of claim 42 , wherein training the correspondingly different portions of the first neural network further comprise selecting a second neural network of the two or more second neural networks from the first neural network to train, wherein the second neural network is selected based, at least in part, on information to be inferenced using the first neural network.

44 . The machine-readable medium of claim 43 , wherein the set of instructions, which if performed by the one or more processors, further cause the one or more processors to:

receive model parameters from each of the two or more second neural networks; and

aggregate the model parameters and use information from the aggregated model parameters to generate the first neural network.

45 . The machine-readable medium of claim 44 , wherein the set of instructions, which if performed by the one or more processors, further cause the one or more processors to use results from the comparing to indicate which operation to select from each layer of the first neural network to select a second neural network for the data point.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2021
From: ROTH, HOLGER REINHARD; YANG, DONG; LI, WENQI; MYRONENKO, ANDRIY; ZHU, WENTAO; XU, ZIYUE; WANG, XIAOSONG; XU, DAGUANG
To: NVIDIA CORPORATION
Reel/Frame 054934/0117 →
Continuity (1)
Related Publication 20210374502A1 · Dec 2, 2021
References Cited (57)
US 11373115B2 · Kopp · 2022 [cited by examiner]
US 20170344829A1 · Lan et al. · 2017 [cited by applicant]
US 20190138934A1 · Prakash et al. · 2019 [cited by applicant]
US 20190311298A1 · Kopp et al. · 2019 [cited by applicant]
US 20190385043A1 · Choudhary · 2019 [cited by examiner]
US 20200104706A1 · Sandler et al. · 2020 [cited by applicant]
US 20200272899A1 · Dunne · 2020 [cited by examiner]
US 20210117780A1 · Malik · 2021 [cited by examiner]
CN 108876702A · 2018 [cited by applicant]
CN 110569779A · 2019 [cited by applicant]
CN 110942154A · 2020 [cited by applicant]
EP 3396600A1 · 2018 [cited by applicant]
Mengwei Xu, Neural Architecture Search over Decentralized Data, Feb. 15, 2020, arXiv. [cited by examiner]
Zhuotun Zhu, V-NAS: Neural Architecture Search for Volumetric Medical Image Segmentation, Aug. 4, 2019, arXiv. [cited by examiner]
Cai et al., “ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware,”Feb. 23, 2019, 13 pages. [cited by applicant]
Elsken et al., “Neural Architecture Search: A survey,” Sep. 5, 2018, 17 pages. [cited by applicant]
Ganin et al. “Unsupervised Domain Adaptation by Backpropagation,” Jun. 1, 2015, 10 pages. [cited by applicant]
Ganin et al., “Domain-Adversarial Training of Neural Networks,” The Journal of Machine Learning Research, 17(1): 2016, 35 pages. [cited by applicant]
Ginsburg et al., “Stochastic Gradient Methods with Layer-Wise Adaptive Moments for Training of Deep Networks,” Sep. 18, 2019, 13 pages. [cited by applicant]
Guo et al., “Single Path One-Shot Neural Architecture Search with Uniform Sampling,” Apr. 6, 2019, 14 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2021/034367, mailed Sep. 9, 2021, filed May 26, 2021, 22 pages. [cited by applicant]
Isensee et al., “nnU-Net: Breaking the Spell on Successful Medical Image Segmentation,” Apr. 17, 2019, 8 pages. [cited by applicant]
Isensee et al., “No New-Net,” International Conference on Medical Image Computing and Computer Assisted Intervention, Sep. 27, 2018, 10 pages. [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks,” IEEE Conference on Computer Vision and Pattern Recognition, 2017, 10 pages. [cited by applicant]
Kamnitsas et al., “DeepMedic for Brain Tumor Segmentation,” International Workshop on Brainlesion, Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries, Springer, 2016, 12 pages. [cited by applicant]
Kamnitsas et al., “Unsupervised Domain Adaptation in Brain Lesion Segmentation with Adversarial Networks,” International Conference on Information Processing in Medical Imaging, Decemebr 28, 2016, 13 pages. [cited by applicant]
Li et al., “On the Compactness, Efficiency, and Representation of 3D Convolutional Networks: Brain Parcellation as a Pretext Task,” International Conference on Information Processing in Medical Imaging, Jul. 6, 2017, 13… [cited by applicant]
Li et al., “Privacy-preserving Federated Brain Tumour Segmentation, ” Oct. 2, 2019, 8 pages. [cited by applicant]
Liang et al., “Think Locally, Act Globally: Federated Learning with Local and Global Representations,” Jan. 6, 2020, 21 pages. [cited by applicant]
Litjens et al., “Computer-aided Detection of Prostate Cancer in MRI,” IEEE Transactions on Medical Imaging, 33(5): 2014, 10 pages. [cited by applicant]
Litjens et al., Evaluation of Prostate Segmentation Algorithms for MRI: The PROMISE 12 Challenge, Medical Image Analysis, 18(2): 2014, 15 pages. [cited by applicant]
Liu et al., “Darts: Differentiable Architecture Search,” Jun. 24, 2018, 13 pages. [cited by applicant]
Mayer et al., “Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools,” Mar. 27, 2019, 35 pages. [cited by applicant]
McMahan et al., “Communication-Efficient Learning of Deep Networks from Decentralized Data,” Feb. 17, 2016, 11 pages. [cited by applicant]
Milletari et al., “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” 2016 Fourth International Conference on 3D Vision (3DV), Oct. 25, 2016, 11 pages. [cited by applicant]
Ronneberger et al., “U-net: Convolutional networks for biomedical image segmentation,” International Conference on Medical Image Computing and Computer-Assisted Intervention, Oct. 5, 2015, 8 pages. [cited by applicant]
Shaw et al., “SqueezeNAS: Fast Neural Architecture Search for Faster Semantic Segmentation,” In Proceedings of the IEEE International Conference on Computer Vision Workshops, 2019, 11 pages. [cited by applicant]
Sheller et al., “Multi-Institutional Deep Learning Modeling without Sharing Patient Data: A Feasibility Study on Brain Tumor Segmentation,” MICCAI Brainlesion Workshop, Oct. 22, 2018, 14 pages. [cited by applicant]
Shin et al., Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning, IEEE Transactions on Medical Imaging, 35(5): May 2016, 14 pages. [cited by applicant]
Simpson et al, “A Large Annotated Medical Image Dataset for the Development and Evaluation of Segmentation Algorithms,” Feb. 25, 2019, 15 pages. [cited by applicant]
Society of Automotive Engineers on-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers on-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Tan et al., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” International Conference on Machine Learning, 2019, 10 pages. [cited by applicant]
Wu et al., “Fbnet: Hardware-Aware Efficient Convnet Design via Differentiable Neural Architecture Search,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 9 pages. [cited by applicant]
Xu et al., “Neural Architecture Search over Decentralized Data,” Feb. 15, 2020, 10 pages. [cited by applicant]
Yang et al., “Technique to Perform Neural Network Architecture Search,” U.S. Appl. No. 16/803,925, filed Feb. 27, 2020. [cited by applicant]
Yu et al., “C2FNAS: Coarse-to-Fine Neural Architecture Search for 3D Medical Image Segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Zhang et al., Generalizing Deep Learning for Medical Image Segmentation to Unseen Domains via Deep Stacked Transformation, IEEE Transactions on Medical Imaging, 2020, 10 pages. [cited by applicant]
Zhang et al., “Task Driven Generative Modeling for Unsupervised Domain Adaptation: Application to X-ray Image Segmentation,” International Conference on Medical Image Computing and Computer-Assisted Intervention, Jun. 1… [cited by applicant]
Zhu et al., “Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks,” In Proceedings of the IEEE International Conference on Computer Vision, 2017, 10 pages. [cited by applicant]
Zhu et al., “V-NAS: Neural Architecture Search for Volumetric Medical Image Segmentation,” 2019 International Conference on 3D Vision (3DV), Sep. 16, 2019, 9 pages. [cited by applicant]
Çiçek et al., “3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation,” International Conference on Medical Image Computing and Computer-Assisted Intervention, 2016, 8 pages. [cited by applicant]
Office Action for Chinese Application No. 202180045946.8, mailed Jun. 19, 2025, 25 pages. [cited by applicant]
Roth et al., “Federated Whole Prostate Segmentation in MRI with Personalized Neural Architectures,” 2021, 11 pages. [cited by applicant]
Office Action for Chinese Application No. 202180045946.8, mailed Jan. 7, 2026, 28 pages. [cited by applicant]
He et al., “FedNAS: Federated Deep Learning via Neural Architecture Search,” Apr. 18, 2020, 6 Pages. [cited by applicant]