IP Library › Granted Patent US 12,639,576
Granted Patent B2
US 12,639,576 · App. 19/320,807 · Granted May 26, 2026

Architectural augmentation of neural networks using evaluated specialty node units

Inventors: James K. Baker (Maitland, FL); Bradley J. Baker (Berwyn, PA)
Assignee: D5AI LLC
G06N3/082G06F16/9024G06F18/214G06F18/217G06F18/24G06N3/04G06N3/045G06N3/048G06N3/08G06N3/084G06N20/00G06N20/20H04L67/142G06F18/2148G06N5/046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,576
App. No.
19/320,807
Granted
May 26, 2026
Kind
B2
Abstract

A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.

Claims (52)

1 . A computer-implemented method comprising:

iteratively training, at least partially, by a programmed computer system that comprises one or more processor cores, via machine learning, a neural network to perform a training task; and

after the training:

generating, by the programmed computer system, a plurality of candidate specialty node units for the neural network, each comprising one or more nodes, and wherein each candidate specialty node unit has a specialty type selected from the group consisting of: a sparse node unit, a template node unit, a local autoencoder unit, a discriminator unit, a node unit with memory, a program node unit, and a second-order function node unit;

evaluating, by the programmed computer system, performance of the plurality of candidate specialty node units on the training task using a subset of training data;

selecting, by the programmed computer system, one or more of the candidate specialty node units based on the evaluation;

inserting, by the programmed computer system, the selected one or more candidate specialty node units into the neural network to create a modified neural network, wherein each of the selected on or more candidate specialty node units, when inserted, is functionally integrated into the neural network at an insertion location determined based on task-specific criteria; and

resuming, by the programmed computer system, iterative training of the modified neural network using training data.

2 . The method of claim 1 , wherein evaluating performance of the plurality of candidate specialty node units comprises computing a change in loss of the neural network when each candidate specialty node unit is provisionally inserted into the neural network.

3 . The method of claim 1 , wherein evaluating performance of the plurality of candidate specialty node units comprises computing a change in accuracy of the neural network when each candidate specialty node unit is provisionally inserted into the neural network.

4 . The method of claim 1 , wherein selecting one or more of the candidate specialty node units comprises excluding a candidate specialty node unit whose insertion results in degradation of neural network performance beyond a threshold value.

5 . The method of claim 1 , wherein the insertion location for a selected specialty node unit is determined based on gradient magnitude at nodes in the neural network.

6 . The method of claim 1 , wherein the insertion location for a selected specialty node unit is determined based on identification of a bottleneck in the neural network.

7 . The method of claim 1 , wherein the insertion location for a selected specialty node unit is determined based on activation sparsity of nodes in the neural network.

8 . The method of claim 1 , further comprising training each specialty node unit using a subset of the training data selected to emphasize a subdomain of the training task.

9 . The method of claim 8 , further comprising assigning a specialization tag to each selected specialty node unit based on the subdomain emphasized during training of that specialty node unit.

10 . The method of claim 1 , wherein a program node unit comprises a symbolic computational graph external to the neural network.

11 . The method of claim 1 , wherein a second-order function node unit is configured to compute a transformation of a parameter of another node within the neural network.

12 . The method of claim 1 , further comprising:

identifying a region within the neural network exhibiting sparse activation or a gradient bottleneck; and

training a specialty node unit using training data corresponding to a subdomain of the training task aligned with that region,

wherein the trained specialty node unit is inserted into the identified region of the neural network.

13 . The method of claim 1 , further comprising freezing weights of at least one inserted specialty node unit during resumed training of the modified neural network.

14 . The method of claim 1 , wherein generating the plurality of candidate specialty node units comprises instantiating the candidate specialty node units from a predefined library of reusable module templates.

15 . The method of claim 1 , further comprising logging insertion decisions and corresponding task performance results to a module registry configured to support reuse of specialty node units in other neural network training workflows.

16 . A computer system comprising:

one or more processor cores; and

a non-transitory computer-readable medium storing instructions that, when executed by the one or more processor cores, cause the computer system to perform operations comprising:

iteratively training, at least partially, via machine learning, a neural network to perform a training task;

after the training:

generating a plurality of candidate specialty node units for the neural network, each comprising one or more nodes, and wherein each candidate specialty node unit has a specialty type selected from the group consisting of: a sparse node unit, a template node unit, a local autoencoder unit, a discriminator unit, a node unit with memory, a program node unit, and a second-order function node unit;

evaluating performance of the plurality of candidate specialty node units on the training task using a subset of training data;

selecting one or more of the candidate specialty node units based on the evaluation;

inserting the selected one or more candidate specialty node units into the neural network to create a modified neural network, wherein each of the selected on or more candidate specialty node units, when inserted, is functionally integrated into the neural network at an insertion location determined based on task-specific criteria; and

resuming iterative training of the modified neural network using training data.

17 . The computer system of claim 16 , wherein the non-transitory computer-readable medium stores instructions that, when executed by the one or more processor cores, cause the computer system to evaluate performance of the plurality of candidate specialty node units by computing a change in loss of the neural network when each candidate specialty node unit is provisionally inserted into the neural network.

18 . The computer system of claim 16 , wherein the non-transitory computer-readable medium stores instructions that, when executed by the one or more processor cores, cause the computer system to evaluate performance of the plurality of candidate specialty node units by computing a change in accuracy of the neural network when each candidate specialty node unit is provisionally inserted into the neural network.

19 . The computer system of claim 16 , wherein the non-transitory computer-readable medium stores instructions that, when executed by the one or more processor cores, cause the computer system to select one or more of the candidate specialty node units by excluding a candidate specialty node unit whose insertion results in degradation of neural network performance beyond a threshold value.

20 . The computer system of claim 16 , wherein the insertion location for a selected specialty node unit is determined based on gradient magnitude at nodes in the neural network.

21 . The computer system of claim 16 , wherein the insertion location for a selected specialty node unit is determined based on identification of a bottleneck in the neural network.

22 . The computer system of claim 16 , wherein the insertion location for a selected specialty node unit is determined based on activation sparsity of nodes in the neural network.

23 . The computer system of claim 16 , wherein the instructions further cause the computer system to train each specialty node unit using a subset of the training data selected to emphasize a subdomain of the training task.

24 . The computer system of claim 23 , wherein the instructions further cause the computer system to assign a specialization tag to each selected specialty node unit based on the subdomain emphasized during training of that specialty node unit.

25 . The computer system of claim 16 , wherein a program node unit comprises a symbolic computational graph external to the neural network.

26 . The computer system of claim 16 , wherein a second-order function node unit is configured to compute a transformation of a parameter of another node within the neural network.

27 . The computer system of claim 16 , wherein the instructions further cause the computer system to:

identify a region within the neural network exhibiting sparse activation or a gradient bottleneck; and

train a specialty node unit using training data corresponding to a subdomain of the training task aligned with that region,

wherein the trained specialty node unit is inserted into the identified region of the neural network.

28 . The computer system of claim 16 , wherein the instructions further cause the computer system to freeze weights of at least one inserted specialty node unit during resumed training of the modified neural network.

29 . The computer system of claim 16 , wherein the non-transitory computer-readable medium stores instructions that, when executed by the one or more processor cores, cause the computer system to generate the plurality of candidate specialty node units by instantiating the candidate specialty node units from a predefined library of reusable module templates.

30 . The computer system of claim 16 , wherein the instructions further cause the computer system to log insertion decisions and corresponding task performance results to a module registry configured to support reuse of specialty node units in other neural network training workflows.

Continuity (8)
Continuation 19098299 · Apr 2, 2025
Continuation 18947318 · Nov 14, 2024
Continuation 18352044 · Jul 13, 2023
Continuation 16929900 · Jul 15, 2020
Continuation 16767966
Provisional Application 62647085 · Mar 23, 2018
Provisional Application 62623773 · Jan 30, 2018
Related Publication 20260004133A1 · Jan 1, 2026
References Cited (70)
US 5701398A · Glier et al. · 1997 [cited by applicant]
US 10832137B2 · Baker et al. · 2020 [cited by applicant]
US 10832138B2 · Choi · 2020 [cited by applicant]
US 10839294B2 · Baker · 2020 [cited by applicant]
US 10929757B2 · Baker et al. · 2021 [cited by applicant]
US 11010671B2 · Baker et al. · 2021 [cited by applicant]
US 11074499B2 · Burr · 2021 [cited by applicant]
US 11087217B2 · Baker et al. · 2021 [cited by applicant]
US 11093830B2 · Baker et al. · 2021 [cited by applicant]
US 11151455B2 · Baker et al. · 2021 [cited by applicant]
US 11321612B2 · Baker et al. · 2022 [cited by applicant]
US 11386330B2 · Baker · 2022 [cited by applicant]
US 11410050B2 · Baker · 2022 [cited by applicant]
US 11461655B2 · Baker et al. · 2022 [cited by applicant]
US 11461661B2 · Baker · 2022 [cited by applicant]
US 11610130B2 · Baker · 2023 [cited by applicant]
US 11748624B2 · Baker et al. · 2023 [cited by applicant]
US 12182712B2 · Baker et al. · 2024 [cited by applicant]
US 12288161B2 · Baker et al. · 2025 [cited by applicant]
US 20040015459A1 · Jaeger · 2004 [cited by applicant]
US 20040073096A1 · Kates et al. · 2004 [cited by applicant]
US 20040128004A1 · Adams et al. · 2004 [cited by applicant]
US 20040260662A1 · Staelin et al. · 2004 [cited by applicant]
US 20050283450A1 · Matsugu et al. · 2005 [cited by applicant]
US 20080065572A1 · Abe et al. · 2008 [cited by applicant]
US 20080069437A1 · Baker · 2008 [cited by applicant]
US 20080154820A1 · Kirshenbaum et al. · 2008 [cited by applicant]
US 20150019467A1 · Nugent · 2015 [cited by applicant]
US 20150231680A1 · Jones · 2015 [cited by examiner]
US 20160110642A1 · Matsuda et al. · 2016 [cited by applicant]
US 20160342904A1 · Merkel · 2016 [cited by applicant]
US 20170105791A1 · Yates et al. · 2017 [cited by applicant]
US 20180129959A1 · Gustafson · 2018 [cited by examiner]
US 20190325343A1 · Feng et al. · 2019 [cited by applicant]
US 20200242466A1 · Mohassel · 2020 [cited by applicant]
US 20250232175A1 · Baker et al. · 2025 [cited by applicant]
WO 2018175098A1 · 2018 [cited by applicant]
WO 2018194960A1 · 2018 [cited by applicant]
WO 2018226492A1 · 2018 [cited by applicant]
WO 2018226527A1 · 2018 [cited by applicant]
WO 2018231708A2 · 2018 [cited by applicant]
WO 2019005507A1 · 2019 [cited by applicant]
WO 2019005611A1 · 2019 [cited by applicant]
WO 2019067236A1 · 2019 [cited by applicant]
WO 2019067248A1 · 2019 [cited by applicant]
WO 2019067281A1 · 2019 [cited by applicant]
WO 2019067542A1 · 2019 [cited by applicant]
WO 2019067831A1 · 2019 [cited by applicant]
WO 2019067960A1 · 2019 [cited by applicant]
WO 2019152308A1 · 2019 [cited by applicant]
WO 2020005471A1 · 2020 [cited by applicant]
WO 2020009881A1 · 2020 [cited by applicant]
WO 2020009912A1 · 2020 [cited by applicant]
WO 2020018279A1 · 2020 [cited by applicant]
WO 2020028036A1 · 2020 [cited by applicant]
WO 2020033645A1 · 2020 [cited by applicant]
WO 2020036847A1 · 2020 [cited by applicant]
WO 2020041026A1 · 2020 [cited by applicant]
WO 2020046719A1 · 2020 [cited by applicant]
WO 2020046721A1 · 2020 [cited by applicant]
Glacet et al., The impact of Edge Deletions on the Number of Errors in Networks, In:OPODIS'11 Proceedings of the 15th International Conference on Principles of Distributed System, Dec. 13-16, 2011; https://www.researchg… [cited by applicant]
International Search Report and Written Opinion for International PCT Application No. PCT/US2019/015389, dated Apr. 25, 2019. [cited by applicant]
International Preliminary Report on Patentability for International PCT Application No. PCT/US2019/015389, dated Sep. 4, 2020. [cited by applicant]
Fu, Li-Min, and Li-Chen Fu. “Mapping rule-based systesm into neural architecture.” Knowledge-Based Systems 3, No. 1 (1990):48-56 (Year: 1990). [cited by applicant]
Petrowski, A. “Choosing among several parallel implementations of the backpropagation algorithm.” In Proceedings of 1994 IEEE International Conference on Neural Networks (ICNN'94), vol. 3, pp. 1981-1986. IEEE, 1994. (Ye… [cited by applicant]
Lehtokangas, M. (1999). Modelling with constructive backpropagation. Neural Networks, 12(4-5), 707-716. (Year: 1999). [cited by applicant]
Rojas, Raul. “The backpropagation algorithm.” In Neural networks, pp. 149-182. Springer, Berlin, Heidelberg, 1996. (Year: 1996). [cited by applicant]
Valle M. Analog VLSI implementation of artificial neural networks with supervised on-chip learning. Analog Integrated Circuits and Signal Processing. Dec. 2002;33:263-87. (Year: 2002). [cited by applicant]
Sumpter BG, Getino C, Noid DW. Theory and applications of neural computing in chemical science. Annual Review of Physical Chemistry. Oct. 1994; 45(1):439-81. (Year: 1994). [cited by applicant]
Sethi K. Entropy nets: from decision trees to neural networks. Proceedings of the IEEE. Oct. 1990;78(10):1605-13. (Year: 1990). [cited by applicant]