IP Library › Granted Patent US 12,020,160
Granted Patent B2
US 12,020,160 · App. 15/875,575 · Granted Jun 25, 2024

Generation of neural network containing middle layer background

Inventor: Takeshi Inagaki (Sagamihara, JP)
Assignee: International Business Machines Corporation
G06N3/082G06F18/214G06F18/24133G06N3/045G06N3/088G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,020,160
App. No.
15/875,575
Granted
Jun 25, 2024
Kind
B2
Abstract

A method, computer program product and system for generating a neural network. Initial neural networks are prepared, each of which includes an input layer containing one or more input nodes, a middle layer containing one or more middle nodes, and an output layer containing one or more output nodes. A new neural network is generated that includes a new middle layer containing one or more middle nodes based on the middle nodes of the middle layers of the initial neural networks.

Claims (48)

1. A computer program product for generating a neural network, the computer program product comprising a computer readable storage medium having program code embodied therewith, the program code comprising the programming instructions for:

preparing a plurality of initial neural networks, each of which comprises an input layer containing one or more input nodes, a middle layer containing one or more middle nodes, and an output layer containing one or more output nodes;

performing supervised training of the output layer of each of the plurality of initial neural networks using a set of training data; and

generating a new neural network comprising a new middle layer containing one or more middle nodes based on the middle nodes of the middle layers of the plurality of initial neural networks.

2. The computer program product as recited in claim 1 , wherein the plurality of initial neural networks comprises N initial neural networks, N being an integer larger than 1, and wherein the generating of the new neural network comprises the programming instructions for:

selecting one or more of the middle nodes of the N initial neural networks; and

including the selected one or more middle nodes in the new middle layer of the new neural network.

3. The computer program product as recited in claim 2 , wherein the selecting of the one or more of the middle nodes of the N initial neural networks comprises the programming instructions for:

obtaining K different sets of training data, K being an integer more than 1;

performing supervised training on the N initial neural networks with each of the K different sets of training data to obtain K training results for each of the N initial neural networks; and

selecting at least one of the middle nodes in the middle layer of the N initial neural networks using the K training results, such that selected middle nodes contribute to an output from the output layer to a greater degree than non-selected middle nodes.

4. The computer program product as recited in claim 2 , wherein the middle layer of each of the plurality of initial neural network comprises L middle nodes, L being an integer larger than 2, and wherein the number of the middle nodes in the new middle layer is equal to or less than L.

5. The computer program product as recited in claim 2 , wherein the generating of the new neural network further comprises the programming instructions for:

performing unsupervised training on the selected middle nodes, the unsupervised training comprising biasing the middle nodes such that certain middle nodes are avoided.

6. The computer program product as recited in claim 1 , wherein the preparing of the plurality of initial neural networks comprises the programming instructions for:

obtaining N initial conditions, N being an integer larger than 1, each condition corresponding to one of the initial neural networks; and

performing unsupervised training of the middle layer of each initial neural network using the corresponding initial condition.

7. The computer program product as recited in claim 1 , wherein the preparing of the plurality of initial neural networks comprises the programming instructions for:

obtaining M initial conditions, M being an integer larger than 2, each condition corresponding to one of M candidate neural networks;

performing unsupervised training of the middle layer of each candidate neural network using the corresponding initial condition;

performing supervised training of the output layer of each candidate neural network using a set of training data;

evaluating a performance of each candidate neural network; and

selecting N initial neural networks from among the M candidate neural networks using the performances, N being an integer larger than 1 and smaller than M.

8. A system, comprising:

a memory unit for storing a computer program for generating a neural network; and

a processor coupled to the memory unit, wherein the processor is configured to execute the program instructions of the computer program comprising:

preparing a plurality of initial neural networks, each of which comprises an input layer containing one or more input nodes, a middle layer containing one or more middle nodes, and an output layer containing one or more output nodes;

performing supervised training of the output layer of each of the plurality of initial neural networks using a set of training data; and

generating a new neural network comprising a new middle layer containing one or more middle nodes based on the middle nodes of the middle layers of the plurality of initial neural networks.

9. The system as recited in claim 8 , wherein the plurality of initial neural networks comprises N initial neural networks, N being an integer larger than 1, and wherein the generating of the new neural network comprises:

selecting one or more of the middle nodes of the N initial neural networks; and

including the selected one or more middle nodes in the new middle layer of the new neural network.

10. The system as recited in claim 9 , wherein the selecting of the one or more of the middle nodes of the N initial neural networks comprises:

obtaining K different sets of training data, K being an integer more than 1;

performing supervised training on the N initial neural networks with each of the K different sets of training data to obtain K training results for each of the N initial neural networks; and

selecting at least one of the middle nodes in the middle layer of the N initial neural networks using the K training results, such that selected middle nodes contribute to an output from the output layer to a greater degree than non-selected middle nodes.

11. The system as recited in claim 9 , wherein the middle layer of each of the plurality of initial neural network comprises L middle nodes, L being an integer larger than 2, and wherein the number of the middle nodes in the new middle layer is equal to or less than L.

12. The system as recited in claim 9 , wherein the generating of the new neural network further comprises:

performing unsupervised training on the selected middle nodes, the unsupervised training comprising biasing the middle nodes such that certain middle nodes are avoided.

13. The system as recited in claim 8 , wherein the preparing of the plurality of initial neural networks comprises:

obtaining N initial conditions, N being an integer larger than 1, each condition corresponding to one of the initial neural networks; and

performing unsupervised training of the middle layer of each initial neural network using the corresponding initial condition.

14. The system as recited in claim 8 , wherein the preparing of the plurality of initial neural networks comprises:

obtaining M initial conditions, M being an integer larger than 2, each condition corresponding to one of M candidate neural networks;

performing unsupervised training of the middle layer of each candidate neural network using the corresponding initial condition;

performing supervised training of the output layer of each candidate neural network using a set of training data;

evaluating a performance of each candidate neural network; and

selecting N initial neural networks from among the M candidate neural networks using the performances, N being an integer larger than 1 and smaller than M.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2018
From: INAGAKI, TAKESHI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 045095/0713 →
Continuity (1)
Related Publication 20190228310A1 · Jul 25, 2019