IP Library › Granted Patent US 11,244,227
Granted Patent B2
US 11,244,227 · App. 16/290,250 · Granted Feb 8, 2022

Discrete feature representation with class priority

Inventor: Masataro Asai (Tokyo, JP)
Assignee: International Business Machines Corporation
G06N3/082G06K9/628G06N3/0481G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,244,227
App. No.
16/290,250
Granted
Feb 8, 2022
Kind
B2
Abstract

A discrete neural network is trained by training a neural network having an output layer so as to output discrete values. The output layer includes a plurality of nodes. Each node corresponding to one of a plurality of classes. The training includes activating the nodes by priority according to the corresponding class.

Claims (38)

1. A computer-implemented method for neural network training, comprising:

training a neural network having an output layer that utilizes continuous values so that the output layer of the neural network will tend to output discrete values such that a temperature parameter τ is utilized for annealing during the training by decreasing τ from a positive value to 0, wherein a softmax value z i,j approaches a discrete value from a continuous value as τ approaches 0, wherein the output layer includes a plurality of nodes, each node corresponding to one of a plurality of classes;

assigning a priority to at least one class of the plurality of classes; and

activating the nodes by priority according to the corresponding class of the plurality of classes.

2. The method of claim 1 , wherein the training further includes regularizing each class by minimizing a network loss function that applies a penalty term associated with the priority of the corresponding class.

3. The method of claim 2 , wherein the network loss function includes activation levels of the plurality of nodes except nodes of one class.

4. The method of claim 2 , wherein the network loss function includes activation levels of the plurality of nodes except nodes of two or more classes.

5. The method of claim 2 , wherein the network loss function includes weighted activation levels of the plurality of nodes of all classes, the weighted activation levels being weighted according to the priority of the corresponding class.

6. The method of claim 1 , wherein the output layer comprises a plurality of sets of nodes, each set of nodes corresponding to one of a plurality of variables, and each set of the plurality of sets includes one of the plurality of nodes corresponding to each class,

wherein the method further comprises:

identifying a variable of the plurality of variables of which only a particular class is activated regardless of input data to the neural network, and

deleting a set of nodes corresponding to the identified variable.

7. The method of claim 1 , wherein the output layer comprises a plurality of sets of nodes, each set of nodes corresponding to one of a plurality of variables, and each set of the plurality of sets includes one of the plurality of nodes corresponding to each class,

wherein the method further comprises:

identifying a class that is not activated throughout the plurality of variables regardless of input data to the neural network, and

deleting nodes corresponding to the identified class throughout the plurality of variables.

8. The method of claim 6 , wherein, during the training of the neural network, the plurality of nodes of each set calculate the softmax value based at least on logit values of outputs from nodes in a previous layer connected to the output layer and a sample of a predetermined distribution.

9. The method of claim 6 , wherein, during the training of the neural network, the plurality of nodes of each set calculate the softmax value base at least on logit values of outputs from nodes of a previous layer connected to the output layer, a sample of Gumbel distribution, and a temperature parameter.

10. The method of claim 8 , further comprises:

replacing the output layer used at the training with an argmax layer.

11. The method of claim 1 , wherein the neural network is a Variational Autoencoder (VAE), and the output layer is included in an encoder of the VAE.

12. The method of claim 11 , wherein output from the encoder is used for input to a problem solver.

13. An apparatus comprising

a processor or a programmable circuitry; and

one or more computer readable mediums collectively including instructions that, when executed by the processor or the programmable circuitry, cause the processor or the programmable circuitry to perform operations including:

training a neural network having an output layer that utilizes continuous values so that the output layer of the neural network will tend to output discrete values such that a temperature parameter τ is utilized for annealing during the training by decreasing τ from a positive value to 0, wherein a softmax value z i,j approaches a discrete value from a continuous value as τ approaches 0, wherein the output layer includes a plurality of nodes, each node corresponding to one of a plurality of classes;

assigning a priority to at least one class of the plurality of classes; and

activating the nodes by priority according to the corresponding class of the plurality of classes.

14. The apparatus of claim 13 , wherein the training further includes regularizing each class by minimizing a network loss function that applies a penalty term associated with the priority of the corresponding class.

15. The apparatus of claim 14 , wherein the network loss function includes activation levels of the plurality of nodes except nodes of one class.

16. The apparatus of claim 14 , wherein the network loss function includes activation levels of the plurality of nodes except nodes of two or more classes.

17. A computer program product including one or more computer readable storage mediums collectively storing program instructions that are executable by a processor or programmable circuitry to cause the processor or programmable circuitry to perform operations comprising:

training a neural network having an output layer that utilizes continuous values so that the output layer of the neural network will tend to output discrete values such that a temperature parameter τ is utilized for annealing during the training by decreasing τ from a positive value to 0, wherein a softmax value z i,j approaches a discrete value from a continuous value as τ approaches 0, wherein the output layer includes a plurality of nodes, each node corresponding to one of a plurality of classes;

assigning a priority to at least one class of the plurality of classes; and

activating the nodes by priority according to the corresponding class of the plurality of classes.

18. The computer program product of claim 17 , wherein the training further includes regularizing each class by minimizing a network loss function that applies a penalty term associated with the priority of the corresponding class.

19. The computer program product of claim 18 , wherein the network loss function includes activation levels of the plurality of nodes except nodes of one class.

20. The computer program product of claim 18 , wherein the network loss function includes activation levels of the plurality of nodes except nodes of two or more classes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2019
From: ASAI, MASATARO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048481/0232 →
Continuity (1)
Related Publication 20200279164A1 · Sep 3, 2020
Cited By (1)
US 12,737,692