Machine learning through multiple layers of novel machine trained processing nodes
Some embodiments of the invention provide efficient, expressive machined-trained networks for performing machine learning. The machine-trained (MT) networks of some embodiments use novel processing nodes with novel activation functions that allow the MT network to efficiently define with fewer processing node layers a complex mathematical expression that solves a particular problem (e.g., face recognition, speech recognition, etc.). In some embodiments, the same activation function (e.g., a cup function) is used for numerous processing nodes of the MT network, but through the machine learning, this activation function is configured differently for different processing nodes so that different nodes can emulate or implement two or more different functions (e.g., two or more Boolean logical operators, such as XOR and AND). The activation function in some embodiments is a periodic function that can be configured to implement different functions (e.g., different sinusoidal functions).
1 . A method comprising:
determining, based on first inference-time data and based on first weight data representing weights associated with a first layer of a machine-trained (MT) network, second inference-time data representing a first set of values;
determining, based on the second inference-time data and configuration data associated with the first layer of the MT network, third inference-time data representing a second set of values, wherein the configuration data was determined during training of the MT network and defines activation functions for each node of the first layer of the MT network, and wherein determining the third inference-time data comprises:
determining, using a first activation function of a first node of the first layer, a first value based on the second inference-time data and a third set of values, indicated by the configuration data, representing parameters of the first activation function, wherein the parameters at least partially define a first piecewise linear function, and
determining, using a second activation function of a second node of the first layer, a second value based on the second inference-time data and a fourth set of values, indicated by the configuration data, representing parameters of the second activation function, wherein the parameters at least partially define a second piecewise linear function different than the first piecewise linear function; and
generating, based on the third inference-time data, output data representing output of the MT network.
2 . The method of claim 1 , wherein the third set of values represents parameters at least partially defining a first piecewise linear cup function, and wherein the fourth set of values represents parameters at least partially defining a second piecewise linear cup function different than the first piecewise linear cup function.
3 . The method of claim 1 , wherein the third set of values represents parameters including a first parameter indicating a domain limitation for a first linear equation of the first piecewise linear function.
4 . The method of claim 1 , wherein the third set of values represents parameters including a first parameter and a second parameter indicating domain limitations for a first linear equation of the first piecewise linear function.
5 . The method of claim 1 , wherein the third set of values represents parameters including a first parameter indicating a parameter value for a first linear equation associated with a first portion of a domain of the first piecewise linear function.
6 . The method of claim 1 , wherein the third set of values represents parameters including a first parameter indicating a slope value for a first linear equation associated with a first portion of a domain of the first piecewise linear function.
7 . The method of claim 1 , wherein the third set of values represents parameters including a first parameter indicating a y-intercept value for a first linear equation associated with a first portion of a domain of the first piecewise linear function.
8 . The method of claim 1 , wherein the third set of values represents parameters including a first parameter indicating a constant range value for a first portion of a domain of the first piecewise linear function.
9 . The method of claim 1 , wherein the first piecewise linear function being defined to have:
a first domain interval associated with a linear function having a negative slope,
a second domain interval associated with a linear function having zero slope, and
a third domain interval associated with a linear function having a positive slope.