Neural cluster and neural processing SoC including the same
According to one or more embodiments of the present disclosure, a neural cluster includes a plurality of neural core units each including a neural core configured to process a neural network operation, a plurality of shared memory units each including a shared memory shared by the plurality of neural core units, and a mesh network for connecting the plurality of neural core units and the plurality of shared memory units, wherein the plurality of shared memory units are arranged in a central portion of the neural cluster, and the plurality of neural core units are arranged symmetrically above and below a shared memory area where the plurality of shared memory units are arranged.
1 . A neural cluster, comprising:
a plurality of neural core units each including a neural core configured to process a neural network operation;
a plurality of shared memory units each including a shared memory shared by the plurality of neural core units; and
a mesh network for connecting the plurality of neural core units and the plurality of shared memory units,
wherein the plurality of shared memory units are arranged in a central portion of the neural cluster, and the plurality of neural core units are arranged symmetrically above and below a shared memory area where the plurality of shared memory units are arranged,
wherein the mesh network includes a plurality of routers each arranged at an intersection formed by a row line and a column line, and a mesh network bus, and
wherein each of the plurality of routers is connected to one of the plurality of neural core units or one of the plurality of shared memory units through the mesh network bus, and wherein each of the plurality of routers includes:
a first port configured to receive a first data packet having a first flag;
a second port configured to receive a second data packet having a second flag and having the same destination as the first data packet;
a third port connected to the destination of the first data packet and the second data packet; and
an arbiter configured to determine whether an atomic transfer is applied to each of the first data packet and the second data packet, based on the first flag and the second flag.
2 . The neural cluster of claim 1 , wherein each of the plurality of routers is further connected to one or more adjacent routers through the mesh network bus.
3 . The neural cluster of claim 1 , wherein each of the plurality of neural core units and each of the plurality of shared memory units include a network interface unit,
wherein the mesh network bus includes a data channel, a response channel, and a request channel, and
wherein the network interface unit is configured to map an AXI-AW channel, an AXI-W channel, an AXI-B channel, an AXI-AR channel, and an AXI-R channel according to an AMBA (Advanced Microcontroller Bus Architecture) AXI (Advanced extensible Interface) protocol to the data channel, the response channel, and the request channel.
4 . The neural cluster of claim 3 , wherein the AXI-AW channel, the AXI-W channel, and the AXI-R channel are mapped to the data channel, the AXI-B channel is mapped to the response channel, and the AXI-AR channel is mapped to the request channel.
5 . The neural cluster of claim 1 , wherein the arbiter is configured to:
check a first input enable signal generated from the first port upon reception of the first data packet;
check a second input enable signal generated from the second port upon reception of the second data packet; and
first check the first flag based on a predetermined criterion if the first input enable signal and the second input enable signal are activated.
6 . The neural cluster of claim 5 , wherein the arbiter is configured to mask the second input enable signal and transmit the first data packet to the third port if the first flag indicates the atomic transfer.
7 . The neural cluster of claim 6 , wherein the arbiter is configured to maintain the transmission of one or more data packets, which have one or more flags indicating the atomic transfer received after the first data packet at the first port, to the third port until a data packet having a flag indicating a non-atomic transfer is received at the first port.
8 . A neural cluster, comprising:
a first neural core unit including a first neural core configured to process a neural network operation;
a second neural core unit including a second neural core configured to process a neural network operation;
a first shared memory unit including a first shared memory shared by the first neural core unit and the second neural core unit;
a second shared memory unit including a second shared memory shared by the first neural core unit and the second neural core unit; and
a mesh network for connecting the first neural core unit, the second neural core unit, the first shared memory unit, and the second shared memory unit,
wherein the first neural core is configured to generate a first data access request in a first cycle and generate a second data access request in a second cycle,
wherein the second neural core is configured to generate a third data access request in the first cycle and generate a fourth data access request in the second cycle,
wherein the first to fourth data access requests are respectively interleaved and transmitted as distributed to the first shared memory and the second shared memory, and
wherein the first shared memory unit and the second shared memory unit are arranged in a central portion of the neural cluster, and the first neural core unit and the second neural core unit are respectively arranged symmetrically above and below a shared memory area where the first shared memory unit and the second shared memory unit are arranged.
9 . The neural cluster of claim 8 , wherein a size of data accessed according to the first to fourth data access requests is the same as each other.
10 . The neural cluster of claim 9 , wherein the first to fourth data access requests are respectively interleaved according to an interleaving unit and transmitted as distributed to the first shared memory and the second shared memory, and the interleaving unit is changeable.
11 . The neural cluster of claim 10 , wherein if the interleaving unit is the same as the size of the data accessed according to each of the first to fourth data access requests, the first data access request and the third data access request are transmitted to the first shared memory, and the second data access request and the fourth data access request are transmitted to the second shared memory.
12 . The neural cluster of claim 10 , wherein the first neural core is configured to further generate a fifth data access request in a third cycle and further generate a sixth data access request in a fourth cycle,
wherein the second neural core is configured to further generate a seventh data access request in the third cycle and further generate an eighth data access request in the fourth cycle, and
wherein if the interleaving unit is twice the size of the data accessed according to each of the first to fourth data access requests, the first to fourth data access requests are transmitted to the first shared memory, and the fifth to eighth data access requests are transmitted to the second shared memory.
13 . The neural cluster of claim 8 , wherein each of the first neural core unit and the second neural core unit further includes a network interface unit.
14 . The neural cluster of claim 13 , wherein the first neural core is configured to further generate a first system address with the first data access request and further generate a second system address with the second data access request,
wherein the second neural core is configured to further generate a third system address with the third data access request and further generate a fourth system address with the fourth data access request,
wherein the network interface unit of the first neural core unit is configured to parse the first system address and the second system address according to a predetermined parsing rule, wherein the network interface unit of the second neural core unit is configured to parse the third system address and the fourth system address according to the predetermined parsing rule, and
wherein, according to the parsed first to fourth system addresses, the first to fourth data access requests are respectively interleaved and transmitted as distributed to the first shared memory and the second shared memory.
15 . The neural cluster of claim 14 , wherein the second system address is an address consecutive to the first system address, and the fourth system address is an address consecutive to the third system address.
16 . A neural processing SoC (System on a Chip), comprising:
a first neural cluster; and
a second neural cluster,
wherein each of the first neural cluster and the second neural cluster includes:
a plurality of neural core units each including a neural core configured to process a neural network operation;
a plurality of shared memory units each including a shared memory shared by the plurality of neural core units; and
a mesh network for connecting the plurality of neural core units and the plurality of shared memory units,
wherein the plurality of shared memory units are arranged in a central portion of each of the first neural cluster and the second neural cluster, and the plurality of neural core units are arranged symmetrically above and below a shared memory area where the plurality of shared memory units are arranged, and
wherein a plurality of shared memories of the first neural cluster are shared by the plurality of neural core units of the second neural cluster,
wherein the mesh network includes a plurality of routers each arranged at an intersection formed by a row line and a column line, and a mesh network bus, and
wherein each of the plurality of routers is connected to one of the plurality of neural core units or one of the plurality of shared memory units through the mesh network bus, and wherein each of the plurality of routers includes:
a first port configured to receive a first data packet having a first flag;
a second port configured to receive a second data packet having a second flag and having the same destination as the first data packet;
a third port connected to the destination of the first data packet and the second data packet; and
an arbiter configured to determine whether an atomic transfer is applied to each of the first data packet and the second data packet, based on the first flag and the second flag.
17 . A neural processing SoC, comprising:
a first neural cluster; and
a second neural cluster,
wherein each of the first neural cluster and the second neural cluster includes:
a first neural core unit including a first neural core configured to process a neural network operation;
a second neural core unit including a second neural core configured to process a neural network operation;
a first shared memory unit including a first shared memory shared by the first neural core unit and the second neural core unit;
a second shared memory unit including a second shared memory shared by the first neural core unit and the second neural core unit; and
a mesh network for connecting the first neural core unit, the second neural core unit, the first shared memory unit, and the second shared memory unit,
wherein the first neural core is configured to generate a first data access request in a first cycle and generate a second data access request in a second cycle,
wherein the second neural core is configured to generate a third data access request in the first cycle and generate a fourth data access request in the second cycle,
wherein the first to fourth data access requests are respectively interleaved and transmitted as distributed to the first shared memory and the second shared memory,
wherein the first shared memory unit and the second shared memory unit are arranged in a central portion of each of the first neural cluster and the second neural cluster, and the first neural core unit and the second neural core unit are respectively arranged symmetrically above and below a shared memory area where the first shared memory unit and the second shared memory unit are arranged, and
wherein the first shared memory and the second shared memory of the first neural cluster are shared by the first neural core unit and the second neural core unit of the second neural cluster.