IP Library › Granted Patent US 11,301,713
Granted Patent B2
US 11,301,713 · App. 16/758,196 · Granted Apr 12, 2022

Information processing apparatus, information processing method, and non-transitory computer readable medium

Inventor: Salita Sombatsiri (Tokyo, JP)
Assignee: NEC CORPORATION
G06K9/4628G06F17/153G06K9/6267G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,301,713
App. No.
16/758,196
Granted
Apr 12, 2022
Kind
B2
Abstract

An object is to provide an information processing apparatus capable of preventing utilization percentage of PEs from decreasing in a series of processes in CNN. An information processing apparatus ( 1 ) according to the present disclosure includes a PE (Processing Element) Grid ( 20 ) configured to perform a convolution by using a plurality of Kernels for Input matrix data and thereby generate a different Output matrix data for each of the used Kernels, the PE Grid ( 20 ) including a plurality of PEs configured to calculate pixels constituting the Output matrix data, and a Parallelism Controller ( 10 ) configured to determine, based on the Input matrix data or a dimension of the Output matrix data, and the number of the Kernels, whether pixels included in respective Output matrix data should be parallelly calculated or a plurality of pixels included in one Output matrix data should be parallelly calculated.

Claims (26)

1. An information processing apparatus comprising:

a PE (Processing Element) Grid configured to perform a convolution by using a plurality of Kernels for Input matrix data and thereby generate a different Output matrix data for each of the used Kernels, the PE Grid including a plurality of PEs configured to calculate pixels constituting the Output matrix data; and

a Parallelism Controller configured to determine, based on a dimension of the Input matrix data or a dimension of the Output matrix data, and the number of the Kernels, whether pixels included in respective Output matrix data should be parallelly calculated or a plurality of pixels included in one Output matrix data should be parallelly calculated.

2. The information processing apparatus according to claim 1 , wherein the Parallelism Controller determines whether the pixels included in respective Output matrix data should be parallelly calculated or the plurality of pixels included in one Output matrix data should be parallelly calculated so that the number of PEs that perform computations, among the plurality of PEs, increases.

3. The information processing apparatus according to claim 1 , further comprising a Sparse Weight Broadcaster configured to determine Kernels applied to the plurality of PEs, wherein

the Parallelism Controller determines the number of Kernels to be used at the same time in the PE Grid and outputs information about the determined number of Kernels to the Sparse Weight Broadcaster, and

the Sparse Weight Broadcaster assigns Kernels to the plurality of PEs so that the determined number of Kernels parallelly perform computations.

4. The information processing apparatus according to claim 3 , wherein the Sparse Weight Broadcaster outputs index information including Kernel IDs for identifying the Kernels and element IDs for identifying a plurality of elements included in the Kernels, and Weight values associated with the Kernel IDs and the element IDs, to the plurality of PEs.

5. The information processing apparatus according to claim 4 , wherein

the Parallelism Controller outputs address information for identifying data to be extracted from the Input matrix data to the plurality of PEs, and

the plurality of PEs extract data corresponding to the address information from the Input matrix data.

6. The information processing apparatus according to claim 5 , wherein

each of PEs to which the same Kernel is assigned extracts different data from the Input matrix data, and

each of PEs to which different Kernels are assigned extracts the same data from the Input matrix data.

7. The information processing apparatus according to claim 3 , wherein the Sparse Weight Broadcaster assigns the same Kernel to all the PEs included in the PE Grid, assigns different Kernels to all the PEs included in the PE Grid, or divides the plurality of PEs included in the GE Grid into two or more groups and assigns a different Kernel to each of the groups.

8. The information processing apparatus according to claim 3 , wherein

the PEs calculate values of pixels by using Non-zero values included in the Kernels, and

the Sparse Weight Broadcaster assigns the Kernels to respective PEs so that a difference between the numbers of Non-zero values that are used when the respective PEs calculate pixels becomes smaller than a predetermined threshold.

9. An information processing method comprising:

performing a convolution by using a plurality of Kernels for Input matrix data and thereby acquiring a dimension of Output matrix data generated for each of the used Kernels or a dimension of the Input matrix data, and the number of the Kernels;

determining, based on the dimension of the Output matrix data or the dimension of the Input matrix data, and the number of the Kernels, whether pixels included in respective Output matrix data should be parallelly calculated or the plurality of pixels included in one Output matrix data should be parallelly calculated; and

generating the Output matrix data based on the calculation method.

10. A non-transitory computer readable medium storing a program for causing a computer to execute:

performing a convolution by using a plurality of Kernels for Input matrix data and thereby acquire a dimension of Output matrix data generated for each of the used Kernels or a dimension of the Input matrix data, and the number of the Kernels;

determining, based on the dimension of the Output matrix data or the dimension of the Input matrix data, and the number of the Kernels, whether pixels included in respective Output matrix data should be parallelly calculated or the plurality of pixels included in one Output matrix data should be parallelly calculated; and

generating the Output matrix data based on the calculation method.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2020
From: SOMBATSIRI, SALITA
To: NEC CORPORATION
Reel/Frame 052463/0729 →
Continuity (1)
Related Publication 20200302215A1 · Sep 24, 2020