IP Library Granted Patent US 12,423,378
Granted Patent B2
US 12,423,378 · App. 17/228,895 · Granted Sep 23, 2025

Electronic device having graphics processor and acceleration method thereof

Inventors: Wei Zhang (Shanghai, CN); Deming Gu (Shanghai, CN)
Assignee: GLENFLY TECHNOLOGY CO., LTD
G06F17/16G06F17/15G06T1/20G06T1/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,378
App. No.
17/228,895
Granted
Sep 23, 2025
Kind
B2
Abstract

A graphics processor includes a texel unit and an execution unit. The texel unit includes a loading module. The execution unit includes an im2col module to execute an im2col algorithm to expand an original matrix to obtain an expansion matrix according to the size of a kernel. The execution unit multiplies the expansion matrix and the kernel to obtain a feature map matrix. The loading module calculates feature coordinates of each element of the feature map matrix according to the coordinates of the expansion matrix, and obtains the original coordinates of each element of the original matrix according to the feature coordinates, the size of the kernel, a stride, and padding. The loading module reads at least one of the memory blocks covered by the original coordinates of each element of the original matrix, and outputs data corresponding to the original coordinates in the memory blocks.

Claims (36)

1. An electronic device, comprising:

a graphics processor, configured to accelerate a convolution calculation, configured to:

read an original matrix used for the convolution calculation from a memory outside the graphics processor; wherein the memory comprises a plurality of memory blocks, each of which is adjacent to another and is the same size, and the original matrix is stored by a specific data configuration in at least one of the memory blocks;

execute an im2col algorithm, expand the original matrix to obtain an expansion matrix according to a size of a kernel, and define expansion coordinates of each element in the expansion matrix;

multiply the expansion matrix and the kernel to obtain a feature map matrix corresponding to the original matrix;

receive the expansion coordinates, calculate feature coordinates of each element of the feature map matrix according to the expansion coordinates, and obtain original coordinates of each element of the original matrix according to the feature coordinates, the size of the kernel, a stride, and padding; and

read the at least one of the memory blocks covered by the original coordinates of each element of the original matrix, and sends data corresponding to the original coordinates in the at least one of the memory blocks into the im2col algorithm;

wherein the graphics processor comprises:

a return buffer, receiving and storing data of the original matrix, or the data corresponding to the original coordinates in the at least one memory blocks;

a data expander, expanding the original matrix using an im2col operation to obtain the expansion matrix;

a data multiplexer, selecting data required for the convolution calculation in the expansion matrix according to the graphics processor; and

an output merge buffer, combining the data in the expansion matrix selected by the data multiplexer, and outputting the combined data selected by the data multiplexer to a register file.

2. The electronic device as claimed in claim 1 , wherein the graphics processor further comprises the register file, to store the data in the original matrix, data in the expansion matrix, and data in the feature map matrix in the convolution calculation.

3. The electronic device as claimed in claim 2 , wherein the graphics processor executes the convolution calculation according to the data in the original matrix, the data in the expansion matrix, and the data in the feature map matrix in the register file.

4. The electronic device as claimed in claim 1 , wherein the graphics processor further comprises an L1 cache; in the convolution calculation, the L1 cache reads and stores the original matrix for the convolution calculation from the memory for the graphics processor to access.

5. The electronic device as claimed in claim 1 , wherein the graphics processor further comprises a second memory, to store a result of the convolution calculation executed by the graphics processor in the memory.

6. The electronic device as claimed in claim 1 , wherein a size of each memory block included in the memory is a matrix size of 4*8.

7. The electronic device as claimed in claim 1 , wherein the size of the kernel is a matrix size of 3*3, the stride is equal to 1, and the padding is equal to 0.

8. A method for accelerating a convolution calculation, applied to a graphics processor, comprising:

the graphics processor receiving an original matrix from a memory outside the graphics processor; wherein the memory comprises a plurality of memory blocks, each of which is adjacent to another and is the same size, and the original matrix is stored by a specific data configuration in at least one of the memory blocks;

the graphics processor executing an im2col algorithm, and expanding the original matrix to obtain an expansion matrix according to a size of a kernel; wherein each element in the expansion matrix has expansion coordinates;

the graphics processor multiplying the expansion matrix and the kernel to obtain a feature map matrix corresponding to the original matrix;

the graphics processor calculating feature coordinates of each element of the feature map matrix according to the expansion coordinates;

the graphics processor obtaining original coordinates of each element of the original matrix according to the feature coordinates, the size of the kernel, a stride, and padding;

the graphics processor reading the at least one of the memory blocks covered by the original coordinates of each element of the original matrix, and outputting data corresponding to the original coordinates in the at least one of the memory blocks;

wherein the step of executing the im2col algorithm comprises:

receiving and storing by a return buffer data of the original matrix, or the data corresponding to the original coordinates in the at least one memory blocks;

expanding by a data expander the original matrix using an im2col operation to obtain the expansion matrix;

selecting by a data multiplexer data required for the convolution calculation in the expansion matrix according to the graphics processor;

combining by an output merge buffer the data in the expansion matrix selected by the data multiplexer, and outputting the combined data selected by the data multiplexer to a register file.

9. The method as claimed in claim 8 , wherein the step of executing the im2col algorithm further comprises storing by the register file the data in the original matrix, data in the expansion matrix, and data in the feature map matrix in the convolution calculation.

10. The method as claimed in claim 9 , wherein the step of executing the im2col algorithm further comprises executing by the graphics processor the convolution calculation according to the data in the original matrix, the data in the expansion matrix, and the data in the feature map matrix in the register file.

11. The method as claimed in claim 8 , wherein the step of executing the im2col algorithm further comprises reading and storing by an L1 cache the original matrix for the convolution calculation from the memory for the graphics processor to access in the convolution calculation.

12. The method as claimed in claim 8 , wherein the step of executing the im2col algorithm further comprises storing by a second memory a result of the convolution calculation executed by the graphics processor in the memory.

13. The method as claimed in claim 8 , wherein a size of each memory block included in the memory is a matrix size of 4*8.

14. The method as claimed in claim 8 , wherein the size of the kernel is a matrix size of 3*3, the stride is equal to 1, and the padding is equal to 0.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2021
From: ZHANG, WEI; ZHAI, XINGANG
To: SHANGHAI ZHAOXIN SEMICONDUCTOR CO., LTD.
Reel/Frame 055901/0157 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2021
From: SHANGHAI ZHAOXIN SEMICONDUCTOR CO., LTD.
To: GLENFLY TECHNOLOGY CO., LTD
Reel/Frame 055901/0214 →
Priority Claims (1)
CN 202011048270.1 · Sep 29, 2020 · national
Continuity (1)
Related Publication 20220100814A1 · Mar 31, 2022
References Cited (9)
US 10255547B2 · Woolley, Jr. · 2019 [cited by examiner]
US 11308574B2 · Ould-Ahmed-Vall et al. · 2022 [cited by applicant]
US 20210390367A1 · Liu · 2021 [cited by examiner]
CN 100583162C · 2010 [cited by applicant]
CN 106919942A · 2017 [cited by applicant]
CN 107341761A · 2017 [cited by applicant]
TW 201935408A · 2019 [cited by applicant]
WO 2005124692A1 · 2005 [cited by applicant]
Chinese language Notice of Allowance dated Aug. 30, 2022, issued in application No. TW 110103971. [cited by applicant]