IP Library Granted Patent US 12,073,259
Granted Patent B2
US 12,073,259 · App. 17/048,457 · Granted Aug 27, 2024

Apparatus and method for efficient parallel computation

Inventors: Bernhard Frohwitter (Munich, DE); Thomas Lippert (Aschaffenburg, DE)
Assignee: PARTEC AG
G06F9/5094G06F1/329G06F9/4893G06F9/5027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,073,259
App. No.
17/048,457
Granted
Aug 27, 2024
Kind
B2
Abstract

The present invention provides a computing unit for operating in a parallel computing system, the computing unit comprising a plurality of processing elements and an interface for connecting the computing unit to other components of the computing system wherein each processing element has a nominal maximum processing rate NPR and each processing element includes a respective memory unit such that data can be transferred from the memory unit at a predetermined maximum data rate MBW and the interface provides a maximum data transfer rate CBW, wherein in order to provide a predetermined peak calculation performance for the computing unit PP obtainable by a number n processing elements operating at the nominal maximum processing rate such that PP=n×NPR operations per second, the computing unit includes an integer multiple f times n processing elements wherein f is greater than one and each processing element is limited to operate at a processing rate of NPR/f.

Claims (28)

1. A computing unit for operating in a parallel computing system, the computing unit comprising:

a plurality of processing elements, wherein each of the plurality of processing elements comprises a memory unit and has a nominal maximum processing rate (NPR); and

an interface for connecting the computing unit to other components of the parallel computing system;

wherein the computing unit is provided with a predetermined peak calculation performance (PP) that is the product of the NPR and an amount of the plurality of processing elements operating at the NPR,

wherein the amount of the plurality of processing elements is increased based on an integer (f) greater than one,

wherein each of the plurality of processing elements included in the amount that was increased is limited to operate at a processing rate equal to NPR/f,

wherein for each of the plurality of processing elements, data is transferred from the respective memory unit at a predetermined maximum data rate (MBW), and

wherein f is selected such that a predetermined ratio R=MBW/(NPR/f) is obtained.

2. The computing unit according to claim 1 , wherein R is in a range between 0.15 byte/flop and 1 byte/flop.

3. The computing unit according to claim 1 , wherein the MBW is within 30% of a processing element to processing element communication rate.

4. The computing unit according to claim 1 , wherein the plurality of processing elements are graphics processing units.

5. The computing unit according to claim 1 , wherein the plurality of processing elements are connected together by an interface unit, and wherein each of the plurality of processing elements is connected to the interface unit by a plurality S of serial data lanes.

6. The computing unit according to claim 1 , wherein the computing unit is a computer card comprising a multi-core processor arranged to control the plurality of processing elements.

7. The computing unit according to claim 1 , wherein the plurality of processing elements are GPUs.

8. A method of configuring a computing unit for operating in a parallel computing system, the method comprising:

configuring a plurality of processing elements, wherein each of the plurality of processing elements comprises a memory unit and has a nominal maximum processing rate (NPR); and

configuring an interface for connecting the computing unit to other components of the parallel computing system;

wherein the computing unit is provided with a predetermined peak calculation performance (PP) that is the product of the NPR and an amount of the plurality of processing elements operating at the NPR,

wherein the amount of the plurality of processing elements is increased based on an integer (f) greater than one, and

wherein each of the plurality of processing elements included in the amount that was increased is limited to operate at a processing rate equal to NPR/f,

wherein for each of the plurality of processing elements, data is transferred from the respective memory unit at a predetermined maximum data rate (MBW), and

wherein f is selected such that a predetermined ratio R=MBW/(NPR/f) is obtained.

9. The method according to claim 8 , wherein R is in a range between 0.15 byte/flop and 1 byte/flop.

10. The method according to claim 8 , wherein the MBW is within 30% of a processing element to processing element communication rate.

11. The method according to claim 8 , wherein the plurality of processing elements are graphics processing units.

12. The method according to claim 8 , wherein the plurality of processing elements are connected together by an interface unit, and wherein each of the plurality of processing elements is connected to the interface unit by a plurality S of serial data lanes.

13. The method according to claim 8 , wherein the computing unit is a computer card comprising a multi-core processor arranged to control the plurality of processing elements.

14. The method according to claim 8 , wherein the plurality of processing elements are GPUs.

Assignments (2)
CERTIFICATE OF CONVERSION Recorded Jun 7, 2024
From: PARTEC CLUSTER COMPETENCE CENTER GMBH
To: PARTEC AG
Reel/Frame 067663/0874 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2021
From: FROHWITTER, BERNHARD; LIPPERT, THOMAS
To: PARTEC CLUSTER COMPETENCE CENTER GMBH
Reel/Frame 057967/0463 →
Priority Claims (1)
EP 18172497 · May 15, 2018 · regional
Continuity (1)
Related Publication 20210157656A1 · May 27, 2021