IP Library Granted Patent US 11,429,767
Granted Patent B2
US 11,429,767 · App. 16/723,091 · Granted Aug 30, 2022

Accelerator automation framework for heterogeneous computing in datacenters

Inventors: Cody Hao Yu (Los Angeles, CA); Peng Zhang (Los Angeles, CA)
Assignee: XILINX, INC.
G06F30/27G06F9/5061G06F9/544G06F30/34G06N20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,429,767
App. No.
16/723,091
Granted
Aug 30, 2022
Kind
B2
Abstract

Systems and methods for designing an information processing system are described. In one embodiment, a design space is partitioned into a plurality of independent partitions based on a defined set of rules. A unique processing core is assigned to each partition. A plurality of starting points is generated for each partition, where each starting point is associated with a machine learning algorithm. The starting points for each partition may include a performance driven seed and an area-driven seed. A set of feasible designs associated with the information processing system are determined.

Claims (38)

1. A method for designing an information processing system comprising the steps of:

partitioning a design space of possible modifications to a source code into a plurality of independent partitions based on a defined set of rules;

assigning a unique processing core to each partition;

generating a plurality of starting points for each partition, wherein each starting point is associated with a machine learning algorithm; and

determining a set of feasible designs associated with the information processing system;

wherein the plurality of starting points include a performance-driven seed and an area-driven seed; and

wherein the design space comprises a range of values for buffer width, a range of values for loop tiling, a range of values for loop parallelization, and whether or not loop pipelining is enabled.

2. The method of claim 1 , wherein the design space is partitioned based on loop hierarchy associated with the source code.

3. The method of claim 1 , wherein the design space is partitioned based on an outermost kernel loop associated with the source code.

4. The method of claim 1 , wherein the design space is partitioned according to a machine learning model trained to associated design points of the design space to a partition of the plurality of independent partitions according to similarity in at least one of resource utilization and latency.

5. The method of claim 1 , wherein a set of feasible designs is determined using an iterative searching engine.

6. The method of claim 1 , further comprising terminating a process to determine the set of feasible designs based on a predetermined stopping condition.

7. The method of claim 6 , wherein the stopping condition is determined by computing a Shannon entropy associated with the design space.

8. The method of claim 1 , wherein the performance-driven seed is one in which pipelining is enabled for all loops of the source code.

9. The method of claim 1 , wherein the performance-driven seed is one in which parallelization for each loop of the source code is set to a predefined maximum parallelization value.

10. The method of claim 9 , wherein the performance-driven seed is one in which a buffer width is set to a predefined maximum buffer width.

11. The method of claim 10 , wherein the area-driven seed includes all loops of the source code performed sequentially and the buffer width set to a minimum bit width.

12. A system comprising:

one or more processing devices; and

one or more memory devices operably coupled to the one or more processing devices and storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to:

partition a design space of possible modifications to a source code into a plurality of independent partitions based on a defined set of rules;

assign a unique processing core to each partition;

generate a plurality of starting points for each partition, wherein each starting point is associated with a machine learning algorithm; and

determine a set of feasible designs for an information processing system based on the source code;

wherein the plurality of starting points for each partition include a performance-driven seed and an area-driven seed;

wherein the design space comprises a range of values for buffer width, a range of values for loop tiling, a range of values for loop parallelization, and whether or not loop pipelining is enabled;

wherein the performance-driven seed is one in which pipelining is enabled for all loops of the source code, parallelization for each loop of the source code is set to a predefined maximum parallelization value, and a buffer width is set to a predefined maximum buffer width.

13. The system of claim 12 , wherein the executable code, when executed by the one or more processing devices, further causes the one or more processing devices to partition the design space based on loop hierarchy associated with the source code.

14. The system of claim 12 , wherein the executable code, when executed by the one or more processing devices, further causes the one or more processing devices to partition the design space based on an outermost kernel loop associated with the source code.

15. The system of claim 12 , wherein the executable code, when executed by the one or more processing devices, further causes the one or more processing devices to partition the design space according to a machine learning model trained to associated design points of the design space to a partition of the plurality of independent partitions according to similarity in at least one of resource utilization and latency.

16. The system of claim 12 , wherein the executable code, when executed by the one or more processing devices, further causes the one or more processing devices to determine the set of feasible designs using an iterative searching engine.

17. The system of claim 12 , wherein the area-driven seed includes all loops of the source code performed sequentially and the buffer width set to a minimum bit width.

18. A non-transitory computer-readable medium storing executable code that, when executed by one or more processing devices, causes the one or more processing devices to:

partition a design space of possible modifications to a source code into a plurality of independent partitions based on a defined set of rules;

assign a unique processing core to each partition;

generate a plurality of starting points for each partition, wherein each starting point is associated with a machine learning algorithm; and

determine a set of feasible designs for an information processing system based on the source code;

wherein the design space is partitioned based on a loop hierarchy associated with the source code, the loop hierarchy including a plurality of levels of nested loops.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2020
From: FALCON COMPUTING SOLUTIONS, INC.
To: XILINX, INC.
Reel/Frame 054670/0159 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR'S NAME AND ASSIGNEE'S NAME PREVIOUSLY RECORDED ON REEL 051346 FRAME 0744. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT.. Recorded Oct 27, 2020
From: YU, CODY HAO; ZHANG, PENG
To: FALCON COMPUTING SOLUTIONS, INC.
Reel/Frame 054486/0894 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2019
From: YU, CODY HAO; ZHANG, PENG; CONG, JINGSHENG JASON
To: FALCON COMPUTING
Reel/Frame 051346/0744 →
Continuity (2)
Provisional Application 62782875 · Dec 20, 2018
Related Publication 20200202058A1 · Jun 25, 2020