IP Library Granted Patent US 12,293,187
Granted Patent B2
US 12,293,187 · App. 18/524,942 · Granted May 6, 2025

Efficient processing of nested loops for computing device with multiple configurable processing elements using multiple spoke counts

Inventors: Douglas Vanesko (Dallas, TX); Tony M. Brewer (Plano, TX)
Assignee: Micron Technology, Inc.
G06F9/325G06F9/3867G06F9/30065
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,293,187
App. No.
18/524,942
Granted
May 6, 2025
Kind
B2
Abstract

Disclosed in some examples, are methods, systems, devices, and machine-readable mediums which provide for more efficient CGRA execution by assigning different initiation intervals to different PEs executing a same code base. The initiation intervals may be a multiple of each other and the PE with the lowest initiation interval may be used to execute instructions of the code that is to be executed at a greater frequency than other instructions than other instructions that may be assigned to PEs with higher initiation intervals.

Claims (34)

1. A method comprising:

identifying a nested loop in a set of instructions;

configuring a first initiation interval of a first processing element of a set of interconnected processing elements to a first value and a second initiation interval of a second processing element of the set of interconnected processing elements to a second value, the second value a multiple of the first value, the first and second initiation intervals specifying a number of consecutive instructions allowed within a processing pipeline of each respective processing element;

assigning instructions of an inner loop of the nested loop to the first processing element and instructions of an outer loop of the nested loop to the second processing element; and

causing execution of the set of instructions by the first and second processing elements, the causing execution comprising encoding an indication of the configured first initiation interval and second initiation interval and the assignment of instructions into machine code representing the set of instructions or into metadata included along with the machine code.

2. The method of claim 1 , wherein the set of interconnected processing elements is a coarse grained reconfigurable array (CGRA) of a compute-near-memory system comprising a hybrid threading fabric (HTF).

3. The method of claim 1 , wherein assigning instructions of the inner loop comprises assigning at least one same instruction to at least two instruction slots of the first processing element.

4. The method of claim 3 , wherein the at least two instruction slots are selected based on an instruction slot of a preceding instruction in the nested loop and the multiple of the second value over the first value.

5. The method of claim 1 , wherein determining the first initiation interval and the second initiation interval is based on a number of instructions in the inner loop and outer loop and a number of processing elements.

6. The method of claim 1 , wherein causing execution of the instructions comprises configuring the first and second processing elements, loading the instructions according to the assignments, and initiating execution by a dispatch interface.

7. The method of claim 1 , wherein the first processing element is assigned instructions of the inner loop that are executed at a higher frequency than instructions of the outer loop assigned to the second processing element.

8. A non-transitory machine-readable medium, storing instructions, which when executed by a machine, cause the machine to perform operations comprising:

identifying a nested loop in a set of instructions;

configuring a first initiation interval of a first processing element of a set of interconnected processing elements to a first value and a second initiation interval of a second processing element of the set of interconnected processing elements to a second value, the second value a multiple of the first value, the first and second initiation intervals specifying a number of consecutive instructions allowed within a processing pipeline of each respective processing element;

assigning instructions of an inner loop of the nested loop to the first processing element and instructions of an outer loop of the nested loop to the second processing element; and

causing execution of the set of instructions by the first and second processing elements, the causing execution comprising encoding an indication of the configured first initiation interval and second initiation interval and the assignment of instructions into machine code representing the set of instructions or into metadata included along with the machine code.

9. The non-transitory machine-readable medium of claim 8 , wherein the set of interconnected processing elements is a coarse grained reconfigurable array (CGRA) of a compute-near-memory system comprising a hybrid threading fabric (HTF).

10. The non-transitory machine-readable medium of claim 8 , wherein the operations of assigning instructions of the inner loop further comprise assigning at least one same instruction to at least two instruction slots of the first processing element.

11. The non-transitory machine-readable medium of claim 10 , wherein the at least two instruction slots are selected based on an instruction slot of a preceding instruction in the nested loop and the multiple of the second value over the first value.

12. The non-transitory machine-readable medium of claim 8 , wherein the operations further comprise determining the first initiation interval and the second initiation interval is based on a number of instructions in the inner loop and outer loop and a number of processing elements.

13. The non-transitory machine-readable medium of claim 8 , wherein the operations of causing execution of the instructions further comprise configuring the first and second processing elements, loading the instructions according to the assignments, and initiating execution by a dispatch interface.

14. The non-transitory machine-readable medium of claim 8 , wherein the first processing element is assigned instructions of the inner loop that are executed at a higher frequency than instructions of the outer loop assigned to the second processing element.

15. A computing device comprising:

a processor;

a memory, storing instructions which when performed by the processor, cause the processor to perform operations comprising:

identifying a nested loop in a set of instructions;

configuring a first initiation interval of a first processing element of a set of interconnected processing elements to a first value and a second initiation interval of a second processing element of the set of interconnected processing elements to a second value, the second value a multiple of the first value, the first and second initiation intervals specifying a number of consecutive instructions allowed within a processing pipeline of each respective processing element;

assigning instructions of an inner loop of the nested loop to the first processing element and instructions of an outer loop of the nested loop to the second processing element; and

causing execution of the set of instructions by the first and second processing elements, the causing execution comprising encoding an indication of the configured first initiation interval and second initiation interval and the assignment of instructions into machine code representing the set of instructions or into metadata included along with the machine code.

16. The computing device of claim 15 , wherein the set of interconnected processing elements is a coarse grained reconfigurable array (CGRA) of a compute-near-memory system comprising a hybrid threading fabric (HTF).

17. The computing device of claim 15 , wherein the operations of assigning instructions of the inner loop further comprise assigning at least one same instruction to at least two instruction slots of the first processing element.

18. The computing device of claim 17 , wherein the at least two instruction slots are selected based on an instruction slot of a preceding instruction in the nested loop and the multiple of the second value over the first value.

19. The computing device of claim 15 , wherein the operations further comprise determining the first initiation interval and the second initiation interval is based on a number of instructions in the inner loop and outer loop and a number of processing elements.

20. The computing device of claim 15 , wherein the operations of causing execution of the instructions further comprise configuring the first and second processing elements, loading the instructions according to the assignments, and initiating execution by a dispatch interface.

Continuity (2)
Continuation 17399801 · Aug 11, 2021
Related Publication 20240111538A1 · Apr 4, 2024
References Cited (55)
US 8122229B2 · Wallach et al. · 2012 [cited by applicant]
US 8156307B2 · Wallach et al. · 2012 [cited by applicant]
US 8205066B2 · Brewer et al. · 2012 [cited by applicant]
US 8423745B1 · Brewer · 2013 [cited by applicant]
US 8561037B2 · Brewer et al. · 2013 [cited by applicant]
US 9239712B2 · Rong et al. · 2016 [cited by applicant]
US 9710384B2 · Wallach et al. · 2017 [cited by applicant]
US 10990391B2 · Brewer · 2021 [cited by applicant]
US 10990392B2 · Brewer · 2021 [cited by applicant]
US 20080270708A1 · Warner et al. · 2008 [cited by applicant]
US 20120079177A1 · Brewer et al. · 2012 [cited by applicant]
US 20130332711A1 · Leidel et al. · 2013 [cited by applicant]
US 20150100950A1 · Ahn et al. · 2015 [cited by applicant]
US 20150106603A1 · Ahn et al. · 2015 [cited by applicant]
US 20150143350A1 · Brewer · 2015 [cited by applicant]
US 20150206561A1 · Brewer et al. · 2015 [cited by applicant]
US 20190042214A1 · Brewer · 2019 [cited by applicant]
US 20190171604A1 · Brewer · 2019 [cited by applicant]
US 20190220426A1 · Jayasena et al. · 2019 [cited by applicant]
US 20190243700A1 · Brewer · 2019 [cited by applicant]
US 20190303154A1 · Brewer · 2019 [cited by applicant]
US 20190324928A1 · Brewer · 2019 [cited by applicant]
US 20190340019A1 · Brewer · 2019 [cited by applicant]
US 20190340020A1 · Brewer · 2019 [cited by applicant]
US 20190340023A1 · Brewer · 2019 [cited by applicant]
US 20190340024A1 · Brewer · 2019 [cited by applicant]
US 20190340027A1 · Brewer · 2019 [cited by applicant]
US 20190340035A1 · Brewer · 2019 [cited by applicant]
US 20190340154A1 · Brewer · 2019 [cited by applicant]
US 20190340155A1 · Brewer · 2019 [cited by applicant]
US 20200065098A1 · Parandeh Afshar et al. · 2020 [cited by applicant]
US 20200133672A1 · Balasubramanian et al. · 2020 [cited by applicant]
US 20200371800A1 · Chirca et al. · 2020 [cited by applicant]
US 20210055964A1 · Brewer · 2021 [cited by applicant]
US 20210064374A1 · Brewer · 2021 [cited by applicant]
US 20210064435A1 · Brewer · 2021 [cited by applicant]
US 20210149600A1 · Brewer · 2021 [cited by applicant]
US 20210232422A1 · Amiri et al. · 2021 [cited by applicant]
US 20230051544A1 · Vanesko et al. · 2023 [cited by applicant]
WO WO2010051167A1 · 2010 [cited by applicant]
WO WO2013184380A2 · 2013 [cited by applicant]
WO WO2019191740A1 · 2019 [cited by applicant]
WO WO2019191742A1 · 2019 [cited by applicant]
WO WO2019191744A1 · 2019 [cited by applicant]
WO WO2019217287A1 · 2019 [cited by applicant]
WO WO2019217295A1 · 2019 [cited by applicant]
WO WO2019217324A1 · 2019 [cited by applicant]
WO WO2019217326A1 · 2019 [cited by applicant]
WO WO2019217329A1 · 2019 [cited by applicant]
WO WO2019089816A3 · 2020 [cited by applicant]
WO WO2023019052A1 · 2023 [cited by applicant]
“International Application Serial No. PCT/US2022/073863, International Search Report mailed Nov. 11, 2022”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT/US2022/073863, Written Opinion mailed Nov. 11, 2022”, 5 pgs. [cited by applicant]
Kieron, Turkington, et al., “Outer Loop Pipelining for Application Specific Datapaths in FPGAs”, IEEE, (2008), 1268-1280. [cited by applicant]
Shouyi, Yin, et al., “Exploiting Parallelism of Imperfect Nested Loops on Coarse-Grained Reconfigurable Architectures”, IEEE, (2016), 3199-3213. [cited by applicant]