IP Library Granted Patent US 12,730,668
Granted Patent B2
US 12,730,668 · App. 17/963,226 · Granted Sep 8, 2026

Load latency amelioration using bunch buffers

Inventor: Peter Foley (Los Altos Hills, CA)
Assignee: Ascenium, Inc.
G06F9/4843G06F9/544
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,668
App. No.
17/963,226
Granted
Sep 8, 2026
Kind
B2
Abstract

Techniques for task processing based on load latency amelioration using bunch buffers are disclosed. A two-dimensional array of compute elements is accessed. Each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements. Control for the compute elements is provided on a cycle-by-cycle basis. The control is enabled by a stream of wide control words generated by the compiler. Sets of control word bits are loaded into buffers. Each buffer is associated with and coupled to a unique compute element within the array of compute elements. The sets of control word bits provide operational control for the compute element with which it is associated. Operations are executed within the array of elements. The operations are based on a selected set of control word bits which comprise a control word bunch.

Claims (39)

1 . A processor-implemented method for task processing comprising:

accessing a two-dimensional array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements;

providing control for the array of compute elements on a cycle-by-cycle basis, wherein the control is enabled by a stream of wide control words generated by the compiler, wherein the control words that are provided enable parallel execution of tasks, wherein the tasks include two or more tasks that are independent of one another;

loading sets of control word bits into buffers, wherein each buffer is associated with and coupled to a unique compute element within the array of compute elements, and wherein the sets of control word bits provide operational control for the compute element with which it is associated; and

executing operations within the array of compute elements, wherein the operations are based on a selected set of control word bits, wherein the selected set of control word bits comprise a control word bunch.

2 . The method of claim 1 wherein control word bunch enables operational control of a particular compute element for a plurality of cycles.

3 . The method of claim 1 further comprising coupling an iteration counter to each buffer.

4 . The method of claim 3 wherein the iteration counter tracks cycling through the sets of control word bits in the buffer coupled to the iteration counter.

5 . The method of claim 3 further comprising using a pre-stored value in an iteration counter to control operation completion.

6 . The method of claim 3 further comprising generating a task completion signal, based on an iteration counter value.

7 . The method of claim 1 further comprising coupling a pointer register to each buffer.

8 . The method of claim 7 wherein the pointer register indicates a next set of control word bits in a buffer to be executed.

9 . The method of claim 7 wherein the pointer enables operation looping within the compute elements.

10 . The method of claim 9 wherein the operation looping is enabled without additional control word loading.

11 . The method of claim 9 wherein the operation looping accomplishes dataflow processing within statically scheduled compute elements.

12 . The method of claim 1 wherein the stream of wide control words includes two or more data dependent branch operations.

13 . The method of claim 12 wherein the two or more data dependent branch operations require a balanced number of execution cycles.

14 . The method of claim 13 wherein the balanced number of execution cycles is determined by the compiler.

15 . The method of claim 1 further comprising executing a memory operation outside of the array of compute elements.

16 . The method of claim 15 wherein the memory operation is enabled by autonomous compute element operation.

17 . The method of claim 16 wherein the autonomous compute element operation is controlled by one or more sets of control word bits.

18 . The method of claim 1 wherein each buffer enables storing sixteen control word bunches.

19 . The method of claim 1 wherein the buffers comprise operation buffers.

20 . The method of claim 1 wherein the accessing, the providing, the loading, and the executing enable background memory accesses.

21 . The method of claim 20 wherein the background memory accesses reduce load latency.

22 . The method of claim 1 wherein the selected set of control word bits is selected from the stream of control words.

23 . The method of claim 1 wherein the selected set of control word bits is loaded into a bunch buffer.

24 . A computer program product embodied in a non-transitory computer readable medium for task processing, the computer program product comprising code which causes one or more processors to perform operations of:

accessing a two-dimensional array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements;

providing control for the array of compute elements on a cycle-by-cycle basis, wherein the control is enabled by a stream of wide control words generated by the compiler, wherein the control words that are provided enable parallel execution of tasks, wherein the tasks include two or more tasks that are independent of one another;

loading sets of control word bits into buffers, wherein each buffer is associated with and coupled to a unique compute element within the array of compute elements, and wherein the sets of control word bits provide operational control for the compute element with which it is associated; and

executing operations within the array of compute elements, wherein the operations are based on a selected set of control word bits, wherein the selected set of control word bits comprise a control word bunch.

25 . A computer system for task processing comprising:

a memory which stores instructions;

one or more processors coupled to the memory, wherein the one or more processors, when executing the instructions which are stored, are configured to:

access a two-dimensional array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements;

provide control for the array of compute elements on a cycle-by-cycle basis, wherein the control is enabled by a stream of wide control words generated by the compiler, wherein the control words that are provided enable parallel execution of tasks, wherein the tasks include two or more tasks that are independent of one another;

load sets of control word bits into buffers, wherein each buffer is associated with and coupled to a unique compute element within the array of compute elements, and wherein the sets of control word bits provide operational control for the compute element with which it is associated; and

execute operations within the array of compute elements, wherein the operations are based on a selected set of control word bits, wherein the selected set of control word bits comprise a control word bunch.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2024
From: FOLEY, PETER
To: ASCENIUM, INC.
Reel/Frame 066554/0811 →
Continuity (21)
Continuation In Part 17526003 · Nov 15, 2021
Continuation In Part 17465949 · Sep 3, 2021
Provisional Application 63402490 · Aug 31, 2022
Provisional Application 63400087 · Aug 23, 2022
Provisional Application 63393989 · Aug 1, 2022
Provisional Application 63388268 · Jul 12, 2022
Provisional Application 63357030 · Jun 30, 2022
Provisional Application 63340499 · May 11, 2022
Provisional Application 63322245 · Mar 22, 2022
Provisional Application 63318413 · Mar 10, 2022
Provisional Application 63295544 · Dec 31, 2021
Provisional Application 63254557 · Oct 12, 2021
Provisional Application 63232230 · Aug 12, 2021
Provisional Application 63229466 · Aug 4, 2021
Provisional Application 63193522 · May 26, 2021
Provisional Application 63166298 · Mar 26, 2021
Provisional Application 63125994 · Dec 16, 2020
Provisional Application 63114003 · Nov 16, 2020
Provisional Application 63091947 · Oct 15, 2020
Provisional Application 63075849 · Sep 9, 2020
Related Publication 20230031902A1 · Feb 2, 2023
References Cited (65)
US 5594884A · Matoba et al. · 1997 [cited by applicant]
US 5764994A · Craft · 1998 [cited by applicant]
US 7840777B2 · Mykland · 2010 [cited by applicant]
US 8694978B1 · Rus et al. · 2014 [cited by applicant]
US 8856768B2 · Mykland · 2014 [cited by applicant]
US 8869123B2 · Mykland · 2014 [cited by applicant]
US 8949806B1 · Lee et al. · 2015 [cited by applicant]
US 9158544B2 · Mykland · 2015 [cited by applicant]
US 9275002B2 · Manet et al. · 2016 [cited by applicant]
US 9304770B2 · Mykland · 2016 [cited by applicant]
US 9395992B2 · Doing et al. · 2016 [cited by applicant]
US 9424055B2 · Lundvall et al. · 2016 [cited by applicant]
US 9473155B2 · Staszewski et al. · 2016 [cited by applicant]
US 9477470B2 · Mykland · 2016 [cited by applicant]
US 9529715B2 · Kumar et al. · 2016 [cited by applicant]
US 9582277B2 · Muff et al. · 2017 [cited by applicant]
US 9594559B2 · Fontenot et al. · 2017 [cited by applicant]
US 9600287B2 · Gschwind et al. · 2017 [cited by applicant]
US 9633160B2 · Mykland · 2017 [cited by applicant]
US 9652238B2 · Muff et al. · 2017 [cited by applicant]
US 9684511B2 · Shanbhogue et al. · 2017 [cited by applicant]
US 9830164B2 · Yazdani · 2017 [cited by applicant]
US 9851969B2 · Greiner et al. · 2017 [cited by applicant]
US 9886277B2 · Loktyukhin et al. · 2018 [cited by applicant]
US 9898293B2 · Whittaker · 2018 [cited by applicant]
US 9921836B2 · Kumar et al. · 2018 [cited by applicant]
US 9928062B2 · Azagury et al. · 2018 [cited by applicant]
US 9934040B2 · Bonanno et al. · 2018 [cited by applicant]
US 9971605B2 · Henry et al. · 2018 [cited by applicant]
US 9977674B2 · Rupley, II et al. · 2018 [cited by applicant]
US 9977675B2 · Nystad · 2018 [cited by applicant]
US 9977679B2 · Caulfield et al. · 2018 [cited by applicant]
US 9983882B2 · Greiner et al. · 2018 [cited by applicant]
US 9983884B2 · Maiyuran et al. · 2018 [cited by applicant]
US 10089277B2 · Mykland · 2018 [cited by applicant]
US 10592250B1 · Diamant · 2020 [cited by examiner]
US 10861126B1 · Sharma · 2020 [cited by examiner]
US 20020174318A1 · Studdard et al. · 2002 [cited by applicant]
US 20040111710A1 · Chakradhar et al. · 2004 [cited by applicant]
US 20130326190A1 · Chung et al. · 2013 [cited by applicant]
US 20140040598A1 · Fleischer et al. · 2014 [cited by applicant]
US 20140149657A1 · Jakovljevic et al. · 2014 [cited by applicant]
US 20140317383A1 · Park et al. · 2014 [cited by applicant]
US 20150186146A1 · Kushida et al. · 2015 [cited by applicant]
US 20160246602A1 · Radhika et al. · 2016 [cited by applicant]
US 20160313984A1 · Meixner · 2016 [cited by examiner]
US 20180225116A1 · Henry et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180322606A1 · Das et al. · 2018 [cited by applicant]
US 20180341493A1 · Roy et al. · 2018 [cited by applicant]
US 20190004777A1 · Meixner · 2019 [cited by applicant]
US 20190347190A1 · Singh · 2019 [cited by applicant]
US 20190369990A1 · Doerr et al. · 2019 [cited by applicant]
US 20200026498A1 · Sumbul et al. · 2020 [cited by applicant]
US 20200226058A1 · Lee · 2020 [cited by examiner]
US 20200241879A1 · Vorbach et al. · 2020 [cited by applicant]
US 20210026637A1 · Vorbach · 2021 [cited by applicant]
US 20220413850A1 · Gately · 2022 [cited by examiner]
US 20230289182A1 · Hussain · 2023 [cited by examiner]
KR 1020150051083A · 2015 [cited by applicant]
WO WO2011038940A1 · 2011 [cited by applicant]
Musicus, B. R. (1988). The OKI advanced array processor (AAP): Development Software Manual. [cited by applicant]
Podobas, Artur, Kentaro Sano, and Satoshi Matsuoka. “A survey on coarse-grained reconfigurable architectures from a performance perspective.” IEEE Access 8 (2020): 146719-146743. [cited by applicant]
International Search Report dated Feb. 9, 2023 for PCT/US2022/046210. [cited by applicant]
Chang, Kyungwook, and Kiyoung Choi. “Mapping control intensive kernels onto coarse-grained reconfigurable array architecture.” 2008 International SoC Design Conference. vol. 1. IEEE, 2008. [cited by applicant]