IP Library Granted Patent US 12,705,092
Granted Patent B2
US 12,705,092 · App. 18/119,315 · Granted Aug 11, 2026

Highly parallel processing architecture with out-of-order resolution

Inventor: Peter Foley (Los Altos Hills, CA)
Assignee: Ascenium, Inc.
G06F9/4881G06F8/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,092
App. No.
18/119,315
Filed
Mar 9, 2023
Granted
Aug 11, 2026
Kind
B2
Art Unit
2196
USPC
718/102
Abstract

Techniques for task processing based on a highly parallel processing architecture with out-of-order resolution are disclosed. A two-dimensional array of compute elements is accessed. Each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements. The array of compute elements is coupled to supporting logic and to memory, which, along with the array of compute elements, comprise compute hardware. A set of directions is provided to the hardware, through a control word generated by the compiler, for compute element operation. The set of directions is augmented with data access ordering information. The data access ordering is performed by the hardware. A compiled task is executed on the array of compute elements, based on the set of directions that was augmented.

Claims (36)

1 . A processor-implemented method for task processing comprising:

accessing a two-dimensional array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements, and wherein the array of compute elements is coupled to supporting logic and to memory, which together with the array of compute elements comprises compute hardware;

providing a set of directions to the hardware, through a control word generated by the compiler, for compute element operation;

augmenting the set of directions with data access ordering information, wherein data access ordering is performed by the compute hardware, wherein the ordering information includes ordering information for a single architectural cycle, wherein the single architectural cycle contains multiple compute element operations; and

executing a compiled task on the array of compute elements, based on the set of directions that was augmented.

2 . The method of claim 1 wherein the ordering information includes ordering information for load and/or store operations.

3 . The method of claim 2 wherein the load and/or store operations read and/or write data to the memory.

4 . The method of claim 2 wherein the load and/or store ordering information enables the hardware to detect data access hazards.

5 . The method of claim 4 wherein the data access hazards include write-after-read, read-after-write, and write-after-write conflicts.

6 . The method of claim 4 further comprising resolving a data access hazard that was detected.

7 . The method of claim 6 wherein the resolving includes delaying loads and/or stores.

8 . The method of claim 7 wherein data for the load and/or store is held in buffers.

9 . The method of claim 8 wherein the data held in buffers is committed after the data access hazard detection and mitigation window has expired.

10 . The method of claim 2 wherein the load and/or store operations involve a temporal distance of more than one architectural cycle.

11 . The method of claim 10 further comprising using local buffers to delay commitment of data for the load and/or store operations.

12 . The method of claim 11 wherein the data for load operations is read from the memory.

13 . The method of claim 11 wherein the data for store operations is written to the memory.

14 . The method of claim 1 wherein the compute hardware ensures semantic correctness of operations to the memory.

15 . The method of claim 1 wherein the data access ordering enables ordering of memory data.

16 . The method of claim 15 wherein the ordering of memory data enables compute element result sequencing.

17 . The method of claim 1 wherein the control word specifies operations for the array of compute elements.

18 . The method of claim 17 wherein the operations are specified on a physical cycle-by-cycle basis.

19 . The method of claim 18 wherein the physical cycle-by-cycle basis comprises an architectural cycle.

20 . The method of claim 1 wherein the data access ordering information is generated by the hardware during runtime.

21 . A computer program product embodied in a non-transitory computer readable medium for task processing, the computer program product comprising code which causes one or more processors to perform operations of:

accessing a two-dimensional array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements, and wherein the array of compute elements is coupled to supporting logic and to memory, which together with the array of compute elements comprises compute hardware;

providing a set of directions to the hardware, through a control word generated by the compiler, for compute element operation;

augmenting the set of directions with data access ordering information, wherein data access ordering is performed by the compute hardware, wherein the ordering information includes ordering information for a single architectural cycle, wherein the single architectural cycle contains multiple compute element operations; and

executing a compiled task on the array of compute elements, based on the set of directions that was augmented.

22 . A computer system for task processing comprising:

a memory which stores instructions;

one or more processors coupled to the memory, wherein the one or more processors, when executing the instructions which are stored, are configured to:

access a two-dimensional array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements, and wherein the array of compute elements is coupled to supporting logic and to memory, which together with the array of compute elements comprises compute hardware;

provide a set of directions to the hardware, through a control word generated by the compiler, for compute element operation;

augment the set of directions with data access ordering information, wherein data access ordering is performed by the compute hardware, wherein the ordering information includes ordering information for a single architectural cycle, wherein the single architectural cycle contains multiple compute element operations; and

execute a compiled task on the array of compute elements, based on the set of directions that was augmented.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2024
From: FOLEY, PETER
To: ASCENIUM, INC.
Reel/Frame 066555/0029 →
Continuity (24)
Continuation In Part 17526003 · Nov 15, 2021
Continuation In Part 17465949 · Sep 3, 2021
Provisional Application 63447915 · Feb 24, 2023
Provisional Application 63442131 · Jan 31, 2023
Provisional Application 63424960 · Nov 14, 2022
Provisional Application 63424961 · Nov 14, 2022
Provisional Application 63402490 · Aug 31, 2022
Provisional Application 63400087 · Aug 23, 2022
Provisional Application 63393989 · Aug 1, 2022
Provisional Application 63388268 · Jul 12, 2022
Provisional Application 63357030 · Jun 30, 2022
Provisional Application 63340499 · May 11, 2022
Provisional Application 63322245 · Mar 22, 2022
Provisional Application 63318413 · Mar 10, 2022
Provisional Application 63254557 · Oct 12, 2021
Provisional Application 63232230 · Aug 12, 2021
Provisional Application 63229466 · Aug 4, 2021
Provisional Application 63193522 · May 26, 2021
Provisional Application 63166298 · Mar 26, 2021
Provisional Application 63125994 · Dec 16, 2020
Provisional Application 63114003 · Nov 16, 2020
Provisional Application 63091947 · Oct 15, 2020
Provisional Application 63075849 · Sep 9, 2020
Related Publication 20230273818A1 · Aug 31, 2023
References Cited (63)
US 5594884A · Matoba et al. · 1997 [cited by applicant]
US 5764994A · Craft · 1998 [cited by applicant]
US 7840777B2 · Mykland · 2010 [cited by applicant]
US 8694978B1 · Rus et al. · 2014 [cited by applicant]
US 8856768B2 · Mykland · 2014 [cited by applicant]
US 8869123B2 · Mykland · 2014 [cited by applicant]
US 8949806B1 · Lee et al. · 2015 [cited by applicant]
US 9158544B2 · Mykland · 2015 [cited by applicant]
US 9304770B2 · Mykland · 2016 [cited by applicant]
US 9395992B2 · Doing et al. · 2016 [cited by applicant]
US 9424055B2 · Lundvall et al. · 2016 [cited by applicant]
US 9473155B2 · Staszewski et al. · 2016 [cited by applicant]
US 9477470B2 · Mykland · 2016 [cited by applicant]
US 9529715B2 · Kumar et al. · 2016 [cited by applicant]
US 9582277B2 · Muff et al. · 2017 [cited by applicant]
US 9594559B2 · Fontenot et al. · 2017 [cited by applicant]
US 9600287B2 · Gschwind et al. · 2017 [cited by applicant]
US 9633160B2 · Mykland · 2017 [cited by applicant]
US 9652238B2 · Muff et al. · 2017 [cited by applicant]
US 9684511B2 · Shanbhogue et al. · 2017 [cited by applicant]
US 9830164B2 · Yazdani · 2017 [cited by applicant]
US 9851969B2 · Greiner et al. · 2017 [cited by applicant]
US 9886277B2 · Loktyukhin et al. · 2018 [cited by applicant]
US 9898293B2 · Whittaker · 2018 [cited by applicant]
US 9921836B2 · Kumar et al. · 2018 [cited by applicant]
US 9928062B2 · Azagury et al. · 2018 [cited by applicant]
US 9934040B2 · Bonanno et al. · 2018 [cited by applicant]
US 9971605B2 · Henry et al. · 2018 [cited by applicant]
US 9977674B2 · Rupley, II et al. · 2018 [cited by applicant]
US 9977675B2 · Nystad · 2018 [cited by applicant]
US 9977679B2 · Caulfield et al. · 2018 [cited by applicant]
US 9983882B2 · Greiner et al. · 2018 [cited by applicant]
US 9983884B2 · Maiyuran et al. · 2018 [cited by applicant]
US 10089277B2 · Mykland · 2018 [cited by applicant]
US 20020174318A1 · Studdard et al. · 2002 [cited by applicant]
US 20040111710A1 · Chakradhar et al. · 2004 [cited by applicant]
US 20130304996A1 · Venkataraman et al. · 2013 [cited by applicant]
US 20130326190A1 · Chung et al. · 2013 [cited by applicant]
US 20140149657A1 · Jakovljevic et al. · 2014 [cited by applicant]
US 20140317383A1 · Park et al. · 2014 [cited by applicant]
US 20140372730A1 · Qadri et al. · 2014 [cited by applicant]
US 20150186146A1 · Kushida et al. · 2015 [cited by applicant]
US 20160246602A1 · Radhika et al. · 2016 [cited by applicant]
US 20160313984A1 · Meixner · 2016 [cited by examiner]
US 20180225116A1 · Henry et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180322606A1 · Das et al. · 2018 [cited by applicant]
US 20180341493A1 · Roy et al. · 2018 [cited by applicant]
US 20190004777A1 · Meixner · 2019 [cited by applicant]
US 20190188824A1 · Meixner et al. · 2019 [cited by applicant]
US 20190347190A1 · Singh · 2019 [cited by applicant]
US 20190369990A1 · Doerr et al. · 2019 [cited by applicant]
US 20200026498A1 · Sumbul et al. · 2020 [cited by applicant]
US 20200192860A1 · Mykland · 2020 [cited by applicant]
US 20200241879A1 · Vorbach et al. · 2020 [cited by applicant]
US 20200310814A1 · Kothinti Naresh · 2020 [cited by examiner]
US 20210026632A1 · Dooley · 2021 [cited by examiner]
US 20220075651A1 · Harboe · 2022 [cited by examiner]
KR 1020150051083A · 2015 [cited by applicant]
WO WO2011038940A1 · 2011 [cited by applicant]
International Search Report dated Jun. 28, 2023 for PCT U.S. Appl. No. 23/014,863. [cited by applicant]
Musicus, B. R. (1988). The OKI advanced array processor (AAP): Development Software Manual. [cited by applicant]
Chang, Kyungwook, and Kiyoung Choi. “Mapping control intensive kernels onto coarse-grained reconfigurable array architecture.” 2008 International SoC Design Conference. vol. 1. IEEE, 2008. [cited by applicant]