IP Library Granted Patent US 12,493,554
Granted Patent B2
US 12,493,554 · App. 18/387,695 · Granted Dec 9, 2025

Parallel processing using hazard detection and mitigation

Inventor: Peter Foley (Los Altos Hills, CA)
Assignee: Ascenium, Inc.
G06F12/0846G06F9/383
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,554
App. No.
18/387,695
Granted
Dec 9, 2025
Kind
B2
Abstract

Techniques for parallel processing using hazard detection and mitigation are disclosed. An array of compute elements is accessed. Each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements. Control for the compute elements is provided on a cycle-by-cycle basis. Control is enabled by a stream of wide control words generated by the compiler. Memory access operations are tagged with precedence information. The tagging is contained in the control words. The tagging is provided by the compiler at compile time. Memory access operations are monitored. The monitoring is based on the precedence information and a number of architectural cycles of the cycle-by-cycle basis. The tagging is augmented at run time, based on the monitoring. Memory access data is held before promotion, based on the monitoring.

Claims (42)

1 . A processor-implemented method for parallel processing comprising:

accessing an array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements, wherein each compute element is disposed within one or more integrated circuits;

providing control for the compute elements on a cycle-by-cycle basis, wherein the control is enabled by a stream of control words generated by the compiler;

tagging memory access operations with precedence information, wherein the tagging is contained in the control words, wherein the tagging includes tagging store accesses with a unique precedence tag, wherein the unique precedence tag is based on a cycle count that includes a number of cycles within which a store access must occur before an out-of-time exception is thrown, and wherein the tagging is provided by the compiler at compile time;

monitoring the memory access operations, wherein the monitoring is based on the precedence information and a number of architectural cycles of the cycle-by-cycle basis; and

holding memory access data before promotion, based on the monitoring.

2 . The method of claim 1 further comprising augmenting the tagging at run time, based on the monitoring.

3 . The method of claim 1 wherein the holding is accomplished using access buffers coupled to a memory cache.

4 . The method of claim 3 wherein the holding prevents premature data promotion into or out of the memory cache.

5 . The method of claim 3 wherein the memory cache comprises a data cache for the array of compute elements.

6 . The method of claim 3 wherein the memory cache has an access time that is unknown to the compiler.

7 . The method of claim 3 further comprising transferring the memory access data between the array of compute elements and the access buffers using a crossbar switch.

8 . The method of claim 7 wherein a delay for the transferring is unknown to the compiler.

9 . The method of claim 7 wherein the unique precedence tag enables load access priority in the crossbar switch.

10 . The method of claim 1 wherein the precedence information enables hardware ordering of memory access loads to the array of compute elements and memory access stores from the array of compute elements.

11 . The method of claim 10 wherein the precedence information provides semantically correct operation ordering.

12 . The method of claim 1 wherein the precedence information comprises intra-control word precedence and inter-control word precedence.

13 . The method of claim 1 further comprising identifying hazardous loads and stores by comparing load and store addresses to contents of an access buffer.

14 . The method of claim 13 wherein the comparing identifies potential accesses to a same address.

15 . The method of claim 13 further comprising including the precedence information in the comparing.

16 . The method of claim 13 further comprising delaying promoting of data to the access buffer and releasing of the data from the access buffer.

17 . The method of claim 16 wherein the delaying avoids hazards.

18 . The method of claim 17 wherein the avoiding the hazards is based on a comparative precedence value.

19 . The method of claim 17 wherein the hazards include write-after-read conflicts, read-after-write conflicts, and write-after-write conflicts.

20 . The method of claim 13 wherein the identifying enables hazard mitigation.

21 . The method of claim 20 wherein the hazard mitigation includes load-to-store forwarding, store-to-load forwarding, and store-to-store forwarding.

22 . The method of claim 1 wherein the control words include branch operations.

23 . The method of claim 22 further comprising suppressing memory access stores for untaken branch paths.

24 . A computer program product embodied in a non-transitory computer readable medium for parallel processing, the computer program product comprising code which causes one or more processors to perform operations of:

accessing an array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements, wherein each compute element is disposed within one or more integrated circuits;

providing control for the compute elements on a cycle-by-cycle basis, wherein the control is enabled by a stream of control words generated by the compiler;

tagging memory access operations with precedence information, wherein the tagging is contained in the control words, wherein the tagging includes tagging store accesses with a unique precedence tag, wherein the unique precedence tag is based on a cycle count that includes a number of cycles within which a store access must occur before an out-of-time exception is thrown, and wherein the tagging is provided by the compiler at compile time;

monitoring the memory access operations, wherein the monitoring is based on the precedence information and a number of architectural cycles of the cycle-by-cycle basis; and

holding memory access data before promotion, based on the monitoring.

25 . A computer system for parallel processing comprising:

a memory which stores instructions;

one or more processors coupled to the memory, wherein the one or more processors, when executing the instructions which are stored, are configured to:

access an array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements, wherein each compute element is disposed within one or more integrated circuits;

provide control for the compute elements on a cycle-by-cycle basis, wherein the control is enabled by a stream of control words generated by the compiler;

tag memory access operations with precedence information, wherein the tagging is contained in the control words, wherein the tagging includes tagging store accesses with a unique precedence tag, wherein the unique precedence tag is based on a cycle count that includes a number of cycles within which a store access must occur before an out-of-time exception is thrown, and wherein the tagging is provided by the compiler at compile time;

monitor the memory access operations, wherein the monitoring is based on the precedence information and a number of architectural cycles of the cycle-by-cycle basis; and

hold memory access data before promotion, based on the monitoring.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2024
From: FOLEY, PETER
To: ASCENIUM, INC.
Reel/Frame 066802/0865 →
Continuity (19)
Continuation In Part 17526003 · Nov 15, 2021
Continuation In Part 17465949 · Sep 3, 2021
Provisional Application 63536144 · Sep 1, 2023
Provisional Application 63529159 · Jul 27, 2023
Provisional Application 63460909 · Apr 21, 2023
Provisional Application 63447915 · Feb 24, 2023
Provisional Application 63442131 · Jan 31, 2023
Provisional Application 63424960 · Nov 14, 2022
Provisional Application 63424961 · Nov 14, 2022
Provisional Application 63254557 · Oct 12, 2021
Provisional Application 63232230 · Aug 12, 2021
Provisional Application 63229466 · Aug 4, 2021
Provisional Application 63193522 · May 26, 2021
Provisional Application 63166298 · Mar 26, 2021
Provisional Application 63125994 · Dec 16, 2020
Provisional Application 63114003 · Nov 16, 2020
Provisional Application 63091947 · Oct 15, 2020
Provisional Application 63075849 · Sep 9, 2020
Related Publication 20240070076A1 · Feb 29, 2024
References Cited (76)
US 5594884A · Matoba et al. · 1997 [cited by applicant]
US 5764994A · Craft · 1998 [cited by applicant]
US 6988183B1 · Wong · 2006 [cited by examiner]
US 7343477B1 · Thatipelli · 2008 [cited by examiner]
US 7840777B2 · Mykland · 2010 [cited by applicant]
US 8627017B2 · Sheaffer et al. · 2014 [cited by applicant]
US 8677081B1 · Wentzlaff et al. · 2014 [cited by applicant]
US 8694978B1 · Rus et al. · 2014 [cited by applicant]
US 8856768B2 · Mykland · 2014 [cited by applicant]
US 8869123B2 · Mykland · 2014 [cited by applicant]
US 8949806B1 · Lee et al. · 2015 [cited by applicant]
US 9158544B2 · Mykland · 2015 [cited by applicant]
US 9304770B2 · Mykland · 2016 [cited by applicant]
US 9395992B2 · Doing et al. · 2016 [cited by applicant]
US 9424055B2 · Lundvall et al. · 2016 [cited by applicant]
US 9473155B2 · Staszewski et al. · 2016 [cited by applicant]
US 9477470B2 · Mykland · 2016 [cited by applicant]
US 9529715B2 · Kumar et al. · 2016 [cited by applicant]
US 9582277B2 · Muff et al. · 2017 [cited by applicant]
US 9594559B2 · Fontenot et al. · 2017 [cited by applicant]
US 9600287B2 · Gschwind et al. · 2017 [cited by applicant]
US 9633160B2 · Mykland · 2017 [cited by applicant]
US 9652238B2 · Muff et al. · 2017 [cited by applicant]
US 9684511B2 · Shanbhogue et al. · 2017 [cited by applicant]
US 9830164B2 · Yazdani · 2017 [cited by applicant]
US 9851969B2 · Greiner et al. · 2017 [cited by applicant]
US 9886277B2 · Loktyukhin et al. · 2018 [cited by applicant]
US 9898293B2 · Whittaker · 2018 [cited by applicant]
US 9921836B2 · Kumar et al. · 2018 [cited by applicant]
US 9928062B2 · Azagury et al. · 2018 [cited by applicant]
US 9934040B2 · Bonanno et al. · 2018 [cited by applicant]
US 9946547B2 · Yu et al. · 2018 [cited by applicant]
US 9971605B2 · Henry et al. · 2018 [cited by applicant]
US 9977674B2 · Rupley et al. · 2018 [cited by applicant]
US 9977675B2 · Nystad · 2018 [cited by applicant]
US 9977679B2 · Caulfield et al. · 2018 [cited by applicant]
US 9983882B2 · Greiner et al. · 2018 [cited by applicant]
US 9983884B2 · Maiyuran et al. · 2018 [cited by applicant]
US 10089277B2 · Mykland · 2018 [cited by applicant]
US 10540584B2 · McBride et al. · 2020 [cited by applicant]
US 10922146B1 · Minkin et al. · 2021 [cited by applicant]
US 11163486B2 · Bavishi et al. · 2021 [cited by applicant]
US 20020174318A1 · Studdard et al. · 2002 [cited by applicant]
US 20040111710A1 · Chakradhar et al. · 2004 [cited by applicant]
US 20050125786A1 · Dai · 2005 [cited by examiner]
US 20080288744A1 · Gonion · 2008 [cited by examiner]
US 20080288745A1 · Gonion · 2008 [cited by examiner]
US 20080288754A1 · Gonion · 2008 [cited by examiner]
US 20080288759A1 · Gonion · 2008 [cited by examiner]
US 20110320883A1 · Gonion · 2011 [cited by examiner]
US 20130326190A1 · Chung et al. · 2013 [cited by applicant]
US 20140149657A1 · Jakovljevic et al. · 2014 [cited by applicant]
US 20140317383A1 · Park et al. · 2014 [cited by applicant]
US 20150106597A1 · Godard et al. · 2015 [cited by applicant]
US 20150186146A1 · Kushida et al. · 2015 [cited by applicant]
US 20160246602A1 · Radhika et al. · 2016 [cited by applicant]
US 20170083330A1 · Burger · 2017 [cited by examiner]
US 20180225116A1 · Henry et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180322606A1 · Das et al. · 2018 [cited by applicant]
US 20180341493A1 · Roy et al. · 2018 [cited by applicant]
US 20180357172A1 · Lai · 2018 [cited by applicant]
US 20190004777A1 · Meixner · 2019 [cited by applicant]
US 20190347190A1 · Singh · 2019 [cited by applicant]
US 20190369990A1 · Doerr et al. · 2019 [cited by applicant]
US 20200026498A1 · Sumbul et al. · 2020 [cited by applicant]
US 20200241879A1 · Vorbach et al. · 2020 [cited by applicant]
US 20220050624A1 · Bavishi et al. · 2022 [cited by applicant]
US 20220075651A1 · Harboe et al. · 2022 [cited by applicant]
KR 1020150051083A · 2015 [cited by applicant]
WO WO2011038940A1 · 2011 [cited by applicant]
WO WO2020252763A1 · 2019 [cited by applicant]
Musicus, “The OKI Advanced Array Processor (AAP)—Development Software Manual, RLE Technical Report No. 539”, Research Laboratory of Electronics, Massachusetts Institute of Technology, Dec. 1988, 270 pages (Year: 1988). [cited by examiner]
International Search Report dated Feb. 28, 2024 for PCT/US2023/036949. [cited by applicant]
Chang, Kyungwook, and Kiyoung Choi. “Mapping control intensive kernels onto coarse-grained reconfigurable array architecture.” 2008 International SoC Design Conference. vol. 1. IEEE, 2008. [cited by applicant]
Musicus, B. R. (1988). The OKI advanced array processor (AAP): Development Software Manual. [cited by applicant]