IP Library › Granted Patent US 12,743,377
Granted Patent B2
US 12,743,377 · App. 18/785,026 · Granted Sep 22, 2026

Parallel processing architecture with block move support

Inventor: Peter Foley (Los Altos Hills, CA)
Assignee: Ascenium, Inc.
G06F12/084G06F9/544G06F2212/6042G06F2212/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,377
App. No.
18/785,026
Filed
Jul 26, 2024
Granted
Sep 22, 2026
Kind
B2
Art Unit
2132
USPC
711/117
Abstract

Techniques for task processing are disclosed. An array of compute elements is accessed. Each compute element within the array is known to a compiler and is coupled to its neighboring compute elements. The array of compute elements is coupled to at least one data cache. The data cache provides memory storage for the array. Control for the compute elements is provided on a cycle-by-cycle basis. Control is enabled by a stream of wide control words generated by the compiler. A load address and a store address are generated. The load and the store addresses comprise memory block move addresses. The memory block move addresses point to memory storage locations in the data cache. A memory block move is executed, based on the memory block move addresses. The data for the memory block move is transferred outside of the array.

Claims (38)

1 . A processor-implemented method for task processing comprising:

accessing an array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements, wherein the array of compute elements is coupled to at least one data cache, wherein the data cache provides memory storage for the array of compute elements;

providing control for the array of compute elements on a cycle-by-cycle basis, wherein the control is enabled by a stream of wide control words generated by the compiler;

generating a load address and a store address, wherein the load address and the store address comprise memory block move addresses, and wherein the memory block move addresses point to memory storage locations in the at least one data cache;

coupling load buffers located adjacent to at least one edge of the array of compute elements, wherein the load buffers provide storage for data obtained from the load address and a dataless store address; and

executing a memory block move, based on the memory block move addresses, wherein data for the memory block move is transferred outside of the array of compute elements.

2 . The method of claim 1 wherein the load address and the store address are generated in a same cycle.

3 . The method of claim 1 wherein the memory block move comprises a data cache to data cache transfer.

4 . The method of claim 1 wherein a control word from the stream of wide control words includes a load target start address, a store target start address, a block size, and a stride.

5 . The method of claim 4 wherein the generating a load address and a store address encompasses physical address translation of the load target start address and the store target start address, respectively.

6 . The method of claim 1 wherein the memory block move is executed as a pseudo-atomic operation.

7 . The method of claim 6 wherein the pseudo-atomic operation uses memory hazard detection and mitigation.

8 . The method of claim 1 wherein the memory block move that is transferred outside of the array of compute elements is enabled by the load buffers.

9 . The method of claim 1 wherein the load buffers are located adjacent to two opposite edges of the array of compute elements.

10 . The method of claim 1 further comprising coupling a crossbar switch between the load buffers and the at least one data cache.

11 . The method of claim 10 wherein the crossbar switch enables memory access anywhere within the at least one data cache.

12 . The method of claim 1 wherein the array of compute elements comprises a two-dimensional (2D) array.

13 . The method of claim 12 wherein the 2D array includes rows of compute elements and columns of compute elements.

14 . The method of claim 13 wherein the generating a load address and a store address is performed by one or more compute elements within a column of compute elements.

15 . The method of claim 1 wherein successful completion of the memory block move occurs within one architectural cycle.

16 . The method of claim 15 wherein the architectural cycle includes a plurality of clock cycles.

17 . The method of claim 1 wherein the memory block move implements a load-to-store forwarding operation.

18 . The method of claim 17 wherein the load-to-store forwarding operation enables hazard detection and mitigation.

19 . The method of claim 1 wherein the stream of wide control words comprises variable length control words generated by the compiler.

20 . A computer program product embodied in a non-transitory computer readable medium for task processing, the computer program product comprising code which causes one or more processors to perform operations of:

accessing an array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements, wherein the array of compute elements is coupled to at least one data cache, wherein the data cache provides memory storage for the array of compute elements;

providing control for the array of compute elements on a cycle-by-cycle basis, wherein the control is enabled by a stream of wide control words generated by the compiler;

generating a load address and a store address, wherein the load address and the store address comprise memory block move addresses, and wherein the memory block move addresses point to memory storage locations in the at least one data cache;

coupling load buffers located adjacent to at least one edge of the array of compute elements, wherein the load buffers provide storage for data obtained from the load address and a dataless store address; and

executing a memory block move, based on the memory block move addresses, wherein data for the memory block move is transferred outside of the array of compute elements.

21 . A computer system for task processing comprising:

a memory which stores instructions;

one or more processors coupled to the memory, wherein the one or more processors, when executing the instructions which are stored, are configured to:

access an array of compute elements, wherein each compute element within the array of compute elements is known to a compiler and is coupled to its neighboring compute elements within the array of compute elements, wherein the array of compute elements is coupled to at least one data cache, wherein the data cache provides memory storage for the array of compute elements;

provide control for the array of compute elements on a cycle-by-cycle basis, wherein the control is enabled by a stream of wide control words generated by the compiler;

generate a load address and a store address, wherein the load address and the store address comprise memory block move addresses, and wherein the memory block move addresses point to memory storage locations in the at least one data cache;

couple load buffers located adjacent to at least one edge of the array of compute elements, wherein the load buffers provide storage for data obtained from the load address and a dataless store address; and

execute a memory block move, based on the memory block move addresses, wherein data for the memory block move is transferred outside of the array of compute elements.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2026
From: FOLEY, PETER
To: ASCENIUM, INC.
Reel/Frame 075317/0542 →
Continuity (14)
Continuation In Part 17526003 · Nov 15, 2021
Continuation In Part 17465949 · Sep 3, 2021
Provisional Application 63536144 · Sep 1, 2023
Provisional Application 63529159 · Jul 27, 2023
Provisional Application 63254557 · Oct 12, 2021
Provisional Application 63232230 · Aug 12, 2021
Provisional Application 63229466 · Aug 4, 2021
Provisional Application 63193522 · May 26, 2021
Provisional Application 63166298 · Mar 26, 2021
Provisional Application 63125994 · Dec 16, 2020
Provisional Application 63114003 · Nov 16, 2020
Provisional Application 63091947 · Oct 15, 2020
Provisional Application 63075849 · Sep 9, 2020
Related Publication 20240385965A1 · Nov 21, 2024
References Cited (70)
US 5594884A · Matoba et al. · 1997 [cited by applicant]
US 5764994A · Craft · 1998 [cited by applicant]
US 6026239A · Patrick et al. · 2000 [cited by applicant]
US 7840777B2 · Mykland · 2010 [cited by applicant]
US 8677081B1 · Wentzlaff et al. · 2014 [cited by applicant]
US 8694978B1 · Rus et al. · 2014 [cited by applicant]
US 8856768B2 · Mykland · 2014 [cited by applicant]
US 8869123B2 · Mykland · 2014 [cited by applicant]
US 8949806B1 · Lee et al. · 2015 [cited by applicant]
US 9158544B2 · Mykland · 2015 [cited by applicant]
US 9304770B2 · Mykland · 2016 [cited by applicant]
US 9395992B2 · Doing et al. · 2016 [cited by applicant]
US 9424055B2 · Lundvall et al. · 2016 [cited by applicant]
US 9473155B2 · Staszewski et al. · 2016 [cited by applicant]
US 9477470B2 · Mykland · 2016 [cited by applicant]
US 9529715B2 · Kumar et al. · 2016 [cited by applicant]
US 9582277B2 · Muff et al. · 2017 [cited by applicant]
US 9594559B2 · Fontenot et al. · 2017 [cited by applicant]
US 9600287B2 · Gschwind et al. · 2017 [cited by applicant]
US 9633160B2 · Mykland · 2017 [cited by applicant]
US 9652238B2 · Muff et al. · 2017 [cited by applicant]
US 9684511B2 · Shanbhogue et al. · 2017 [cited by applicant]
US 9830164B2 · Yazdani · 2017 [cited by applicant]
US 9851969B2 · Greiner et al. · 2017 [cited by applicant]
US 9886277B2 · Loktyukhin et al. · 2018 [cited by applicant]
US 9898293B2 · Whittaker · 2018 [cited by applicant]
US 9921836B2 · Kumar et al. · 2018 [cited by applicant]
US 9928062B2 · Azagury et al. · 2018 [cited by applicant]
US 9934040B2 · Bonanno et al. · 2018 [cited by applicant]
US 9946547B2 · Yu et al. · 2018 [cited by applicant]
US 9971605B2 · Henry et al. · 2018 [cited by applicant]
US 9977674B2 · Rupley, II et al. · 2018 [cited by applicant]
US 9977675B2 · Nystad · 2018 [cited by applicant]
US 9977679B2 · Caulfield et al. · 2018 [cited by applicant]
US 9983882B2 · Greiner et al. · 2018 [cited by applicant]
US 9983884B2 · Maiyuran et al. · 2018 [cited by applicant]
US 10089277B2 · Mykland · 2018 [cited by applicant]
US 10540584B2 · McBride et al. · 2020 [cited by applicant]
US 10922146B1 · Minkin et al. · 2021 [cited by applicant]
US 11163486B2 · Bavishi et al. · 2021 [cited by applicant]
US 20020169938A1 · Scott et al. · 2002 [cited by applicant]
US 20020174318A1 · Studdard et al. · 2002 [cited by applicant]
US 20030196072A1 · Chinnakonda · 2003 [cited by examiner]
US 20040111710A1 · Chakradhar et al. · 2004 [cited by applicant]
US 20130326190A1 · Chung et al. · 2013 [cited by applicant]
US 20140149657A1 · Jakovljevic et al. · 2014 [cited by applicant]
US 20140317383A1 · Park et al. · 2014 [cited by applicant]
US 20150039855A1 · Pechanek · 2015 [cited by examiner]
US 20150127921A1 · Choi et al. · 2015 [cited by applicant]
US 20150186146A1 · Kushida et al. · 2015 [cited by applicant]
US 20160246602A1 · Radhika et al. · 2016 [cited by applicant]
US 20180225116A1 · Henry et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180322606A1 · Das et al. · 2018 [cited by applicant]
US 20180341493A1 · Roy et al. · 2018 [cited by applicant]
US 20190004777A1 · Meixner · 2019 [cited by applicant]
US 20190347190A1 · Singh · 2019 [cited by applicant]
US 20190369990A1 · Doerr et al. · 2019 [cited by applicant]
US 20200026498A1 · Sumbul et al. · 2020 [cited by applicant]
US 20200241879A1 · Vorbach et al. · 2020 [cited by applicant]
US 20210263854A1 · Ingalls · 2021 [cited by examiner]
US 20220050624A1 · Bavishi et al. · 2022 [cited by applicant]
US 20220121450A1 · Brewer · 2022 [cited by examiner]
KR 1020150051083A · 2015 [cited by applicant]
WO WO2011038940A1 · 2011 [cited by applicant]
WO WO2020252763A1 · 2019 [cited by applicant]
WO WO0233565A2 · 2022 [cited by applicant]
International Search Report dated Nov. 13, 2024 for PCT/US2024/039674. [cited by applicant]
Chang, Kyungwook, and Kiyoung Choi. “Mapping control intensive kernels onto coarse-grained reconfigurable array architecture.” 2008 International SoC Design Conference. vol. 1. IEEE, 2008. [cited by applicant]
Musicus, B. R. (1988). The OKI advanced array processor (AAP): Development Software Manual. [cited by applicant]