IP Library Granted Patent US 12,468,535
Granted Patent B2
US 12,468,535 · App. 18/595,866 · Granted Nov 11, 2025

Cooperative instruction prefetch on multicore system

Inventors: Rahul Nagarajan (San Jose, CA); Christopher Leary (Sunnyvale, CA); Thejasvi Magudilu Vijayaraj (Santa Clara, CA); Thomas James Norrie (San Jose, CA)
Assignee: Google LLC
G06F9/3802G06F9/3887
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,535
App. No.
18/595,866
Granted
Nov 11, 2025
Kind
B2
Abstract

Aspects of the disclosure are directed to methods, systems, and apparatuses using an instruction prefetch pipeline architecture that provides good performance without the complexity of a full cache coherent solution deployed in conventional CPUs. The architecture can include components which can be used to construct an instruction prefetch pipeline, including instruction memory (TiMem), instruction buffer (iBuf), a prefetch unit, and an instruction router.

Claims (26)

1 . A hardware circuit, comprising:

a plurality of tiles respectively configured to operate in parallel to stream data from an upstream input to a downstream destination; and

a plurality of task instruction memories, each task instruction memory being arranged in a sequence and coupled to the plurality of tiles via an instruction router that determines a sequence for streaming the data and filters requests for instructions to remove duplicate requests.

2 . The hardware circuit of claim 1 , wherein each tile comprises a tile access core comprising a prefetch unit.

3 . The hardware circuit of claim 1 , wherein each tile comprises a tile execute core comprising a prefetch unit.

4 . The hardware circuit of claim 1 , further comprising an instruction broadcast bus comprising a plurality of data processing lanes, the number of data processing lanes corresponding to the number of task instruction memories.

5 . The hardware circuit of claim 1 , further comprising an instruction request bus comprising a plurality of data processing lanes, the number of data processing lanes corresponding to the number of task instruction memories.

6 . The hardware circuit of claim 1 , wherein each tile comprises a prefetch unit configured to provide a request to at least one task instruction memory during a prefetch window.

7 . The hardware circuit of claim 1 , wherein the instruction router comprises a round robin arbiter configured to arbitrate prefetch read requests.

8 . The hardware circuit of claim 1 , wherein each tile comprises an instruction buffer configured to store instructions for a processing core.

9 . A computer-implemented method comprising:

streaming, respectively by operating a plurality of tiles of a hardware circuit in parallel, data from an upstream input to a downstream destination;

arranging, by an instruction router of the hardware circuit, each task instruction memory of a plurality of task instruction memories in a sequence by coupling to the plurality of tiles to determine a sequence for streaming the data; and

filtering, by the instruction router, requests for instructions to remove duplicate requests.

10 . The method of claim 9 , wherein each tile comprises a tile access core comprising a prefetch unit.

11 . The method of claim 9 , wherein each tile comprises a tile execute core comprising a prefetch unit.

12 . The method of claim 9 , wherein the hardware circuit further comprises an instruction broadcast bus comprising a plurality of data processing lanes, the number of data processing lanes corresponding to the number of task instruction memories.

13 . The method of claim 9 , wherein the hardware circuit further comprises an instruction request bus comprising a plurality of data processing lanes, the number of data processing lanes corresponding to the number of task instruction memories.

14 . The method of claim 9 , further comprising providing, by a prefetch unit of the hardware circuit, a request to at least one task instruction memory during a prefetch window.

15 . The method of claim 9 , further comprising arbitrating, by the instruction router, prefetch read requests via a round robin arbiter.

16 . The method of claim 9 , further comprising storing, by an instruction buffer of the hardware circuit, instructions for a processing core.

17 . A non-transitory computer-readable medium for storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

streaming, respectively by operating a plurality of tiles of a hardware circuit in parallel, data from an upstream input to a downstream destination;

arranging, by an instruction router of the hardware circuit, each task instruction memory of a plurality of task instruction memories in a sequence by coupling to the plurality of tiles to determine a sequence for streaming the data; and

filtering, by the instruction router, requests for instructions to remove duplicate requests.

18 . The non-transitory computer-readable medium of claim 17 , wherein the operations further comprise providing, by a prefetch unit of the hardware circuit, a request to at least one task instruction memory during a prefetch window.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2024
From: NAGARAJAN, RAHUL; LEARY, CHRISTOPHER; VIJAYARAJ, THEJASVI MAGUDILU; NORRIE, THOMAS JAMES
To: GOOGLE LLC
Reel/Frame 066663/0826 →
Continuity (3)
Continuation 17972681 · Oct 25, 2022
Provisional Application 63281960 · Nov 22, 2021
Related Publication 20240211264A1 · Jun 27, 2024
References Cited (70)
US 6704860B1 · Moore · 2004 [cited by applicant]
US 7206922B1 · Steiss · 2007 [cited by examiner]
US 7248266B2 · Tuomi · 2007 [cited by applicant]
US 7249202B2 · Simon et al. · 2007 [cited by applicant]
US 7461210B1 · Wentzlaff et al. · 2008 [cited by applicant]
US 7613886B2 · Yamazaki · 2009 [cited by applicant]
US 7636835B1 · Ramey et al. · 2009 [cited by applicant]
US 7917701B2 · Charra et al. · 2011 [cited by applicant]
US 9086872B2 · Hargil et al. · 2015 [cited by applicant]
US 9323672B2 · Kim et al. · 2016 [cited by applicant]
US 9628146B2 · Van Nieuwenhuyze et al. · 2017 [cited by applicant]
US 9851917B2 · Park et al. · 2017 [cited by applicant]
US 11113233B1 · Volpe · 2021 [cited by examiner]
US 11182110B1 · Ansari · 2021 [cited by examiner]
US 11210760B2 · Nurvitadhi et al. · 2021 [cited by applicant]
US 11340792B2 · Danilov et al. · 2022 [cited by applicant]
US 11360934B1 · Abts · 2022 [cited by examiner]
US 20040133765A1 · Tanaka · 2004 [cited by examiner]
US 20070198901A1 · Ramchandran et al. · 2007 [cited by applicant]
US 20120159130A1 · Smelyanskiy et al. · 2012 [cited by applicant]
US 20120303932A1 · Farabet · 2012 [cited by examiner]
US 20130151482A1 · Tofano · 2013 [cited by applicant]
US 20170004089A1 · Clemons et al. · 2017 [cited by applicant]
US 20170351516A1 · Mekkat et al. · 2017 [cited by applicant]
US 20180067899A1 · Rub · 2018 [cited by applicant]
US 20180081690A1 · Krishna et al. · 2018 [cited by applicant]
US 20180109449A1 · Sebexen et al. · 2018 [cited by applicant]
US 20180217836A1 · Johnson · 2018 [cited by applicant]
US 20180285233A1 · Norrie et al. · 2018 [cited by applicant]
US 20180322387A1 · Sridharan et al. · 2018 [cited by applicant]
US 20190004814A1 · Chen et al. · 2019 [cited by applicant]
US 20190050717A1 · Temam et al. · 2019 [cited by applicant]
US 20190121784A1 · Wilkinson et al. · 2019 [cited by applicant]
US 20190303311A1 · Bilski et al. · 2019 [cited by applicant]
US 20190340722A1 · Aas et al. · 2019 [cited by applicant]
US 20200120154A1 · Ren et al. · 2020 [cited by applicant]
US 20200183738A1 · Champigny · 2020 [cited by applicant]
US 20200192742A1 · Boettcher · 2020 [cited by examiner]
US 20200293488A1 · Ray et al. · 2020 [cited by applicant]
US 20200310797A1 · Corbal · 2020 [cited by examiner]
US 20210081691A1 · Chen et al. · 2021 [cited by applicant]
US 20210109761A1 · Wang et al. · 2021 [cited by applicant]
US 20210194793A1 · Huse · 2021 [cited by applicant]
US 20210303481A1 · Ray et al. · 2021 [cited by applicant]
US 20210326067A1 · Li · 2021 [cited by applicant]
US 20220012060A1 · Fok · 2022 [cited by examiner]
US 20230229524A1 · Dearth et al. · 2023 [cited by applicant]
JP 2018521427A · 2018 [cited by applicant]
JP 2021177366A · 2021 [cited by applicant]
JP 2021532430A · 2021 [cited by applicant]
WO 9707451A2 · 1997 [cited by applicant]
WO 2006106342A2 · 2006 [cited by applicant]
Yang, L et al., Optimal Application Mapping and Scheduling for Network-on-Chips with Computation in STT-RAM Based Router. 2019, IEEE, pp. 1174-1189. (Year: 2019). [cited by examiner]
Paul, S et al., Dynamic task allocation and scheduling with contention-awareness for Network-on-Chip based multicore systems, Jan. 2021, Elsevier, 16 pages. (Year: 2021). [cited by examiner]
Deb, D et al., ECAP: energy-efficient caching for prefetch blocks in tiled chip multiprocessor, 2019, The Institution of Engineering and Technology , pp. 417-428. (Year: 2019). [cited by examiner]
Chhugan et al., “Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture”, PVLDB '08, Aug. 23-28, 2008, Auckland, New Zealand, 12 pages. [cited by applicant]
Yavits et al., “Sparse Matrix Multiplication on an Associative Processor”, retrieved from the Internet on Nov. 19, 2021 <https://arxiv.org/ftp/arxiv/papers/1705/1705.07282.pdf>, 10 pages. [cited by applicant]
Batcher odd-even mergesort, From Wikipedia, the free encyclopedia, retrieved from the Internet on Nov. 19, 2021 <https://en.wikipedia.org/wiki/Batcher_odd%E2%80%93even_mergesort>, 2 pages. [cited by applicant]
Chole et al. SparseCore: An Accelerator for Structurally Sparse CNNs. 2018. 3 pages. Retrieved from the Internet: <https://mlsys.org/Conferences/doc/2018/72.pdf>. [cited by applicant]
Fang et al. Active Memory Operations. Proceedings of the International Conference on Supercomputing—Proceedings of ICS07: 21st ACM International Conference on Supercomputing 2007 Association for Computing Machinery US, … [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/033710 dated Dec. 19, 2022. 17 pages. [cited by applicant]
Prabhakar et al. Plasticine: A Reconfigurable Architecture for Parallel Patterns. Proceedings of the 44th Annual International Symposium on Computer Architecture , ISCA ' 17, ACM Press, New York, New York, USA, Jun. 24,… [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/049353 dated Feb. 6, 2023. 17 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/048357 dated Feb. 22, 2023. 16 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/048919 dated Feb. 22, 2023. 16 pages. [cited by applicant]
Deb, D., et al., COPE: Reducing Cache Pollution and Network Contention by Inter-tile Coordinated Prefetching in NoC-based MPSoCs, 2020, ACM. pp. 17:1-17:31. (Year: 2020). [cited by applicant]
Office Action for Japanese Patent Application No. 2023-570416 dated Jan. 7, 2025. 6 pages. [cited by applicant]
Gaye et al. An Analysis of the Hot Spot Contention and Message Combining on the SSS-MIN. May 25, 1994. D-I vol. J77-D-I, No. 5, pp. 354-363. [cited by applicant]
Notice of Grant for Japanese Patent Application No. 2023-571623 dated Feb. 4, 2025. 3 pages. [cited by applicant]
Notice of Grant for Japanese Patent Application No. 2023-572877 dated Feb. 4, 2025. 3 pages. [cited by applicant]