IP Library › Granted Patent US 12,197,343
Granted Patent B2
US 12,197,343 · App. 18/357,732 · Granted Jan 14, 2025

Streaming engine with multi dimensional circular addressing selectable at each dimension

Inventor: Joseph Zbiciak (San Jose, CA)
Assignee: Texas Instruments Incorporated
G06F12/0897G06F9/3001G06F9/30036G06F9/30038G06F9/30047G06F9/30072G06F9/30101G06F9/3012G06F9/3013G06F9/30145G06F9/345G06F9/3822G06F9/383G06F9/3853G06F9/3877G06F9/3887G06F12/04G06F12/0815G06F12/0862G06F12/0875G06F9/3552G06F2212/1056G06F2212/452G06F2212/454G06F2212/6026
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,343
App. No.
18/357,732
Granted
Jan 14, 2025
Kind
B2
Abstract

A streaming engine employed in a digital data processor may specify a fixed read-only data stream defined by plural nested loops. An address generator produces address of data elements for the nested loops. A steam head register stores data elements next to be supplied to functional units for use as operands. A stream template register independently specifies a linear address or a circular address mode for each of the nested loops.

Claims (63)

1. A device comprising:

a first memory that stores a set of instructions that, when executed by a processor device, cause the processor device to:

store a set of values in a register that includes:

a first addressing mode value for a first loop that specifies whether to loop back to a first end of a first memory region when a second end of the first memory region is reached, wherein, when the first addressing mode value specifies to loop back to the first end of the first memory region, the set of values further specifies a size of the first memory region; and

a second addressing mode value for a second loop that specifies whether to loop back to a first end of a second memory region when a second end of the second memory region is reached, wherein, when the second addressing mode value specifies to loop back to the first end of the second memory region, the set of values further specifies a size of the second memory region; and

retrieve a set of data from a second memory based on the set of values by traversing the first memory region and the second memory region.

2. The device of claim 1 , wherein the first loop is nested within the second loop.

3. The device of claim 1 , wherein the second memory is a cache memory.

4. The device of claim 3 , wherein:

the set of instructions cause the set of values to be provided to a cache control circuit coupled to the cache memory; and

the cache control circuit is configured to retrieve the set of data from the cache memory by generating a set of addresses based on the set of values and providing the set of addresses to the cache memory.

5. The device of claim 4 , wherein:

the processor device includes the register; and

the set of instructions cause the set of values to be provided to the cache control circuit by the processor device.

6. The device of claim 4 , wherein the set of instructions cause the processor device to:

provide a read signal to the cache control circuit; and

based on the read signal, receive a data element of the set of data from the cache control circuit.

7. The device of claim 6 , wherein:

the data element is a first data element; and

the read signal specifies whether to replace the first data element in the cache control circuit with a second data element.

8. The device of claim 1 , wherein the set of instructions includes a stream open instruction that, when executed by the processor device, cause the set of data to be retrieved from the second memory.

9. The device of claim 1 , wherein the set of values includes:

a first value that specifies a first block size;

a second value that specifies a second block size; and

a third value that specifies whether the size of the first memory region is based on:

the first block size and the second block size; or

the first block size, independent of the second block size.

10. The device of claim 9 , wherein the third value specifies whether the size of the first memory region is based on a sum of the first block size, the second block size, and one.

11. The device of claim 1 further comprising:

the processor device coupled to the first memory; and

the second memory coupled to the processor device.

12. A device comprising:

a first memory configured to store a set of data;

a processor that includes a register configured to store a set of values that includes:

a first addressing mode value for a first loop that specifies whether to loop back to a first end of a first memory region when a second end of the first memory region is reached, wherein, when the first addressing mode value specifies to loop back to the first end of the first memory region, the set of values further specifies a size of the first memory region; and

a second addressing mode value for a second loop that specifies whether to loop back to a first end of a second memory region when a second end of the second memory region is reached, wherein, when the second addressing mode value specifies to loop back to the first end of the second memory region, the set of values further specifies a size of the second memory region;

a memory control circuit coupled between the processor and the first memory; and

a second memory configured to store a set of instructions, that when executed by the processor, cause the processor to cause the memory control circuit to provide the set of data from the first memory to the processor based on the set of values by traversing the first memory region and the second memory region.

13. The device of claim 12 , wherein the first loop is nested within the second loop.

14. The device of claim 12 , wherein the first memory is a cache memory.

15. The device of claim 12 , wherein:

the set of instructions cause the processor to provide the set of values to the memory control circuit; and

the memory control circuit is configured to retrieve the set of data from the first memory by generating a set of addresses based on the set of values and providing the set of addresses to the first memory.

16. The device of claim 15 , wherein the set of instructions cause the processor to:

provide a read signal to the memory control circuit; and

based on the read signal, receive a data element of the set of data from the memory control circuit.

17. The device of claim 16 , wherein:

the data element is a first data element; and

the read signal specifies whether to replace the first data element in the memory control circuit with a second data element.

18. The device of claim 12 , wherein the set of values includes:

a first value that specifies a first block size;

a second value that specifies a second block size; and

a third value that specifies whether the size of the first memory region is based on:

the first block size and the second block size; or

the first block size, independent of the second block size.

19. The device of claim 18 , wherein the third value specifies whether the size of the first memory region is based on a sum of the first block size, the second block size, and one.

20. A method comprising:

storing a set of values in a register of a processor, wherein the set of values includes:

a first addressing mode value for a first loop that specifies whether to loop back to a first end of a first memory region when a second end of the first memory region is reached, wherein, when the first addressing mode value specifies to loop back to the first end of the first memory region, the set of values further specifies a size of the first memory region; and

a second addressing mode value for a second loop that specifies whether to loop back to a first end of a second memory region when a second end of the second memory region is reached, wherein, when the second addressing mode value specifies to loop back to the first end of the second memory region, the set of values further specifies a size of the second memory region; and

in response to an instruction:

providing the set of values from the processor to a memory control circuit; and

retrieving, using the memory control circuit, a set of data from a memory based on the set of values by traversing the first memory region and the second memory region.

Continuity (4)
Continuation 17067986 · Oct 12, 2020
Continuation 16437900 · Jun 11, 2019
Continuation 15384451 · Dec 20, 2016
Related Publication 20230359565A1 · Nov 9, 2023
References Cited (35)
US 5499360A · Barbara · 1996 [cited by examiner]
US 6079008A · Clery, III · 2000 [cited by examiner]
US 6289434B1 · Roy · 2001 [cited by examiner]
US 6360348B1 · Yang · 2002 [cited by applicant]
US 6728741B2 · Keay · 2004 [cited by examiner]
US 6738967B1 · Radigan · 2004 [cited by examiner]
US 6754809B1 · Guttag · 2004 [cited by examiner]
US 6829696B1 · Balmer · 2004 [cited by examiner]
US 6832296B2 · Hooker · 2004 [cited by applicant]
US 6839831B2 · Balmer · 2005 [cited by examiner]
US 6986142B1 · Ehlig · 2006 [cited by examiner]
US 7047394B1 · Van Dyke · 2006 [cited by examiner]
US 7426668B2 · Mukherjee · 2008 [cited by examiner]
US 7428680B2 · Mukherjee · 2008 [cited by examiner]
US 7434131B2 · Mukherjee · 2008 [cited by examiner]
US 7549082B2 · Southgate · 2009 [cited by examiner]
US 7978506B2 · Lowrey · 2011 [cited by examiner]
US 8108652B1 · Hui · 2012 [cited by applicant]
US 10318433B2 · Zbiciak · 2019 [cited by examiner]
US 10755014B2 · Chou · 2020 [cited by examiner]
US 10789405B2 · Chou · 2020 [cited by examiner]
US 10810131B2 · Zbiciak · 2020 [cited by examiner]
US 11709779B2 · Zbiciak · 2023 [cited by examiner]
US 20030097652A1 · Roediger · 2003 [cited by examiner]
US 20040177225A1 · Furtek et al. · 2004 [cited by applicant]
US 20040181646A1 · Ben-David et al. · 2004 [cited by applicant]
US 20050283589A1 · Matsuo · 2005 [cited by applicant]
US 20060248322A1 · Southgate · 2006 [cited by examiner]
US 20070101100A1 · Al Sukhni · 2007 [cited by examiner]
US 20100106943A1 · Ukai · 2010 [cited by examiner]
US 20130339700A1 · Blasco-Allue · 2013 [cited by examiner]
J. C. Alves and P. C. Diniz, “Custom FPGA-based micro-architecture for streaming computing,” 2011 VII Southern Conference on Programmable Logic (SPL), Cordoba, Argentina, 2011, pp. 51-56, doi: 10.1109/SPL.2011.5782624. … [cited by examiner]
H. Kim, S. Ahn, Y. Oh, B. Kim, W. W. Ro and W. J. Song, “Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor Cores,” 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MI… [cited by examiner]
Severance, et al.; “Embedded supercomputing in FPGAs with the VectorBlox MXP Matrix Processor”; 2013 International Conference on Hardware/Software Codesign and System Synthesis (CODES+ISSS), Montreal, QC, Canada 2013; p… [cited by applicant]
Farahini, et al; “Atomic stream computer unit based on micro-thread level parallelism”; 2015 IEEE 26th International Conference on Application-specific Systems, Architectures and Processors (ASAP); Toronto, ON, Canada, … [cited by applicant]