IP Library Granted Patent US 8,954,674
Granted Patent B2
US 8,954,674 · App. 14/048,291 · Granted Feb 10, 2015

Scatter-gather intelligent memory architecture for unstructured streaming data on multiprocessor systems

Inventors: Daehyun Kim (San Jose, CA); Christopher J. Hughes (Santa Clara, CA); Yen-Kuang Chen (Cupertino, CA); Partha Kundu (Palo Alto, CA)
Assignee: Intel Corporation
G11C7/1075G06F12/0817G06F12/0862G06F12/0877G06F12/0891G06F12/0897G06F12/08G06F12/0815G11C7/1072G06F12/0811G06F2212/6026
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,954,674
App. No.
14/048,291
Granted
Feb 10, 2015
Kind
B2
Abstract

A scatter/gather technique optimizes unstructured streaming memory accesses, providing off-chip bandwidth efficiency by accessing only useful data at a fine granularity, and off-loading memory access overhead by supporting address calculation, data shuffling, and format conversion.

Claims (26)

1. A processor comprising:

a core including:

a computation processor; and

a scatter/gather engine coupled to the computation processor, the scatter/gather engine to generate sub-cache line sized non-sequential data accesses to a memory based on an access pattern communicated to the scatter/gather engine from and defined by an application, to communicate sub-cache line sized data with the memory, wherein the scatter/gather engine includes an access processor, an access pattern generator, and a cache interface, wherein the access pattern generator is to generate the sub-cache line sized non-sequential data accesses to the memory based on an indirect access pattern; and

a cache coupled to the core, wherein data is to be transferred between the cache and the memory using full-cache line sized transfers.

2. The processor of claim 1 , wherein the access processor is to compute memory addresses for the sub-cache line sized non-sequential data accesses and perform data format conversion.

3. The processor of claim 1 , wherein the access processor is to generate the sub-cache line sized non-sequential data accesses to the memory based on the indirect access pattern.

4. The processor of claim 1 , wherein the access pattern generator is to generate the sub-cache line sized non-sequential data accesses to the memory based on a stride-based access pattern.

5. The processor of claim 1 , wherein the scatter/gather engine further comprises:

a stream port coupled to the computation processor, the stream port including a buffer capable of storing ordered data accessible by the access processor and the computation processor.

6. The processor of claim 1 , wherein the cache interface in conjunction with the cache is to provide data coherence when a same data is accessed through both the cache and the scatter/gather engine.

7. The processor of claim 1 , further comprising:

a memory controller coupled to the scatter/gather engine and the memory, the memory controller to support both cache line and sub-cache line sized data accesses to the memory.

8. A machine-readable non-transitory medium having stored thereon instructions, which if performed by a machine cause the machine to perform a method comprising:

transferring full-cache line sized data between a cache of a multicore processor and an off-chip memory;

generating, by a scatter/gather engine of the multicore processor, sub-cache line sized non-sequential data accesses to the off-chip memory based on an indirect access pattern defined and communicated by an application to transfer sub-cache line sized data between the multicore processor and the off-chip memory, the sub-cache line sized data having fewer bits than the full-cache line sized data transfers; and

computing memory addresses for the sub-cache line sized non-sequential data accesses, performing data format conversion, and generating the sub-cache line sized non-sequential data access addresses for the off-chip memory based on the indirect access pattern.

9. The machine-readable non-transitory medium of claim 8 , wherein the method further comprises:

allocating a stream port in the scatter/gather engine to handle the sub-cache line sized data; and

directing access to the off-chip memory through the allocated stream port.

10. The machine-readable non-transitory medium of claim 8 , wherein the method further comprises:

allocating a stream port in the scatter/gather engine to a thread in the multicore processor; and

in response to a thread context switch, releasing the stream port after write data stored in the stream port has been written to the off-chip memory.

11. The machine-readable non-transitory medium of claim 8 , wherein the method further comprises enforcing data coherency when a same data is accessed through both the cache and the scatter/gather engine.

12. The machine-readable non-transitory medium of claim 8 , wherein the method further comprises enforcing data coherency via a mutual exclusion of the data in a buffer in the scatter/gather engine or in the cache.

13. The machine-readable non-transitory medium of claim 12 , wherein the method further comprises enforcing the data coherency via an address range check in a directory.

Continuity (4)
Continuation 13782515 · Mar 1, 2013
Continuation 13280117 · Oct 24, 2011
Continuation 11432753 · May 10, 2006
Related Publication 20140040542A1 · Feb 6, 2014