IP Library Granted Patent US 9,323,672
Granted Patent B2
US 9,323,672 · App. 14/585,573 · Granted Apr 26, 2016

Scatter-gather intelligent memory architecture for unstructured streaming data on multiprocessor systems

Inventors: Daehyun Kim (San Jose, CA); Christopher J. Hughes (Santa Clara, CA); Yen-Kuang Chen (Cupertino, CA); Partha Kundu (Palo Alto, CA)
Assignee: Intel Corporation
G06F12/0806G06F12/08G06F12/0811G06F12/0815G06F12/0817G06F12/0862G06F12/0877G06F12/0891G06F12/0897G11C7/1072G11C7/1075G06F2212/6026G06F2212/62Y02B60/1225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,323,672
App. No.
14/585,573
Granted
Apr 26, 2016
Kind
B2
Abstract

A scatter/gather technique optimizes unstructured streaming memory accesses, providing off-chip bandwidth efficiency by accessing only useful data at a fine granularity, and off-loading memory access overhead by supporting address calculation, data shuffling, and format conversion.

Claims (32)

1. A processor comprising:

a plurality of cores including a first core having:

a computation processor; and

a programmable engine coupled to the computation processor to perform address calculation and memory accesses, the programmable engine to generate sub-cache line sized non-sequential data accesses to a memory based on an access pattern to communicate sub-cache line sized data with the memory, the access pattern explicitly provided by an application, wherein the programmable engine is to store a stride access pattern data structure to be allocated for a first matrix of a matrix multiply function; and

a cache coupled to the first core, wherein data is to be transferred between the cache and the memory via a full-cache line sized transfer.

2. The processor of claim 1 , wherein the programmable engine includes an access processor, an access pattern generator, and a cache interface.

3. The processor of claim 2 , wherein the access processor is to compute memory addresses for the sub-cache line sized non-sequential data accesses and perform data format conversion.

4. The processor of claim 2 , wherein the access processor is to prepare an operand while the computation processor is to perform a computation.

5. The processor of claim 2 , wherein the access processor is to generate the sub-cache line sized non-sequential data accesses to the memory based on an indirect access pattern.

6. The processor of claim 5 , wherein the indirect access pattern comprises a sparse matrix dense vector multiplication.

7. The processor of claim 2 , wherein the access pattern generator is to generate the sub-cache line sized non-sequential data accesses to the memory based on a stride-based access pattern.

8. The processor of claim 2 , wherein the computation processor and the access processor are to be associated with a same process identifier.

9. The processor of claim 2 , wherein the programmable engine further comprises:

a stream port coupled to the computation processor, the stream port including a buffer to store ordered data accessible by the access processor and the computation processor.

10. The processor of claim 9 , wherein the buffer comprises a first port and a second port to enable the computation processor and the access processor to concurrently access the buffer.

11. The processor of claim 9 , wherein the computation processor is to allocate the stream port, map the access pattern to the stream port, and wherein the access processor is to execute a memory access handler to perform a scatter/gather operation to place the sub-cache line sized data in the stream port.

12. The processor of claim 11 , wherein the memory access handler is to receive a port value to identify the stream port and a pattern value to identify the access pattern.

13. The processor of claim 1 , wherein the computation processor is to perform data computations for the first matrix using a stream port of the access processor.

14. A system comprising:

a multicore processor having a plurality of cores including a first core having a computation processor and a programmable engine to perform address calculation and memory accesses to generate sub-cache line sized non-sequential data accesses to a memory based on an access pattern and communicate sub-cache line sized data with the memory, a first memory controller and a second memory controller, the access pattern explicitly provided by an application, wherein the programmable engine is to store a stride access pattern data structure to be allocated for a first matrix of a matrix multiply function; and

the memory coupled to the multicore processor, the memory including a first plurality of channels associated with the first memory controller and a second plurality of channels associated with the second memory controller.

15. The system of claim 14 , wherein each of the first plurality of channels are to be used to perform scatter/gather operations.

16. The system of claim 14 , wherein the memory is to support cache line size data transfer and sub-cache line size data transfer.

17. A non-transitory machine-accessible storage medium including instructions that when executed cause a system to perform a method comprising:

performing, in a programmable engine coupled to a computation processor of a core, address calculation and memory accesses;

generating, in the programmable engine, sub-cache line sized non-sequential data accesses to a memory based on an access pattern;

storing, in the programmable engine, a stride access pattern data structure to be allocated for a first matrix of a matrix multiply function; and

communicating sub-cache line sized data with the memory, the access pattern explicitly provided by an application, wherein the memory supports cache line size data transfer and sub-cache line size data transfer.

18. The non-transitory machine-accessible storage medium of claim 17 , wherein the method further comprises:

computing, by an access processor of the programmable engine, memory addresses for the sub-cache line sized non-sequential data accesses; and

performing, in the access processor, data format conversion.

19. The non-transitory machine-accessible storage medium of claim 18 , wherein the method further comprises preparing, by the access processor, an operand while the computation processor performs a computation.

Continuity (5)
Continuation 14048291 · Oct 8, 2013
Continuation 13782515 · Mar 1, 2013
Continuation 13280117 · Oct 24, 2011
Continuation 11432753 · May 10, 2006
Related Publication 20150178200A1 · Jun 25, 2015