IP Library › Granted Patent US 9,658,856
Granted Patent B2
US 9,658,856 · App. 14/976,228 · Granted May 23, 2017

Coalescing adjacent gather/scatter operations

Inventors: Andrew T. Forsyth (Kirkland, WA); Brian J. Hickmann (Sherwood, OR); Jonathan C. Hall (Hillsboro, OR); Christopher J. Hughes (Santa Clara, CA)
Assignee: Intel Corporation
G06F9/3853G06F9/30018G06F9/30043G06F9/30098G06F9/30105G06F9/30145G06F9/3804G06F9/3836G06F9/3887G06F12/0875G06F12/1027G06F13/4282G06F15/8007G06F9/3824G06F2212/1016G06F2212/452G06F2212/68
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,658,856
App. No.
14/976,228
Granted
May 23, 2017
Kind
B2
Abstract

According to one embodiment, a processor includes an instruction decoder to decode a first instruction to gather data elements from memory, the first instruction having a first operand specifying a first storage location and a second operand specifying a first memory address storing a plurality of data elements. The processor further includes an execution unit coupled to the instruction decoder, in response to the first instruction, to read contiguous a first and a second of the data elements from a memory location based on the first memory address indicated by the second operand, and to store the first data element in a first entry of the first storage location and a second data element in a second entry of a second storage location corresponding to the first entry of the first storage location.

Claims (34)

1. A system on a chip (SoC) comprising:

an integrated memory controller unit;

a communication device; and

a processor, the processor comprising:

a plurality of 64-bit general-purpose registers;

a plurality of 128-bit single instruction, multiple data (SIMD) registers;

a data cache to cache data;

an instruction cache to cache instructions;

an instruction fetch unit coupled to the instruction cache to fetch the instructions;

a decode unit coupled to the instruction fetch unit, the decode unit to decode the instructions, including a first instruction, the first instruction to indicate a 128-bit operand size, the first instruction having a first field to specify a first 128-bit SIMD destination register of the plurality of 128-bit SIMD registers, the first instruction having a second field to specify a 64-bit general-purpose register of the plurality of 64-bit general-purpose registers to store a base address, and the first instruction to indicate a data element width of 64-bits; and

an execution unit coupled to the decode unit, coupled to the plurality of 128-bit SIMD registers, and coupled to the plurality of 64-bit general-purpose registers, the execution unit to execute the first instruction to:

load a first structure and a second structure from a memory based on the base address, the first structure to include a first 64-bit data element, a second 64-bit data element, and a third 64-bit data element, the second structure to include a first 64-bit data element, a second 64-bit data element, and a third 64-bit data element, wherein the first 64-bit data element, the second 64-bit data element, and the third 64-bit data element of the first structure are to be consecutive elements in the memory, and wherein the first 64-bit data element, the second 64-bit data element, and the third 64-bit data element of the second structure are to be consecutive elements in the memory; and

store the first 64-bit data element of the first structure as a first 64-bit data element of the first 128-bit SIMD destination register, the second 64-bit data element of the first structure as a first 64-bit data element of a second 128-bit SIMD destination register, the third 64-bit data element of the first structure as a first 64-bit data element of a third 128-bit SIMD destination register, the first 64-bit data element of the second structure as a second 64-bit data element of the first 128-bit SIMD destination register, the second 64-bit data element of the second structure as a second 64-bit data element of the second 128-bit SIMD destination register, and the third 64-bit data element of the second structure as a second 64-bit data element of the third 128-bit SIMD destination register, wherein the first 64-bit data element of the first 128-bit SIMD destination register is to include least significant bits of the first 128-bit SIMD destination register, wherein the first 64-bit data element of the second 128-bit SIMD destination register is to include least significant bits of the second 128-bit SIMD destination register, and wherein the first 64-bit data element of the third 128-bit SIMD destination register is to include least significant bits of the third 128-bit SIMD destination register.

2. The SoC of claim 1 , wherein the first instruction has a data element width field to indicate the data element width of 64-bits.

3. The SoC of claim 1 , wherein a single bit of the first instruction is to indicate the 128-bit operand size.

4. The SoC of claim 1 , wherein the first, second, and third 128-bit SIMD source registers are a sequence of registers.

5. The SoC of claim 1 , wherein the processor has a reduced instruction set computing (RISC) architecture.

6. The SoC of claim 1 , further comprising display logic to couple to one or more displays.

7. The SoC of claim 1 , further comprising a graphics processing unit (GPU).

8. The SoC of claim 1 , further comprising an image processor.

9. A system on a chip (SoC) comprising:

an integrated memory controller unit;

a communication device; and

a processor, the processor comprising:

a plurality of 64-bit general-purpose registers;

a plurality of 128-bit single instruction, multiple data (SIMD) registers;

a data cache to cache data;

an instruction cache to cache instructions;

an instruction fetch unit coupled to the instruction cache to fetch the instructions;

a decode unit coupled to the instruction fetch unit, the decode unit to decode the instructions, including a first instruction, the first instruction having a single bit to indicate a 128-bit operand size, the first instruction having a first field to specify a first 128-bit SIMD destination register of the plurality of 128-bit SIMD registers, the first instruction having a second field to specify a 64-bit general-purpose register of the plurality of 64-bit general-purpose registers to store a 64-bit base address, the first instruction having a data element width field to indicate a data element width of 64-bits, and the first instruction to indicate an immediate offset to the 64-bit base address; and

an execution unit coupled to the decode unit, coupled to the plurality of 128-bit SIMD registers, and coupled to the plurality of 64-bit general-purpose registers, the execution unit to execute the first instruction to:

load a first structure and a second structure from a memory based on the 64-bit base address, the first structure to include a first 64-bit data element, a second 64-bit data element, and a third 64-bit data element, the second structure to include a first 64-bit data element, a second 64-bit data element, and a third 64-bit data element, wherein the first 64-bit data element, the second 64-bit data element, and the third 64-bit data element of the first structure are to be consecutive elements in the memory, and wherein the first 64-bit data element, the second 64-bit data element, and the third 64-bit data element of the second structure are to be consecutive elements in the memory; and

store the first 64-bit data element of the first structure as a first 64-bit data element of the first 128-bit SIMD destination register, the second 64-bit data element of the first structure as a first 64-bit data element of a second 128-bit SIMD destination register, the third 64-bit data element of the first structure as a first 64-bit data element of a third 128-bit SIMD destination register, the first 64-bit data element of the second structure as a second 64-bit data element of the first 128-bit SIMD destination register, the second 64-bit data element of the second structure as a second 64-bit data element of the second 128-bit SIMD destination register, and the third 64-bit data element of the second structure as a second 64-bit data element of the third 128-bit SIMD destination register, wherein the first 64-bit data element of the first 128-bit SIMD destination register is to include least significant bits of the first 128-bit SIMD destination register, wherein the first 64-bit data element of the second 128-bit SIMD destination register is to include least significant bits of the second 128-bit SIMD destination register, and wherein the first 64-bit data element of the third 128-bit SIMD destination register is to include least significant bits of the third 128-bit SIMD destination register,

wherein the first, second, and third 128-bit SIMD destination registers are a sequence of registers.

Continuity (2)
Continuation 13997784
Related Publication 20160103790A1 · Apr 14, 2016