IP Library › Granted Patent US 11,907,158
Granted Patent B2
US 11,907,158 · App. 17/135,465 · Granted Feb 20, 2024

Vector processor with vector first and multiple lane configuration

Inventor: Steven Jeffrey Wallach (Dallas, TX)
Assignee: Micron Technology, Inc.
G06F15/8076G06F7/57G06F9/30101
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,907,158
App. No.
17/135,465
Granted
Feb 20, 2024
Kind
B2
Abstract

A vector processor with a vector first and multi-lane configuration. A vector operation for a vector processor can include a single vector or multiple vectors as input. Multiple lanes for the input can be used to accelerate the operation in parallel. And, a vector first configuration can enhance the multiple lanes by reducing the number of elements accessed in the lanes to perform the operation in parallel.

Claims (58)

1. An apparatus, comprising:

a first vector register configured to store a first vector having first data elements of a predetermined number, the first vector register having a first lane and a second lane;

a second vector register configured to store a second vector having second data elements of the predetermined number, the second vector register having a third lane and a fourth lane;

a third register configured to store an index;

a plurality of arithmetic logic units;

a first circuit coupled from the first vector register to a first arithmetic logic unit of the plurality of arithmetic logic units and a second arithmetic logic unit of the plurality of arithmetic logic units, the first circuit configured to provide a first lane input to the first arithmetic logic unit to generate a first output vector, the first circuit configured to provide a second lane input to the second arithmetic logic unit to generate a second output vector, the first lane input and the second lane input including data elements selected according to the index from the first vector, wherein the first circuit comprises:

a first multiplexor configured to receive first vector elements from the first lane of the first vector and output the first lane input to the first arithmetic logic unit; and

a second multiplexor configured to receive second vector elements from the second lane of the first vector and output the second lane input to the second arithmetic logic unit; and

a second circuit coupled from the second vector register to a third arithmetic logic unit of the plurality of arithmetic logic units, the second circuit further configured to provide a third lane input to the first arithmetic logic unit and a fourth lane input to the second arithmetic logic unit, the third lane input and the fourth lane input including data elements selected according to the index from the second vector.

2. The apparatus of claim 1 , wherein the first lane input has a predetermined number of data elements, including a first portion selected from the first vector and a second portion selected from the second vector.

3. The apparatus of claim 2 , wherein in the first lane input the second portion follows the first portion as a vector input to the first arithmetic logic unit.

4. The apparatus of claim 3 , wherein in the first vector register the first portion starts at the index; and in the second vector register the second portion ends before the index.

5. The apparatus of claim 4 , wherein the first circuit includes the first multiplexor configured to select the first portion and the second portion from the first vector register and the second vector register to form the first lane input as the vector input to the first arithmetic logic unit.

6. The apparatus of claim 5 , wherein the apparatus further comprises:

a fourth vector register configured to store a fourth vector having fourth data elements of the predetermined number; and

wherein the second circuit or the first circuit is further configured to couple the second vector register and the fourth vector register to the second arithmetic logic unit, the second circuit or the first circuit including a third multiplexor configured to select data elements according to the index from the second vector register and the fourth vector register to provide a fifth vector as a vector input to the second arithmetic logic unit.

7. The apparatus of claim 6 , wherein the fifth vector has a predetermined number of fifth data elements, including a third portion selected from the second vector register and a fourth portion selected from the fourth vector register.

8. The apparatus of claim 7 , wherein in the fifth vector the fourth portion follows the third portion as the vector input to the second arithmetic logic unit; in the second vector register the third portion starts at the index; and in the fourth vector register the fourth portion ends before the index; the second portion and the third portion are non-overlapping portions in the second vector register; and the second portion and the third portion provide the second data elements stored in the second vector register.

9. The apparatus of claim 8 , wherein the first arithmetic logic unit and the second arithmetic logic unit process the first lane input and the fourth vector respectively in parallel.

10. The apparatus of claim 5 , wherein the apparatus further comprises:

a fourth vector register configured to store a fourth vector having fourth data elements of the predetermined number; and

wherein the second circuit or the first circuit is further configured to couple the fourth vector register and the first vector register to the second arithmetic logic unit, the second circuit or the first circuit including a third multiplexor configured to select data elements according to the index from the fourth vector register and the first vector register to provide a fifth vector as a vector input to the second arithmetic logic unit.

11. The apparatus of claim 10 , wherein the fifth vector has a predetermined number of fifth data elements, including a third portion selected from the fourth vector register and a fourth portion selected from the first vector register.

12. The apparatus of claim 11 , wherein in the fifth vector the fourth portion follows the third portion as the vector input to the second arithmetic logic unit; in the fourth vector register the third portion starts at the index; and in the first vector register the fourth portion ends before the index.

13. The apparatus of claim 12 , wherein the first portion and the fourth portion are non-overlapping portions in the first vector register; the first portion and the fourth portion provide the first data elements stored in the first vector register; and the first arithmetic logic unit and the second arithmetic logic unit process the first lane input and the fifth vector respectively in parallel.

14. The apparatus of claim 5 , wherein the apparatus further comprises:

a second arithmetic logic unit, wherein the second circuit or the first circuit is further configured to couple the second vector register and the first vector register to the second arithmetic logic unit, the second circuit or the first circuit including a third multiplexor configured to select data elements according to the index from the second vector register and the first vector register to provide a fourth vector as a vector input to the second arithmetic logic unit, the first arithmetic logic unit and the second arithmetic logic unit configured to process the first lane input and the fourth vector respectively in parallel.

15. An apparatus, comprising:

a plurality of vector registers configured to store a plurality of vectors of data elements respectively;

a register configured to store an index;

a plurality of processing units configured to operate in parallel; and

a circuit controlled by the register and coupled between the plurality of vector registers and the plurality of processing units, the circuit configured to provide a plurality of input vectors to the plurality of processing units respectively, wherein the circuit is configured to generate each respective input vector in the plurality of input vectors using:

a first portion selected according to the index from a first lane and a second lane of a first vector register in the plurality of vector registers; and

a second portion selected according to the index from a third lane and a fourth lane of a second vector register in the plurality of vector registers,

wherein the circuit comprises:

a first multiplexor configured to receive first data from the first lane of the first vector register and second data from the second lane of the first vector register and output a first lane input to a first processing unit of the plurality of processing units;

a second multiplexor configured to receive third data from the first lane of the first vector register and fourth data from the second lane of the first vector register and output a second lane input to a second processing unit of the plurality of processing units;

a third lane input to the first processing unit from a third multiplexor associated with the third lane of the second vector register, wherein the third multiplexor further outputs the third lane input to a third processing unit of the plurality of processing units; and

a fourth lane input to the second processing unit from a fourth multiplexor associated with the fourth lane of the second vector register.

16. The apparatus of claim 15 , wherein each vector of data elements stored in a respective vector register in the plurality of vector registers has two portions identified by the index; and the circuit is configured to provide the two portions respective in two different input vectors in the plurality of input vectors.

17. The apparatus of claim 16 , wherein the circuit includes a plurality of multiplexors configured to provide the plurality of input vectors to the plurality of processing units respectively in parallel.

18. A method, comprising:

storing, into a plurality of vector registers, a plurality of vectors of data elements respectively;

storing, into a register, an index;

generating, using a circuit controlled by the register and coupled between the plurality of vector registers and a plurality of processing units, a plurality of input vectors to the plurality of processing units respectively, wherein each respective input vector in the plurality of input vectors includes:

a first portion selected according to the index from a first lane and a second lane of a first vector register in the plurality of vector registers; and

a second portion selected according to the index from a third lane and a fourth lane of a second vector register in the plurality of vector registers,

wherein the circuit comprises:

a first multiplexor configured to receive first data from the first lane of the first vector register and second data from the second lane of the first vector register and output a first lane input into a first processing unit of the plurality of processing units;

a second multiplexor configured to receive third data from the first lane of the first vector register and fourth data from the second lane of the first vector register and output a second lane input to a second processing unit of the plurality of processing units;

a third lane input to the first processing unit from a third multiplexor associated with the third lane of the second vector register, wherein the third multiplexor further outputs the third lane input to a third processing unit of the plurality of processing units; and

a fourth lane input to the second processing unit from a fourth multiplexor associated with the fourth lane of the second vector register; and

processing, by the plurality of processing units, the plurality of input vectors in parallel.

19. The method of claim 18 , wherein the generating comprises:

selecting, using a plurality of multiplexors in the circuit and according to the index, data elements from the plurality of vector registers to generate the plurality of input vectors respectively in parallel.

20. The method of claim 18 , wherein the generating comprises:

retrieving a plurality of data elements from the plurality of vector registers respectively in parallel; and

directing, using a plurality of multiplexors in the circuit and according to the index, the plurality of data elements to the plurality of processing units respectively in parallel.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2020
From: WALLACH, STEVEN JEFFREY
To: MICRON TECHNOLOGY, INC.
Reel/Frame 054757/0303 →
Continuity (2)
Continuation 16356146 · Mar 18, 2019
Related Publication 20210117375A1 · Apr 22, 2021