IP Library › Granted Patent US 10,275,247
Granted Patent B2
US 10,275,247 · App. 14/672,156 · Granted Apr 30, 2019

Apparatuses and methods to accelerate vector multiplication of vector elements having matching indices

Inventors: Asit K. Mishra (Hillsboro, OR); Deborah T. Marr (Portland, OR)
Assignee: Intel Corporation
G06F9/30036G06F9/3001G06F9/30018
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,275,247
App. No.
14/672,156
Granted
Apr 30, 2019
Kind
B2
Abstract

Methods and apparatuses relating to accelerating vector multiplication. In one embodiment, an apparatus includes a first buffer to store a first cache line of indices for elements of a first vector, a second buffer to store a second cache line of indices for elements of a second vector, a comparison unit to compare each index of the first cache line of indices with each index of the second cache line of indices, a plurality of multipliers to each multiply an element from the first vector and an element from the second vector for an index match from the comparison unit to produce a product, and an adder to add together the product from each of the plurality of multipliers.

Claims (42)

1. An apparatus comprising:

a first buffer to store a first cache line of indices for elements of a first vector;

a second buffer to store a second cache line of indices for elements of a second vector;

a comparison unit to compare each index of the first cache line of indices with each index of the second cache line of indices;

a plurality of multipliers to each multiply an element from the first vector and an element from the second vector for an index match from the comparison unit to produce a product; and

an adder to add together the product from each of the plurality of multipliers.

2. The apparatus of claim 1 , wherein a first streamer is to provide an index and its element from a data storage device to the first buffer and a second streamer is to provide an index and its element from the data storage device to the second buffer.

3. The apparatus of claim 1 , wherein the indices of the first cache line and the second cache line are not in index order.

4. The apparatus of claim 1 , wherein the comparison unit is to compare each index of the first cache line of indices with each index of the second cache line of indices in a single clock cycle of a processor.

5. The apparatus of claim 1 , wherein the first cache line of indices for elements of the first vector also includes each index's element.

6. The apparatus of claim 1 , wherein the plurality of multipliers are a plurality of multiplier-accumulator units.

7. The apparatus of claim 1 , further comprising logic to notify a requesting processor core that operations on all elements of the first vector and the second vector are completed.

8. The apparatus of claim 1 , wherein the comparison unit is to return each index of the first cache line to the first buffer and each index of the second cache line to the second buffer for non-matching indices.

9. A method comprising:

retrieving a first cache line of indices for elements of a first vector stored in a first buffer;

retrieving a second cache line of indices for elements of a second vector stored in a second buffer;

comparing each index of the first cache line of indices with each index of the second cache line of indices with a comparison unit;

multiplying an element from the first vector and an element from the second vector for each of a plurality of multipliers for an index match from the comparison unit to produce a product; and

adding together the product from each of the plurality of multipliers with an adder.

10. The method of claim 9 , further comprising:

providing an index and its element from a data storage device to the first buffer with a first streamer, and

providing an index and its element from the data storage device to the second buffer with a second streamer.

11. The method of claim 9 , wherein the indices of the first cache line and the second cache line are not in index order.

12. The method of claim 9 , wherein the comparing is in a single clock cycle of a processor.

13. The method of claim 9 , wherein the first cache line of indices for elements of the first vector also includes each index's element.

14. The method of claim 9 , wherein the plurality of multipliers are a plurality of multiplier-accumulator units.

15. The method of claim 9 , further comprising notifying a requesting processor core that operations on all elements of the first vector and the second vector are completed.

16. The method of claim 9 , further comprising returning each index of the first cache line to the first buffer and each index of the second cache line to the second buffer for non-matching indices.

17. A system comprising:

a data storage device to store a first vector and a second vector;

a first buffer to store a first cache line of indices for elements of the first vector from the data storage device;

a second buffer to store a second cache line of indices for elements of the second vector from the data storage device;

a comparison unit to compare each index of the first cache line of indices with each index of the second cache line of indices;

a plurality of multipliers to each multiply an element from the first vector and an element from the second vector for an index match from the comparison unit to produce a product; and

an adder to add together the product from each of the plurality of multipliers.

18. The system of claim 17 , wherein a first streamer is to provide an index and its element from the data storage device to the first buffer and a second streamer is to provide an index and its element from the data storage device to the second buffer.

19. The system of claim 17 , wherein the indices of the first cache line and the second cache line are not in index order.

20. The system of claim 17 , wherein the comparison unit is to compare each index of the first cache line of indices with each index of the second cache line of indices in a single clock cycle of a processor.

21. The system of claim 17 , wherein the first cache line of indices for elements of the first vector also includes each index's element.

22. The system of claim 17 , wherein the plurality of multipliers are a plurality of multiplier-accumulator units.

23. The system of claim 17 , further comprising logic to notify a requesting processor core that operations on all elements of the first vector and the second vector are completed.

24. The system of claim 17 , wherein the comparison unit is to return each index of the first cache line to the first buffer and each index of the second cache line to the second buffer for non-matching indices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2015
From: MISHRA, ASIT K.; MARR, DEBORAH T.
To: INTEL CORPORATION
Reel/Frame 036179/0437 →
Continuity (1)
Related Publication 20160283240A1 · Sep 29, 2016
Cited By (2)
US 12,254,061 US 12,461,799