IP Library Granted Patent US 8,631,224
Granted Patent B2
US 8,631,224 · App. 11/854,630 · Granted Jan 14, 2014

SIMD dot product operations with overlapped operands

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,631,224
App. No.
11/854,630
Granted
Jan 14, 2014
Kind
B2
Abstract

A data processing system includes a plurality of general purpose registers, and processor circuitry for executing one or more instructions, including a vector dot product instruction for simultaneously performing at least two dot products. The vector dot product instruction identifies a first and second source register, each for storing a plurality of vector elements, where a first dot product is to be performed between a first subset of vector elements of the first source register and a first subset of vector elements of the second source register, and a second dot product is to be performed between a second subset of vector elements of the first source register and a second subset of vector elements of the second source register. The first and second subsets of the second source register are different and at least two vector elements of the first and second subsets of the second source register overlap.

Claims (27)

1. A data processing system, comprising:

a plurality of general purpose registers;

processor circuitry for executing one or more instructions, the one or more instructions comprising a vector dot product instruction for simultaneously performing at least two dot products, the vector dot product instruction identifying a first source register from the plurality of general purpose registers, and a second source register from the plurality of general purpose registers, each of the first source register and the second source register for storing a plurality of vector elements, wherein a first dot product of the at least two dot products is to be performed between a first subset of vector elements of the first source register and a first subset of vector elements of the second source register, and a second dot product of the at least two dot products is to be performed between a second subset of vector elements of the first source register and a second subset of vector elements of the second source register, wherein the first and second subsets of the second source register are different and wherein at least two vector elements of the first and second subsets of the second source register overlap, and wherein each of the at least two overlapping vector elements of the second source register is used in both the first dot product and the second dot product with different vector elements of the first source register.

2. The data processing system of claim 1 , wherein the vector dot product instruction further identifies a destination register for storing a result of the first dot product and a result of the second dot product.

3. The data processing system of claim 1 , wherein the processor circuitry further comprises an accumulator, and wherein the vector dot product instruction further identifies a destination register for storing a sum of a result of the first dot product and a first value of the accumulator and a sum of a result of the second dot product and a second value of the accumulator.

4. The data processing system of claim 1 , wherein the first and second subsets of the first source register are a same subset.

5. The data processing system of claim 1 , wherein the first subset of vector elements of the first source register corresponds to same vector element locations as the first subset of vector elements of the second source register.

6. The data processing system of claim 1 , wherein the vector dot product instruction further indicates an offset for use in at least indicating which vector elements of the first source register are to be included in the first subset of vector elements of the first source register.

7. The data processing system of claim 6 , wherein the vector dot product instruction further indicates a second offset for use in at least indicating which vector elements of the second source register are to be included in the first subset of vector elements of the second source register.

8. The data processing system of claim 1 , wherein the vector dot product instruction further indicates an offset for use in at least indicating which vector elements of the second source register are to be included in the first subset of vector elements of the second source register.

9. The data processing system of claim 1 , wherein the first subset of vector elements of the first source register and the second subset of vector elements of the first source register have overlapping vector elements, wherein each of the overlapping vector elements of the first source register is used in both the first dot product and the second dot product with different vector elements of the second source register.

10. A data processing system, comprising:

a plurality of general purpose registers; and

processor circuitry for executing one or more instructions, the one or more instructions comprising a vector dot product instruction for simultaneously performing at least two dot products, the vector dot product instruction identifying a first source register from the plurality of general purpose registers, and a second source register from the plurality of general purpose registers, each of the first source register and the second source register for storing a plurality of vector elements, wherein a first dot product of the at least two dot products is to be performed between a first subset of five vector elements of the first source register and a first subset of five vector elements of the second source register, and a second dot product of the at least two dot products is to be performed between a second subset of five vector elements of the first source register and a second subset of five vector elements of the second source register, wherein four vector elements of the first and second subsets of the second source register overlap, and wherein each of the four overlapping vector elements of the second source register is used in both the first dot product and the second dot product with different vector elements of the first source register.

11. The data processing system of claim 10 , wherein the vector dot product instruction further identifies a destination register for storing a result of the first dot product and a result of the second dot product.

12. The data processing system of claim 10 , wherein the processor circuitry further comprises an accumulator, and wherein the vector dot product instruction further identifies a destination register for storing a sum of a result of the first dot product and a first value of the accumulator and a sum of the a result of the second dot product and a second value of the accumulator.

13. The data processing system of claim 10 , wherein the first and second subsets of the first source register are a same subset.

14. The data processing system of claim 10 , wherein the first subset of vector elements of the first source register corresponds to same vector element locations as the first subset of vector elements of the second source register.

15. The data processing system of claim 10 , wherein each of the first and second source registers identified by the vector dot product instruction is for storing eight vector elements, and wherein the processor circuitry comprises ten multipliers, five of which for performing the first dot product and the other five of which for performing the second dot product.

16. The data processing system of claim 10 , wherein the vector dot product instruction further indicates an offset for use in at least indicating which vector elements of the first or second source register are to be included in the first subset of vector elements of the first or second source register.

17. A method for performing simultaneous dot product operations, comprising:

providing a plurality of general purpose registers; and

providing processor circuitry for executing one or more instructions, the one or more instructions comprising a vector dot product instruction for simultaneously performing at least two dot products, the vector dot product instruction identifying a first source register from the plurality of general purpose registers, and a second source register from the plurality of general purpose registers, each of the first source register and the second source register for storing a plurality of vector elements, wherein a first dot product of the at least two dot products is to be performed between a first subset of vector elements of the first source register and a first subset of vector elements of the second source register, and a second dot product of the at least two dot products is to be performed between a second subset of vector elements of the first source register and a second subset of vector elements of the second source register, wherein the first and second subsets of the second source register are different and wherein at least two vector elements of the first and second subsets of the second source register overlap, and wherein each of the at least two overlapping vector elements of the second source register is used in both the first dot product and the second dot product with different vector elements of the first source register.

18. The method of claim 17 , wherein the vector dot product instruction further identifies a destination register for storing a result of the first dot product and a result of the second dot product.

19. The method of claim 17 , wherein the processor circuitry further comprises an accumulator, and wherein the vector dot product instruction further identifies a destination register for storing a sum of a result of the first dot product and a first value of the accumulator and a sum of the a result of the second dot product and a second value of the accumulator.

20. The method of claim 17 , wherein the first and second subsets of the first source register are a same subset.

21. The method of claim 17 , wherein the vector dot product instruction further indicates an offset for use in at least indicating which vector elements of the first or second source register are to be included in the first subset of vector elements of the first or second source register.

Assignments (22)
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE APPLICATION 11759915 AND REPLACE IT WITH APPLICATION 11759935 PREVIOUSLY RECORDED ON REEL 040925 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Feb 17, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: NXP, B.V. F/K/A FREESCALE SEMICONDUCTOR, INC.
Reel/Frame 052917/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE APPLICATION 11759915 AND REPLACE IT WITH APPLICATION 11759935 PREVIOUSLY RECORDED ON REEL 040928 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Jan 17, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: NXP B.V.
Reel/Frame 052915/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE APPLICATION 11759915 AND REPLACE IT WITH APPLICATION 11759935 PREVIOUSLY RECORDED ON REEL 037486 FRAME 0517. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT AND ASSUMPTION OF SECURITY INTEREST IN PATENTS. Recorded Dec 10, 2019
From: CITIBANK, N.A.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 053547/0421 →
RELEASE OF SECURITY INTEREST Recorded Sep 10, 2019
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: NXP B.V.
Reel/Frame 050744/0097 →
CORRECTIVE ASSIGNMENT TO CORRECT THE TO CORRECT THE APPLICATION NO. FROM 13,883,290 TO 13,833,290 PREVIOUSLY RECORDED ON REEL 041703 FRAME 0536. ASSIGNOR(S) HEREBY CONFIRMS THE THE ASSIGNMENT AND ASSUMPTION OF SECURITY INTEREST IN PATENTS.. Recorded Feb 20, 2019
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: SHENZHEN XINGUODU TECHNOLOGY CO., LTD.
Reel/Frame 048734/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE NATURE OF CONVEYANCE PREVIOUSLY RECORDED AT REEL: 040632 FRAME: 0001. ASSIGNOR(S) HEREBY CONFIRMS THE MERGER AND CHANGE OF NAME. Recorded Sep 21, 2017
From: FREESCALE SEMICONDUCTOR INC.
To: NXP USA, INC.
Reel/Frame 044209/0047 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE PATENTS 8108266 AND 8062324 AND REPLACE THEM WITH 6108266 AND 8060324 PREVIOUSLY RECORDED ON REEL 037518 FRAME 0292. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT AND ASSUMPTION OF SECURITY INTEREST IN PATENTS. Recorded Feb 1, 2017
From: CITIBANK, N.A.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 041703/0536 →
CHANGE OF NAME Recorded Nov 8, 2016
From: FREESCALE SEMICONDUCTOR, INC.
To: NXP USA, INC.
Reel/Frame 040632/0001 →
RELEASE OF SECURITY INTEREST Recorded Nov 7, 2016
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: NXP B.V.
Reel/Frame 040928/0001 →
RELEASE OF SECURITY INTEREST Recorded Sep 21, 2016
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: NXP, B.V., F/K/A FREESCALE SEMICONDUCTOR, INC.
Reel/Frame 040925/0001 →
SUPPLEMENT TO THE SECURITY AGREEMENT Recorded Jun 16, 2016
From: FREESCALE SEMICONDUCTOR, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 039138/0001 →
ASSIGNMENT AND ASSUMPTION OF SECURITY INTEREST IN PATENTS Recorded Jan 13, 2016
From: CITIBANK, N.A.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 037518/0292 →
ASSIGNMENT AND ASSUMPTION OF SECURITY INTEREST IN PATENTS Recorded Jan 12, 2016
From: CITIBANK, N.A.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 037486/0517 →
PATENT RELEASE Recorded Dec 21, 2015
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: FREESCALE SEMICONDUCTOR, INC.
Reel/Frame 037354/0704 →
PATENT RELEASE Recorded Dec 21, 2015
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: FREESCALE SEMICONDUCTOR, INC.
Reel/Frame 037356/0143 →
PATENT RELEASE Recorded Dec 21, 2015
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: FREESCALE SEMICONDUCTOR, INC.
Reel/Frame 037356/0553 →
SECURITY AGREEMENT Recorded Nov 6, 2013
From: FREESCALE SEMICONDUCTOR, INC.
To: CITIBANK, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 031591/0266 →
SECURITY AGREEMENT Recorded Jun 18, 2013
From: FREESCALE SEMICONDUCTOR, INC.
To: CITIBANK, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 030633/0424 →
SECURITY AGREEMENT Recorded May 13, 2010
From: FREESCALE SEMICONDUCTOR, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 024397/0001 →
SECURITY AGREEMENT Recorded Mar 15, 2010
From: FREESCALE SEMICONDUCTOR, INC.
To: CITIBANK, N.A.
Reel/Frame 024085/0001 →
SECURITY AGREEMENT Recorded Feb 16, 2008
From: FREESCALE SEMICONDUCTOR, INC.
To: CITIBANK, N.A.
Reel/Frame 020518/0215 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2007
From: MOYER, WILLIAM C.
To: FREESCALE SEMICONDUCTOR, INC.
Reel/Frame 019822/0026 →