IP Library Granted Patent US 10,990,394
Granted Patent B2
US 10,990,394 · App. 15/719,109 · Granted Apr 27, 2021

Systems and methods for mixed instruction multiple data (xIMD) computing

Inventor: Jeffrey L. Nye (Austin, TX)
Assignee: Intel Corporation
G06F9/30058G06F9/3887G06F9/3891
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,990,394
App. No.
15/719,109
Granted
Apr 27, 2021
Kind
B2
Abstract

An integrated circuit may include a mixed instruction multiple data (xIMD) computing system. The xIMD computing system may include a plurality of data processors, each data processor representative of a lane of a single instruction multiple data (SIMD) computing system, wherein the plurality of data processors are configured to use a first dominant lane for instruction execution and to fork a second dominant lane when a data dependency instruction that does not share a taken/not-taken state with the first dominant lane is encountered during execution of a program by the xIMD computing system.

Claims (30)

1. An integrated circuit, comprising:

a mixed instruction multiple data (xIMD) computing system, comprising:

a plurality of data processors, each data processor representative of a lane of a single instruction multiple data (SIMD) computing system, wherein the plurality of data processors are configured to use a first dominant lane for instruction execution and to fork a second dominant lane when a data dependency instruction that does not share a taken/not-taken state with the first dominant lane is encountered during execution of a program by the xIMD computing system.

2. The integrated circuit of claim 1 , wherein a first set of the plurality of data processors are configured to use a first program counter controlled by the first dominant lane, and wherein a second set of the plurality of data processors are configured to use a second program counter controlled by the second dominant lane during execution of the dependency instruction.

3. The integrated circuit of claim 2 , wherein the first set of the plurality of data processors comprise data processors representative of lanes sharing the taken/not-taken state with the first dominant lane.

4. The integrated circuit of claim 2 , wherein the second set of the plurality of data processors comprise data processors representative of lanes not sharing the taken/not-taken state with the first dominant lane.

5. The integrated circuit of claim 1 , wherein the xIMD computing system comprises a control processor communicatively coupled to the plurality of data processors and configured to perform a loop function, a result mask generation, a coefficient distribution, a reduction result gathering from outputs of the plurality of data processors, or a combination thereof, for the xIMD computing system.

6. The integrated circuit of claim 5 , wherein the control processor comprises a dual issue processor.

7. The integrated circuit of claim 1 , wherein the plurality of data processors are configured to perform SIMD processing, multiple instruction multiple data (MIMD) processing, or a combination thereof, via a SYNC instruction, a DIS instruction, a REMAP instruction, or a combination thereof.

8. The system of claim 1 , comprising a cluster having a plurality of xIMD systems, wherein the xIMD system is included in the cluster.

9. The integrated circuit of claim 1 , wherein the xIMD system is included in a field programmable gate array (FPGA).

10. A system, comprising:

a processor configured to:

receive circuit design data for a mixed instruction multiple data (xIMD) computing system, the xIMD computing system comprising:

a plurality of data processors, each data processor representative of a lane of a single instruction multiple data (SIMD) computing system, wherein the plurality of data processors are configured to use a first dominant lane for instruction execution and to fork a second dominant lane when a data dependency instruction that does not share a taken/not-taken state with the first dominant lane is encountered during execution of a program by the xIMD computing system; and

implement the circuit design data by generating the xIMD computing system as a bitstream.

11. The system of claim 10 , wherein a first set of the plurality of data processors are configured to use a first program counter controlled by the first dominant lane, and wherein a second set of the plurality of data processors are configured to use a second program counter controlled by the second dominant lane during execution of the dependency instruction.

12. The system of claim 10 , wherein the first set of the plurality of data processors comprise data processors representative of lanes sharing the taken/not-taken state with the first dominant lane, and wherein the second set of the plurality of data processors comprise data processors representative of lanes not sharing the taken/not-taken state with the first dominant lane.

13. The system of claim 12 , wherein the xIMD computing system comprises a control processor communicatively coupled to the plurality of data processors and configured to perform a loop function, a result mask generation, a coefficient distribution, a reduction result gathering from outputs of the plurality of data processors, or a combination thereof, for the xIMD computing system.

14. The system of claim 13 , wherein the plurality of data processors are configured to perform SIMD processing, multiple instruction multiple data (MIMD) processing, or a combination thereof, via a SYNC instruction, a DIS instruction, a REMAP instruction, or a combination thereof.

15. The system of claim 10 , wherein the circuit design data comprises a cluster having a plurality of xIMD systems, wherein the xIMD system is included in the cluster, and wherein the processor is configured to implement the circuit design data by generating the cluster.

16. The system of claim 10 , wherein the processor is configured to receive a program logic file comprising software instructions to be executed by a second processor, wherein the second processor is configured to use the xIMD system for data processing, and wherein the processor is configured to generate an executable file for execution by the second processor based on the program logic file.

17. A method, comprising:

receiving a plurality of data inputs;

executing a program via a mixed instruction multiple data (xIMD) computing system to process the plurality of data inputs, wherein the xIMD computing system comprises:

a plurality of data processors, each data processor representative of a lane of a single instruction multiple data (SIMD) computing system, wherein the plurality of data processors are configured to use a first dominant lane for instruction execution and to fork a second dominant lane when a data dependency instruction that does not share a taken/not-taken state with the first dominant lane is encountered during execution of a program by the xIMD computing system; and

outputting program results to a user based on the executing of the program via the xIMD computing system.

18. The method of claim 17 , wherein a first set of the plurality of data processors are configured to use a first program counter controlled by the first dominant lane, and wherein a second set of the plurality of data processors are configured to use a second program counter controlled by the second dominant lane during execution of the dependency instruction.

19. The method of claim 17 , wherein the first set of the plurality of data processors comprise data processors representative of lanes sharing the taken/not-taken state with the first dominant lane, and wherein the second set of the plurality of data processors comprise data processors representative of lanes not sharing the taken/not-taken state with the first dominant lane.

20. The method of claim 17 , wherein the xIMD computing system comprises a control processor communicatively coupled to the plurality of data processors and configured to perform a loop function, a result mask generation, a coefficient distribution, a reduction result gathering from outputs of the plurality of data processors, or a combination thereof, for the xIMD computing system.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2022
From: INTEL CORPORATION
To: TAHOE RESEARCH, LTD.
Reel/Frame 061175/0176 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2017
From: NYE, JEFFREY L.
To: INTEL CORPORATION
Reel/Frame 043780/0196 →
Continuity (1)
Related Publication 20190095208A1 · Mar 28, 2019