IP Library Granted Patent US 9,436,475
Granted Patent B2
US 9,436,475 · App. 13/723,981 · Granted Sep 6, 2016

System and method for executing sequential code using a group of threads and single-instruction, multiple-thread processor incorporating the same

Inventors: Gautam Chakrabarti (Santa Clara, CA); Yuan Lin (Santa Clara, CA); Jaydeep Marathe (Santa Clara, CA); Okwan Kwon (West Lafayette, IN); Amit Sabne (West Lafayette, IN)
Assignee: NVIDIA CORPORATION
G06F9/38G06F8/40G06F9/5016G06F9/52G06F9/522G06F12/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,436,475
App. No.
13/723,981
Granted
Sep 6, 2016
Kind
B2
Abstract

A system and method for executing sequential code in the context of a single-instruction, multiple-thread (SIMT) processor. In one embodiment, the system includes: (1) a pipeline control unit operable to create a group of counterpart threads of the sequential code, one of the counterpart threads being a master thread, remaining ones of the counterpart threads being slave threads and (2) lanes operable to: (2 a ) execute certain instructions of the sequential code only in the master thread, corresponding instructions in the slave threads being predicated upon the certain instructions and (2 b ) broadcast branch conditions in the master thread to the slave threads.

Claims (43)

1. A system for executing sequential code, comprising:

a pipeline control unit configured to create a group of counterpart threads of said sequential code in a single instruction, multiple thread (SIMT) processor, one of said counterpart threads being a master thread, remaining ones of said counterpart threads being slave threads; and

lanes configured to:

execute certain instructions of said sequential code only in said master thread, corresponding instructions in said slave threads being predicated upon said certain instructions, and

broadcast branch conditions in said master thread to said slave threads,

wherein all of said lanes of said SIMT processor execute in lock-step.

2. The system as recited in claim 1 wherein local memories associated with lanes executing said slave threads are further configured to store said branch conditions.

3. The system as recited in claim 1 wherein said certain instructions are selected from the group consisting of:

load instructions,

store instructions, and

exception inducing instructions.

4. The system as recited in claim 1 wherein a lane executing said master thread is further configured to broadcast said branch conditions before execution of a branch instruction in said master thread and lanes executing said slave threads further configured to execute corresponding branch instructions in said slave threads only after said lane broadcasts said branch conditions.

5. The system as recited in claim 1 wherein said pipeline control unit is further configured to predicate said corresponding instructions using a condition based on a thread identifier.

6. The system as recited in claim 1 wherein said sequential code is part of a vector operation.

7. A method of executing sequential code, comprising:

creating a group of counterpart threads of said sequential code in a single instruction, multiple thread (SIMT) processor, one of said counterpart threads being a master thread, remaining ones of said counterpart threads being slave threads;

executing certain instructions of said sequential code only in said master thread, corresponding instructions in said slave threads being predicated upon said certain instructions; and

broadcasting branch conditions in said master thread to said slave threads,

wherein lanes of said SIMT processor execute in lock-step.

8. The method as recited in claim 7 further comprising storing said branch conditions in local memories associated with said slave threads.

9. The method as recited in claim 7 wherein said certain instructions are selected from the group consisting of:

load instructions,

store instructions, and

exception inducing instructions.

10. The method as recited in claim 7 wherein said broadcasting is carried out before execution of a branch instruction in said master thread, said method further comprising:

executing corresponding branch instructions in said slave threads only after said broadcasting is carried out.

11. The method as recited in claim 7 wherein said executing comprises predicating said corresponding instructions using a condition based on a thread identifier.

12. The method as recited in claim 7 wherein said sequential code is part of a vector operation.

13. A single-instruction, multiple-thread (SIMT) processor, comprising:

lanes;

local memories associated with corresponding ones of said lanes;

shared memory device by said lanes; and

a pipeline control unit configured to create a group of counterpart threads of said sequential code and cause said group to be executed in said lanes, one of said counterpart threads being a master thread, remaining ones of said counterpart threads being slave threads, said lanes configured to:

execute certain instructions of said sequential code only in said master thread, corresponding instructions in said slave threads being predicated upon said certain instructions, and

broadcast branch conditions in said master thread to said slave threads,

wherein all of said lanes operate of said SIMT processor execute in lock-step.

14. The SIMT processor as recited in claim 13 wherein said local memories associated with lanes executing said slave threads are further configured to store said branch conditions.

15. The SIMT processor as recited in claim 13 wherein said certain instructions are selected from the group consisting of:

load instructions,

store instructions, and

exception inducing instructions.

16. The SIMT processor as recited in claim 13 wherein a lane executing said master thread is further configured to broadcast said branch conditions before execution of a branch instruction in said master thread and lanes executing said slave threads further configured to execute corresponding branch instructions in said slave threads only after said lane broadcasts said branch conditions.

17. The SIMT processor as recited in claim 13 wherein said pipeline control unit is further configured to predicate said corresponding instructions using a condition based on a thread identifier.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2012
From: CHAKRABARTI, GAUTAM; LIN, YUAN; MARATHE, JAYDEEP; KWON, OKWAN; SABNE, AMIT
To: NVIDIA CORPORATION
Reel/Frame 029518/0117 →
Continuity (2)
Provisional Application 61722661 · Nov 5, 2012
Related Publication 20140129812A1 · May 8, 2014