IP Library › Granted Patent US 11,604,649
Granted Patent B2
US 11,604,649 · App. 17/363,561 · Granted Mar 14, 2023

Techniques for efficiently transferring data to a processor

Inventors: Andrew Kerr (Santa Clara, CA); Jack Choquette (Palo Alto, CA); Xiaogang Qiu (San Jose, CA); Omkar Paranjape (Austin, TX); Poornachandra Rao (Cedar Park, TX); Shirish Gadre (Fremont, CA); Steven J. Heinrich (Madison, AL); Manan Patel (San Jose, CA); Olivier Giroux (Santa Clara, CA); Alan Kaatz (Santa Clara, CA)
Assignee: NVIDIA Corporation
G06F9/30043G06F9/3009G06F9/321G06F9/3838G06F9/3871G06F9/522G06F9/542G06F9/544G06F9/546G06F12/0808G06F12/0888G06F9/3004G06F2212/621
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,604,649
App. No.
17/363,561
Granted
Mar 14, 2023
Kind
B2
Abstract

A technique for block data transfer is disclosed that reduces data transfer and memory access overheads and significantly reduces multiprocessor activity and energy consumption. Threads executing on a multiprocessor needing data stored in global memory can request and store the needed data in on-chip shared memory, which can be accessed by the threads multiple times. The data can be loaded from global memory and stored in shared memory using an instruction which directs the data into the shared memory without storing the data in registers and/or cache memory of the multiprocessor during the data transfer.

Claims (21)

1. A GPU execution method comprising:

decoding a memory access instruction format;

in response to the decoding, initiating a memory access to GPU shared memory including selectively bypassing at least one of a register file and a cache; and

tracking completion of the memory access as an asynchronous copy/direct memory access operation.

2. The method of claim 1 , further including writing to shared memory while dynamically selecting whether to bypass each register file at runtime.

3. The method of claim 1 , wherein the memory access includes reading a sector of global memory and writing it to parts of plural sectors in shared memory.

4. The method of claim 1 , wherein the GPU comprises a collection of processing cores connected to share the shared memory local to and associated with the collection of processing cores, the shared memory including a first portion allocated to registers for the processing cores, a second portion allocated to the cache which is an L1 cache for the processing cores, and a third portion.

5. A processing system comprising:

a processor configured to concurrently execute a plurality of threads;

a plurality of registers, each of the registers assigned to executing threads; and

on-chip memory including a cache and a shared memory,

wherein the system is configured to:

decode a memory access instruction format;

in response to the decoding, initiate a memory access to the shared memory including selectively bypassing a register file and the cache; and

track completion of the memory access as an asynchronous copy/direct memory access operation.

6. The processing system of claim 5 , wherein the system is further configured to write to shared memory while dynamically selecting whether to bypass each register file at runtime.

7. The processing system of claim 5 , wherein the memory access includes reading a sector of global memory and writing it to parts of plural sectors in shared memory.

8. The processing system of claim 5 , wherein the cache is configured to allow the plurality of executing threads to access tagged data stored in the cache, and the shared memory is configured to allow the plurality of executing threads to access untagged data stored in the shared memory.

9. The processing system of claim 6 , wherein the shared memory is a software managed cache.

10. The processing system of claim 6 , wherein the cache memory and the shared memory are a unified physical random access memory.

11. The processing system of claim 10 , wherein the unified physical random access memory includes a register file including the plurality of registers dynamically assignable to the executing threads.

Continuity (4)
Division 16712083 · Dec 12, 2019
Provisional Application 62927417 · Oct 29, 2019
Provisional Application 62927511 · Oct 29, 2019
Related Publication 20210326137A1 · Oct 21, 2021