IP Library › Granted Patent US 11,995,029
Granted Patent B2
US 11,995,029 · App. 17/428,527 · Granted May 28, 2024

Multi-tile memory management for detecting cross tile access providing multi-tile inference scaling and providing page migration

Inventors: Lakshminarayanan Striramassarma (Folsom, CA); Prasoonkumar Surti (Folsom, CA); Varghese George (Folsom, CA); Ben Ashbaugh (Folsom, CA); Aravindh Anantaraman (Folsom, CA); Valentin Andrei (San Jose, CA); Abhishek Appu (El Dorado Hills, CA); Nicolas Galoppo Von Borries (Portland, OR); Altug Koker (El Dorado Hills, CA); Mike Macpherson (Portland, OR); Subramaniam Maiyuran (Gold River, CA); Nilay Mistry (Bangalore, IN); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Selvakumar Panneer (Portland, OR); Vasanth Ranganathan (El Dorado Hills, CA); Joydeep Ray (Folsom, CA); Ankur Shah (Folsom, CA); Saurabh Tangri (Folsom, CA)
Assignee: Intel Corporation
G06F15/7839G06F7/5443G06F7/575G06F7/588G06F9/3001G06F9/30014G06F9/30036G06F9/3004G06F9/30043G06F9/30047G06F9/30065G06F9/30079G06F9/3887G06F9/5011G06F9/5077G06F12/0215G06F12/0238G06F12/0246G06F12/0607G06F12/0802G06F12/0804G06F12/0811G06F12/0862G06F12/0866G06F12/0871G06F12/0875G06F12/0882G06F12/0888G06F12/0891G06F12/0893G06F12/0895G06F12/0897G06F12/1009G06F12/128G06F15/8046G06F17/16G06F17/18G06T1/20G06T1/60H03M7/46G06F9/3802G06F9/3818G06F9/3867G06F2212/1008G06F2212/1021G06F2212/1044G06F2212/302G06F2212/401G06F2212/455G06F2212/60G06N3/08G06T15/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,995,029
App. No.
17/428,527
Granted
May 28, 2024
Kind
B2
Abstract

Multi-tile Memory Management for Detecting Cross Tile Access, Providing Multi-Tile Inference Scaling with multicasting of data via copy operation, and Providing Page Migration are disclosed herein. In one embodiment, a graphics processor for a multi-tile architecture includes a first graphics processing unit (GPU) having a memory and a memory controller, a second graphics processing unit (GPU) having a memory and a cross-GPU fabric to communicatively couple the first and second GPUs. The memory controller is configured to determine whether frequent cross tile memory accesses occur from the first GPU to the memory of the second GPU in the multi-GPU configuration and to send a message to initiate a data transfer mechanism when frequent cross tile memory accesses occur from the first GPU to the memory of the second GPU.

Claims (30)

1. A graphics processor having a multi-tile architecture, comprising:

a first graphics processing unit (GPU) having a memory and a memory controller;

a second graphics processing unit (GPU) having a memory; and

a cross-GPU fabric to communicatively couple the first and second GPUs, wherein the memory controller is configured to determine whether frequent cross tile memory accesses occur between the first GPU and the second GPU in the multi-GPU configuration and to cause initiation of a data transfer between the memory of the first GPU and the memory of the second GPU when frequent cross tile memory accesses occur between the first GPU and the second GPU, wherein the memory controller is configured to detect transfer patterns automatically including accesses to page N of the memory of the second GPU and to start transferring pages N+1 and N+2 prior to requests for pages N+1 and N+2.

2. The graphics processor of claim 1 , further comprising:

a hardware counter to count cross tile memory accesses between the first GPU and the second GPU.

3. The graphics processor of claim 2 , wherein the memory controller is configured to determine whether frequent cross tile memory accesses occur between the first GPU and the second GPU in the multi-GPU configuration using data from the hardware counter.

4. The graphics processor of claim 3 , wherein the memory controller is configured to cause data that is being accessed frequently by the second GPU to be transferred or copied to the memory of the second GPU.

5. The graphics processor of claim 1 , wherein the memory controller is configured to cause data that is being accessed frequently by the first GPU to be transferred or copied to the memory of the first GPU.

6. The graphics processor of claim 1 , wherein the memory controller is configured to detect transfer patterns automatically including accesses between the first and second GPUs.

7. A graphics processing unit (GPU) of a multi-GPU architecture, comprising:

processing resources to perform graphics operations;

a memory; and

a memory controller, wherein the memory controller is configured to determine whether frequent cross tile memory accesses occur between the GPU and a remote memory of a remote GPU in the multi-GPU configuration and to cause initiation of a data transfer between the memory of the GPU and the remote memory of the remote GPU when frequent cross tile memory accesses occur between the GPU and the remote memory of the remote GPU, wherein the memory controller is configured to detect transfer patterns automatically including accesses to page N of the remote memory and to start transferring pages N+1 and N+2 prior to requests for pages N+1 and N+2.

8. The GPU of claim 7 , further comprising:

a hardware counter to count cross tile memory accesses from the GPU to the remote memory of the remote GPU.

9. The GPU of claim 8 , wherein the memory controller is configured to determine whether frequent cross tile memory accesses occur between the GPU and the remote memory of the remote GPU in the multi-GPU configuration using data from the hardware counter.

10. The GPU of claim 9 , wherein the memory controller is configured to cause data that is being accessed frequently by the remote GPU to be transferred or copied to the remote memory.

11. The GPU of claim 7 , wherein the memory controller is configured to cause data that is being accessed frequently by the GPU to be transferred or copied to the memory of the GPU.

12. The GPU of claim 7 , wherein the memory controller is configured to detect transfer patterns automatically between the GPU and the remote GPU.

13. A computer-implemented method to provide a data transfer mechanism for a multiple GPU configuration, the computer-implemented method comprises:

monitoring cross tile memory accesses from a local GPU to one or more remote GPUs in the multi-GPU configuration;

determining, with a memory controller, whether frequent cross tile memory accesses occur from a local GPU to one or more remote GPUs in the multi-GPU configuration; and

sending a message to initiate the data transfer mechanism between a memory of the local GPU and a remote memory of a remote GPU when frequent cross tile memory accesses occur from the local GPU to the remote memory of the remote GPU in the multi-GPU configuration, wherein the data transfer mechanism to transfer or copy the data that is being accessed frequently by the local GPU to the memory of the local GPU and to local memory of at least one other GPU.

14. The computer-implemented method of claim 13 , further comprising:

receiving, with a graphics driver, the message from the memory controller and to provide the data transfer mechanism in response to receiving the message.

15. The computer-implemented method of claim 13 , wherein the data transfer mechanism accesses a page table to provide a translation of virtual addresses to physical addresses.

16. The computer-implemented method of claim 13 , wherein the data transfer mechanism to transfer or copy the data that is being accessed frequently by the local GPU to multiple tiles or GPUs to enable split frame rendering with a first GPU handling rendering for a first portion of a display and a second GPU handling rendering for a second different portion of the display.

17. The computer-implemented method of claim 13 , further comprising:

performing a page allocation to the memory of the local GPU when a first access to a page in a remote GPU memory occurs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2021
From: STRIRAMASSARMA, LAKSHMINARAYANAN; SURTI, PRASOONKUMAR; GEORGE, VARGHESE; ASHBAUGH, BEN; ANANTARAMAN, ARAVINDH; ANDREI, VALENTIN; APPU, ABHISHEK; GALOPPO VON BORRIES, NICOLAS; KOKER, ALTUG; MACPHERSON, MIKE; MAIYURAN, SUBRAMANIAM; MISTRY, NILAY; OULD-AHMED-VALL, ELMOUSTAPHA; PANNEER, SELVAKUMAR; RANGANATHAN, VASANTH; RAY, JOYDEEP; SHAH, ANKUR; TANGRI, SAURABH
To: INTEL CORPORATION
Reel/Frame 057260/0450 →
Continuity (4)
Provisional Application 62819337 · Mar 15, 2019
Provisional Application 62819435 · Mar 15, 2019
Provisional Application 62819361 · Mar 15, 2019
Related Publication 20220114096A1 · Apr 14, 2022
Cited By (3)
US 12,554,674 US 12,561,277 US 12,737,317