IP Library › Granted Patent US 12,197,374
Granted Patent B2
US 12,197,374 · App. 17/359,321 · Granted Jan 14, 2025

Peer-to-peer link sharing for upstream communications from XPUS to a host processor

Inventors: Rahul Pal (Bangalore, IN); Nayan Amrutlal Suthar (Bangalore, IN); David M. Puffer (Tempe, AZ); Ashok Jagannathan (Bangalore, IN)
Assignee: Intel Corporation
G06F13/4221G06F12/0815G06T1/20G06F2213/0026
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,374
App. No.
17/359,321
Granted
Jan 14, 2025
Kind
B2
Abstract

A processor unit comprising a first controller to couple to a host processing unit over a first link; a second controller to couple to a second processor unit over a second link, wherein the second processor unit is to couple to the host central processing unit via a third link; and circuitry to determine whether to send a cache coherent request to the host central processing unit over the first link or over the second link via the second processing unit.

Claims (35)

1. A processor unit comprising:

a first controller to couple to a host processor unit over a first link;

a second controller to couple to a second processor unit over a second link, wherein the second processor unit is to couple to the host processor unit via a third link; and

circuitry to determine whether to send a cache coherent request to the host processor unit over the first link or over the second link via the second processor unit.

2. The processor unit of claim 1 , wherein the first link and the third link are each links according to a Compute Express Link protocol.

3. The processor unit of claim 1 , wherein the circuitry is to determine whether to send the cache coherent request over the first link or over the second link based on an amount of available upstream bandwidth over the first link.

4. The processor unit of claim 3 , wherein the circuitry is to determine the amount of available upstream bandwidth over the first link based on a number of link credits available.

5. The processor unit of claim 3 , wherein the circuitry is to determine the amount of available upstream bandwidth over the first link based on a raw upstream bandwidth metric.

6. The processor unit of claim 1 , wherein the circuitry is to determine whether to send the cache coherent request over the first link or over the second link based on an amount of available bandwidth over the second link.

7. The processor unit of claim 1 , wherein the circuitry is to determine whether to send the cache coherent request over the first link or over the second link based on an amount of available upstream bandwidth over the third link.

8. The processor unit of claim 7 , wherein the circuitry is to determine the amount of available upstream bandwidth over the third link based on a number of host-bound requests received by the processor unit from the second processor unit, wherein the processor unit is to send the host-bound requests to the host processor unit over the first link.

9. The processor unit of claim 1 , further comprising second circuitry to:

track memory requests received from the second processor unit for memory of the host processor unit; and

respond to snoop requests associated with such memory from the host processor unit.

10. The processor unit of claim 1 , wherein the processor unit and the second processor unit are each graphics processing units.

11. A method comprising:

communicating, by a first processor unit, with a host processor unit over a first link;

communicating, by the first processor unit, with a second processor unit over a second link, wherein the second processor unit is to couple to the host processor unit via a third link; and

determining whether to send a cache coherent request to the host processor unit over the first link or over the second link via the second processor unit.

12. The method of claim 11 , further comprising determining whether to send the cache coherent request over the first link or over the second link based on an amount of available upstream bandwidth over the first link.

13. The method of claim 11 , further comprising determining whether to send the cache coherent request over the first link or over the second link based on an amount of available bandwidth over the second link.

14. The method of claim 11 , further comprising determining whether to send the cache coherent request over the first link or over the second link based on an amount of available upstream bandwidth over the third link.

15. The method of claim 11 , further comprising:

tracking memory requests received from the second processor unit for memory of the host processor unit; and

responding to snoop requests associated with such memory from the host processor unit.

16. A system comprising:

a host processor unit; and

a plurality of processor units, a processor unit of the plurality of processor units coupled to the host processor unit via a first link and to other processor units of the plurality of processor units via a plurality of second links, the other processor units coupled to the host processor unit via a plurality of third links;

wherein the processor unit is to determine whether to send a cache coherent request to the host processor unit over the first link or over one of the second links via one of the other processor units.

17. The system of claim 16 , wherein the processor unit is to determine whether to send the cache coherent request over the first link or over one of the second links based on an amount of available upstream bandwidth over the first link.

18. The system of claim 16 , wherein the processor unit is to determine whether to send the cache coherent request over the first link or over one of the second links based on an amount of available upstream bandwidths over the second links.

19. The system of claim 16 , wherein the processor unit is to send a plurality of cache coherent requests to the host processor unit via a first plurality of the other processor units.

20. The system of claim 16 , wherein the processor unit is to:

track memory requests received from a second processor unit of the plurality of processor units, the memory requests for memory of the host processor unit; and

respond to snoop requests associated with such memory from the host processor unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2021
From: PAL, RAHUL; PUFFER, DAVID M.; JAGANNATHAN, ASHOK; SUTHAR, NAYAN AMRUTLAL
To: INTEL CORPORATION
Reel/Frame 056781/0768 →
Continuity (1)
Related Publication 20210318980A1 · Oct 14, 2021
References Cited (13)
US 10990562B2 · Butcher · 2021 [cited by examiner]
US 11947472B2 · Shah · 2024 [cited by examiner]
US 20090198956A1 · Arimilli · 2009 [cited by examiner]
US 20180143932A1 · Lawless et al. · 2018 [cited by applicant]
US 20180167310A1 · Kamble · 2018 [cited by examiner]
US 20190050365A1 · Kopzon · 2019 [cited by examiner]
US 20200097421A1 · Le · 2020 [cited by examiner]
US 20200192798A1 · Natu · 2020 [cited by applicant]
US 20200327084A1 · Choudhary et al. · 2020 [cited by applicant]
US 20200379930A1 · Brownell et al. · 2020 [cited by applicant]
EPO; Extended European Search Report issued in EP Patent Application No. 22165672.1, dated Sep. 28, 2022; 9 pages. [cited by applicant]
Compute Express Link, “Specification—Oct. 2020, Revision 2.0”, 628 pages. [cited by applicant]
Office Action for European Patent Application No. 22165672.1, mailed on Jul. 26, 2023. [cited by applicant]