IP Library Granted Patent US 12,499,071
Granted Patent B2
US 12,499,071 · App. 18/680,970 · Granted Dec 16, 2025

System decoder for training accelerators

Inventors: Francesc Guim Bernat (Barcelona, ES); Da-Ming Chiang (San Jose, CA); Kshitij A. Doshi (Tempe, AZ); Suraj Prabhakaran (Aachen, DE); Mark A. Schmisseur (Phoenix, AZ)
Assignee: Intel Corporation
G06F13/4068G06F9/45533G06F9/5027G06F9/54G06F13/362G06F13/4265G06F13/4282G06N3/02G06N3/04G06N3/045G06N3/08G06F2213/0026
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,071
App. No.
18/680,970
Filed
May 31, 2024
Granted
Dec 16, 2025
Kind
B2
Examiner
SUN, MICHAEL
Art Unit
2183
USPC
712/220
Abstract

There is disclosed an example of an artificial intelligence (AI) system, including: a first hardware platform; a fabric interface configured to communicatively couple the first hardware platform to a second hardware platform; a processor hosted on the first hardware platform and programmed to operate on an AI problem; and a first training accelerator, including: an accelerator hardware; a platform inter-chip link (ICL) configured to communicatively couple the first training accelerator to a second training accelerator on the first hardware platform without aid of the processor; a fabric ICL to communicatively couple the first training accelerator to a third training accelerator on a second hardware platform without aid of the processor; and a system decoder configured to operate the fabric ICL and platform ICL to share data of the accelerator hardware between the first training accelerator and second and third training accelerators without aid of the processor.

Claims (99)

1 . Accelerator hardware for use in a computing platform, the computing platform comprising multiple central processing units (CPUs) and forwarding hardware, the accelerator hardware comprising:

at least one graphics processing unit (GPU) and at least one other GPU, the at least one GPU and the at least one other GPU being configurable to execute respective multiple virtualized instances, the respective multiple virtualized instances being configurable to generate respective data;

wherein:

when the computing platform is in operation:

the at least one GPU is configurable to transfer, via the forwarding hardware, the respective data of a certain one of the respective multiple virtualized instances of the at least one GPU to another certain one of the respective multiple virtualized instances of the at least one other GPU;

data transfer between the at least one GPU and the at least one other GPU is to be in accordance with at least one communication protocol;

the at least one GPU and at least one of the multiple CPUs are configurable for communication with each other in accordance with at least one other communication protocol; and

the at least one communication protocol and the at least one other communication protocol are different from each other, at least in part.

2 . The accelerator hardware of claim 1 , wherein:

the respective data are associated with cloud computing.

3 . The accelerator hardware of claim 1 , wherein:

the computing platform comprises a physical chassis; and

the physical chassis comprises the multiple CPUs and the forwarding hardware.

4 . The accelerator hardware of claim 1 , wherein:

the at least one other communication protocol comprises Peripheral Component Interconnect Express protocol.

5 . At least one non-transitory machine-readable storage medium storing instructions for being executed, at least in part, by accelerator hardware, the accelerator hardware to be used in a computing platform, the computing platform comprising multiple central processing units (CPUs) and forwarding hardware, the accelerator hardware comprising at least one graphics processing unit (GPU) and at least one other GPU, the instructions, when executed, at least in part, by the accelerator hardware, resulting in performance of operations comprising:

configuring the at least one GPU and the at least one other GPU to execute respective multiple virtualized instances, the respective multiple virtualized instances being configurable to generate respective data;

wherein:

when the computing platform is in operation:

the at least one GPU is configurable to transfer, via the forwarding hardware, the respective data of a certain one of the respective multiple virtualized instances of the at least one GPU to another certain one of the respective multiple virtualized instances of the at least one other GPU;

data transfer between the at least one GPU and the at least one other GPU is to be in accordance with at least one communication protocol;

the at least one GPU and at least one of the multiple CPUs are configurable for communication with each other in accordance with at least one other communication protocol; and

the at least one communication protocol and the at least one other communication protocol are different from each other, at least in part.

6 . The at least one non-transitory machine-readable storage medium of claim 5 , wherein:

the respective data are associated with cloud computing.

7 . The at least one non-transitory machine-readable storage medium of claim 5 , wherein:

the computing platform comprises a physical chassis; and

the physical chassis comprises the multiple CPUs and the forwarding hardware.

8 . The at least one non-transitory machine-readable storage medium of claim 5 , wherein:

the at least one other communication protocol comprises Peripheral Component Interconnect Express protocol.

9 . Computing platform comprising:

multiple central processing units (CPUs);

forwarding hardware; and

accelerator hardware comprising at least one graphics processing unit (GPU) and at least one other GPU, the at least one GPU and the at least one other GPU being configurable to execute respective multiple virtualized instances, the respective multiple virtualized instances being configurable to generate respective data;

wherein:

when the computing platform is in operation:

the at least one GPU is configurable to transfer, via the forwarding hardware, the respective data of a certain one of the respective multiple virtualized instances of the at least one GPU to another certain one of the respective multiple virtualized instances of the at least one other GPU;

data transfer between the at least one GPU and the at least one other GPU is to be in accordance with at least one communication protocol;

the at least one GPU and at least one of the multiple CPUs are configurable for communication with each other in accordance with at least one other communication protocol; and

the at least one communication protocol and the at least one other communication protocol are different from each other, at least in part.

10 . The computing platform of claim 9 , wherein:

the respective data are associated with cloud computing.

11 . The computing platform of claim 9 , wherein:

the computing platform comprises a physical chassis; and

the physical chassis comprises the multiple CPUs and the forwarding hardware.

12 . The computing platform of claim 9 , wherein:

the at least one other communication protocol comprises Peripheral Component Interconnect Express protocol.

13 . At least one non-transitory machine-readable storage medium storing instructions for being executed, at least in part, by a computing platform, the computing platform comprising multiple central processing units (CPUs), forwarding hardware, and accelerator hardware, the accelerator hardware comprising at least one graphics processing unit (GPU) and at least one other GPU, the instructions when executed, at least in part, by the computing platform resulting in performance of operations comprising:

configuring the at least one GPU and the at least one other GPU to execute respective multiple virtualized instances, the respective multiple virtualized instances being configurable to generate respective data;

wherein:

when the computing platform is in operation:

the at least one GPU is configurable to transfer, via the forwarding hardware, the respective data of a certain one of the respective multiple virtualized instances of the at least one GPU to another certain one of the respective multiple virtualized instances of the at least one other GPU;

data transfer between the at least one GPU and the at least one other GPU is to be in accordance with at least one communication protocol;

the at least one GPU and at least one of the multiple CPUs are configurable for communication with each other in accordance with at least one other communication protocol; and

the at least one communication protocol and the at least one other communication protocol are different from each other, at least in part.

14 . The at least one non-transitory machine-readable storage medium of claim 13 , wherein:

the respective data are associated with cloud computing.

15 . The at least one non-transitory machine-readable storage medium of claim 13 , wherein:

the computing platform comprises a physical chassis; and

the physical chassis comprises the multiple CPUs and the forwarding hardware.

16 . The at least one non-transitory machine-readable storage medium of claim 13 , wherein:

the at least one other communication protocol comprises Peripheral Component Interconnect Express protocol.

17 . Computing platform for use with accelerator hardware and at least one network, the computing platform being for use in providing, when the computing platform is in operation, at least one service, the computing platform comprising:

at least one central processing unit (CPU) core;

physical accelerator logic;

multiple physical network interface controllers (NICs); and

forwarding hardware;

wherein:

the at least one CPU core, the accelerator hardware, the physical accelerator logic, and the multiple physical NICs are comprised in multiple circuit boards to be communicatively coupled together in the computing platform;

the multiple physical NICs are for use in network traffic communication via the at least one network;

the accelerator hardware comprises accelerator circuitry comprising field programmable gate array circuitry, the accelerator circuitry being configurable to execute multiple virtualized instances for use in association with implementation of neural network-related operations for use in association with the providing of the at least one service; and

when the computing platform is in the operation:

the accelerator circuitry is configurable for communication, using the forwarding hardware and at least one communication protocol, with the physical accelerator logic;

the physical accelerator logic is also configurable for communication, using at least one other communication protocol, with the at least one CPU core; and

the least one communication protocol and the at least one other communication protocol are different from each other, at least in part.

18 . The computing platform of claim 17 , wherein:

the at least one service is associated, at least in part, with cloud computing.

19 . At least one non-transitory machine-readable storage medium storing instructions for being executed, at least in part, by a computing platform, the computing platform to be used with accelerator hardware and at least one network, the computing platform being configurable for use in providing, when the computing platform is in operation, at least one service, the computing platform comprising at least one central processing unit (CPU) core, physical accelerator logic, multiple physical network interface controllers (NICs), and forwarding hardware, the accelerator hardware comprising accelerator circuitry that comprises field programmable gate array circuitry, the instructions when executed, at least in part, by the computing platform resulting in performance of operations comprising:

configuring the accelerator circuitry to execute multiple virtualized instances for use in association with implementation of neural network-related operations for use in association with the providing of the at least one service;

wherein:

the at least one CPU core, the accelerator hardware, the physical accelerator logic, and the multiple physical NICs are comprised in multiple circuit boards to be communicatively coupled together in the computing platform;

the multiple physical NICs are for use in network traffic communication via the at least one network; and

when the computing platform is in the operation:

the accelerator circuitry is configurable for communication, using the forwarding hardware and at least one communication protocol, with the physical accelerator logic;

the physical accelerator logic is also configurable for communication, using at least one other communication protocol, with the at least one CPU core; and

the least one communication protocol and the at least one other communication protocol are different from each other, at least in part.

20 . The at least one non-transitory machine-readable storage medium of claim 19 , wherein:

the at least one service is associated, at least in part, with cloud computing.

21 . A method implemented using a computing platform, the computing platform being for use with accelerator hardware and at least one network, the computing platform being for use in providing, when the computing platform is in operation, at least one service, the computing platform comprising at least one central processing unit (CPU) core, physical accelerator logic, multiple physical network interface controllers (NICs), and forwarding hardware, the accelerator hardware comprising accelerator circuitry that comprises field programmable gate array circuitry, the method comprising:

configuring the accelerator circuitry to execute multiple virtualized instances for use in association with implementation of neural network-related operations for use in association with the providing of the at least one service;

wherein:

the at least one CPU core, the accelerator hardware, the physical accelerator logic, and the multiple physical NICs are comprised in multiple circuit boards to be communicatively coupled together in the computing platform;

the multiple physical NICs are for use in network traffic communication via the at least one network; and

when the computing platform is in the operation:

the accelerator circuitry is configurable for communication, using the forwarding hardware and at least one communication protocol, with the physical accelerator logic;

the physical accelerator logic is also configurable for communication, using at least one other communication protocol, with the at least one CPU core; and

the least one communication protocol and the at least one other communication protocol are different from each other, at least in part.

22 . The method of claim 21 , wherein:

the at least one service is associated, at least in part, with cloud computing.

Continuity (5)
Continuation 17973268 · Oct 25, 2022
Continuation 17584092 · Jan 25, 2022
Continuation 17125439 · Dec 17, 2020
Continuation 15848218 · Dec 20, 2017
Related Publication 20240320179A1 · Sep 26, 2024
References Cited (42)
US 7664931B2 · Erforth · 2010 [cited by examiner]
US 8738860B1 · Griffin et al. · 2014 [cited by applicant]
US 10310760B1 · Dreier · 2019 [cited by examiner]
US 11216314B2 · Harwood et al. · 2022 [cited by applicant]
US 11282161B2 · Ray et al. · 2022 [cited by applicant]
US 11315013B2 · Savic et al. · 2022 [cited by applicant]
US 20060031447A1 · Holt et al. · 2006 [cited by applicant]
US 20170213154A1 · Hammond et al. · 2017 [cited by applicant]
US 20170323418A1 · Dror · 2017 [cited by examiner]
US 20190069437A1 · Adrian et al. · 2019 [cited by applicant]
US 20190102311A1 · Gupta et al. · 2019 [cited by applicant]
US 20190205745A1 · Sridharan et al. · 2019 [cited by applicant]
US 20190286972A1 · El Husseini · 2019 [cited by examiner]
US 20190312772A1 · Zhao et al. · 2019 [cited by applicant]
US 20200311559A1 · Chattopadhyay et al. · 2020 [cited by applicant]
US 20230094125A1 · Rogers et al. · 2023 [cited by applicant]
Final Office Action from U.S. Appl. No. 17/973,268 notified Dec. 21, 2023, 25 pgs. [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 17/973,268 notified Jun. 7, 2023, 23 pgs. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/848,218, dated Oct. 29, 2021. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/125,439, dated Nov. 1, 2021. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/584,092, dated Oct. 14, 2022. [cited by applicant]
Notice of Allowance from U.S. Appl. No. 17/973,268 notified Mar. 18, 2024, 10 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 15/848,218, dated Jun. 24, 2021. [cited by applicant]
Cutress, Ian, “NVIDIA's DGX-2: Sixteen Tesla V100s, 30 TB of NVMe, only $400K”, Anandtech, Mar. 27, 2018, retrieved online via https://www.anandtech.com/show/12587/nvidias-dgx2-sixteen-v100-gpus-30-tb-of-nvme-only-400k,… [cited by applicant]
Gomperts, et al., “Development and implementation of parameterized FPGA-based general purpose neural networks for online applications”, IEEE Transactions on Industrial Informatics 7.1: 78-89 (year: 2010). [cited by applicant]
Intel, “Accelerate Your Data Center with Intel FPGAs”, Solution Brief, Cloud Service Providers, Intel FPGAs; Case Study; 3 pages. [cited by applicant]
Intel Corp., “Intel FPGA Programmable Accelaration Card (Intel”, Platform Selector Guide, 1 page. [cited by applicant]
Intel Corp., “Intel FPGA Programmable Acceleration Card (PAC) D5005”, Product Brief; 1 page. [cited by applicant]
Intel Corp., “Intel FPGA Programmable Acceleration Card (PAC) D5005”, ss-1088-1.0; 1 page. [cited by applicant]
Intel Corp., “Intel FPGA Programmable Acceleration CArd D5005”, retrieved online via https://www.intel.com/content/www/us/en/programmable/products/boards_and_kits/dev-kits/altera/intel-fpga-pac-d5005/overview.html; 2 pa… [cited by applicant]
Intel Corp., “Intel OpenVINO with FPGA Support Through the Intel FPGA Deep Learning Acceleration Suite”, Intel FPGA Deep Learning Acceleration Suite enables Intel FPGAs for accelarated AI optimized for performance, powe… [cited by applicant]
Intel Corp., “Intel Programmable Accelaration Card (Intel PAC) with Intel Arria 10 GX FPGAs”, Product Brief; 1 page. [cited by applicant]
Intel Corp., “Intel Xeon Scalable Platform”, The Future-Forward Platform Foundation for Agile Digital Services; Product Brief; Intel Corporation; 2017, 14 pages. [cited by applicant]
Microway, “Detailed Specifications of the “Skylake-SP”—Intel Xeon Processor Scalable Family CPUs”, retrieved online via https://www.microway.com/knowledge-center-articles/detailed-specifications-of-the-skylake-sp-intel-… [cited by applicant]
NVIDIA, “NVIDIA A100 Tensor Core GPU Architecuture Unprecedented Acceleration at Every Scale”, NVIDIA Corporation, white paper, 2020, 83 pages. [cited by applicant]
NVIDIA, “NVIDIA DGX A100 The Universal System for AI Infrastructure”, NVIDIA Corporation, NVIDIA DGX A100, Data Sheet, May 2020, 2 pages. [cited by applicant]
NVIDIA, “NVIDIA DGX-1 Deep Learning System”, DGX-1 Data Sheet Apr. 2016. [cited by applicant]
NVIDIA, “NVIDIA DGX-1 With Tesla V100 System Architecture The Fastest Platform for Deep Learning”, www.nvidia.com; White Paper; WP-08437-002_v01, 2017, 40 pages. [cited by applicant]
Oh, Nate, “NVIDIA Ships First Volta-based DGX Systems”, Anandtech, Sep. 7, 2017; retrieved online via https://www.anandtech.com/show/11824/nvidia-ships-first-volta-dgx-systems, 3 pages. [cited by applicant]
Saliou, et al., “Analysis of Firewall Performance Variation to Identify the Limits of Automated Network Reconfigurations”, ECIW2006, 2006, pp. 1-11. [cited by applicant]
Smith, R., et al., “NVIDIA Unveils the DGX-1 HPC Server: 8 Teslas, 3U, Q2 2016”, Anandtech, Apr. 6, 2016. Retrieved online via https://www.anandtech.com/show/10229/nvidia-announces-dgx1-server, 7 pages. [cited by applicant]
Wu, “Programming Models' Support for Heterogeneous Architecture”, Doctoral Thesis, University of Tennessee, 2017. [cited by applicant]