IP Library › Granted Patent US 12,524,281
Granted Patent B2
US 12,524,281 · App. 17/799,098 · Granted Jan 13, 2026

C2MPI: a hardware-agnostic message passing interface for heterogeneous computing systems

Inventors: Michael F. Riera (Hialeah, FL); Fengbo Ren (Tempe, AZ); Masudul Hassan Quraishi (Tempe, AZ); Erfan Bank Tavakoli (Tempe, AZ)
Assignee: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
G06F9/541G06F9/3836G06F9/3877G06F9/44G06F9/4881G06F9/505G06F9/545
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,281
App. No.
17/799,098
Granted
Jan 13, 2026
Kind
B2
Abstract

Compute-centric message passing interface (C 2 MPI) provides a hardware-agnostic message passing interface for heterogeneous computing systems. Hardware-agnostic programming with high performance portability is envisioned to be a bedrock for realizing adoption of emerging accelerator technologies in heterogeneous computing systems, such as high-performance computing (HPC) systems, data center computing systems, and edge computing systems. The adoption of emerging accelerators is key to achieving greater scale and performance in heterogeneous computing systems. Accordingly, embodiments described herein provide a flexible hardware-agnostic environment that allows application developers to develop high-performance applications without knowledge of the underlying hardware.

Claims (45)

1 . A method for providing instructions for a host application to a heterogeneous computing system via a compute-centric message passing interface (C 2 MPI), the method comprising:

providing a first hardware-agnostic instruction to invoke a first child rank corresponding to a first accelerator resource, wherein the first hardware-agnostic instruction specifies a first computational function,

wherein the C 2 MPI includes application parent ranks defined inside a hardware-agnostic region of the host application, and accelerator parent ranks, the application parent ranks and the accelerator parent ranks managing child ranks, the accelerator parent ranks performing hardware management, kernel retrieval, registration, and execution and maintaining of system resources allocated by the application parent ranks, and

wherein the application parent ranks and the accelerator parent ranks allocate and deallocate the child ranks based on a duration of execution time of the host application or the first accelerator resource, one of the child ranks being associated with the first accelerator resource after allocation.

2 . The method of claim 1 , further comprising providing a second hardware-agnostic instruction to send first data to the first child rank for processing.

3 . The method of claim 2 , wherein the second hardware-agnostic instruction is sent in response to receiving the first child rank.

4 . The method of claim 2 , further comprising providing a third hardware-agnostic instruction to receive a first processing result from the first child rank.

5 . The method of claim 1 , further comprising providing a second hardware-agnostic instruction to invoke a second child rank corresponding to a second accelerator resource.

6 . The method of claim 5 , wherein the second hardware-agnostic instruction specifies the first computational function.

7 . The method of claim 5 , wherein the second hardware-agnostic instruction specifies a second computational function different from the first computational function.

8 . The method of claim 5 , further comprising:

providing a third hardware-agnostic instruction to send first data to the first child rank for processing; and

providing a fourth hardware-agnostic instruction to send second data to the second child rank for processing.

9 . The method of claim 8 , wherein the third hardware-agnostic instruction and the fourth hardware agnostic instruction allow for parallel execution of the first child rank and the second child rank.

10 . The method of claim 8 , wherein the third hardware-agnostic instruction and the fourth hardware agnostic instruction require pipeline execution of the first child rank and the second child rank.

11 . A method for executing instructions for an application on a heterogeneous computing system received via a compute-centric message passing interface (C 2 MPI), the method comprising:

receiving, from a host application, a first hardware-agnostic instruction to invoke a first child rank corresponding to a first accelerator resource, wherein the first hardware-agnostic instruction specifies a first computational function; and

locating the first accelerator resource based on the first computational function,

wherein the C 2 MPI includes application parent ranks, inside a hardware-agnostic region of the application, and accelerator parent ranks, the application parent ranks and the accelerator parent ranks managing child ranks, the accelerator parent ranks performing hardware management, kernel retrieval, registration, and execution and maintaining of system resources allocated by the application parent ranks, and

wherein the application parent ranks and the accelerator parent ranks allocate and deallocate the child ranks based on a duration of execution time of the application or the first accelerator resource, one of the child ranks being associated with the first accelerator resource after allocation.

12 . The method of claim 11 , wherein the first accelerator resource corresponds to one or a set of accelerators programmed to execute the first computational function.

13 . The method of claim 12 , wherein the set of accelerators programmed to execute the first computational function comprise a same type of accelerator.

14 . The method of claim 12 , wherein the set of accelerators programmed to execute the first computational function comprise heterogeneous types of accelerators.

15 . The method of claim 11 , further comprising:

invoking the first child rank using the first accelerator resource; and returning the first child rank to the host application.

16 . The method of claim 11 , wherein locating the first accelerator resource comprises selecting a type of accelerator suitable for executing the first computational function.

17 . The method of claim 16 , wherein selecting the type of accelerator suitable for executing the first computational function comprises selecting the type of accelerator having a pre-selected performance for executing the first computational function.

18 . The method of claim 16 , wherein locating the first accelerator resource further comprises:

identifying an available one of the type of accelerator suitable for executing the first computational function; and

reserving the available one of the type of accelerator suitable for executing the first computational function as the first accelerator resource.

19 . The method of claim 11 , further comprising:

receiving a second hardware-agnostic instruction to send first data to the first child rank for processing; and

forwarding the first data to the first accelerator resource.

20 . The method of claim 19 , further comprising receiving a first processing result from the first accelerator resource corresponding to the first data.

21 . The method of claim 20 , further comprising staging the first processing result.

22 . The method of claim 20 , further comprising:

receiving a third hardware-agnostic instruction to receive the first processing result from the child rank; and

forwarding the first processing result to the host application.

23 . A non-transitory computer-readable medium having stored thereon software instructions that, when executed by a processor, cause the processor to:

receive, from a host application, a first hardware-agnostic instruction to invoke a first child rank corresponding to a first accelerator resource, wherein the first hardware-agnostic instruction specifies a first computational function;

locate the first accelerator resource based on the first computational function;

invoke the first child rank; and

return the first child rank to the host application,

wherein the C 2 MPI includes application parent ranks, inside a hardware-agnostic region of the application, and accelerator parent ranks, the application parent ranks and the accelerator parent ranks managing child ranks, the accelerator parent ranks performing hardware management, kernel retrieval, registration, and execution and maintaining of system resources allocated by the application parent ranks, and

wherein the application parent ranks and the accelerator parent ranks allocate and deallocate the child ranks based on a duration of execution time of the application or the first accelerator resource, one of the child ranks being associated with the first accelerator resource after allocation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2022
From: RIERA, MICHAEL F.; REN, FENGBO; QURAISHI, MASUDUL HASSAN; BANK TAVAKOLI, ERFAN
To: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
Reel/Frame 060831/0031 →
Continuity (2)
Provisional Application 62983220 · Feb 28, 2020
Related Publication 20230074426A1 · Mar 9, 2023
References Cited (25)
US 7975001B1 · Stefansson · 2011 [cited by examiner]
US 20100076915A1 · Xu et al. · 2010 [cited by applicant]
US 20110271263A1 · Archer et al. · 2011 [cited by applicant]
US 20120069029A1 · Bourd · 2012 [cited by examiner]
US 20130073752A1 · Blocksome · 2013 [cited by applicant]
US 20190007334A1 · Guim Bernat et al. · 2019 [cited by applicant]
Heath Jr. et al., An overview of signal processing techniques for millimeter wave mimo systems, IEEE Journal of Selected Topics in Signal Processing, vol. 10, No. 3, Apr. 2016, pp. 436-453. [cited by applicant]
Rappaport et al., Wireless communications and applications above 100 ghz: Opportunities and challenges for 6g and beyond, IEEE Access, vol. 7, 2019, pp. 78 729-78 757. [cited by applicant]
Bai et al., Coverage and capacity of millimeter-wave cellular networks, IEEE Communications Magazine, vol. 52, No. 9, Sep. 2014, pp. 70-77. [cited by applicant]
Alrabeiah et al., Deep learning for TDD and FDD massive MIMO: mapping channels in space and frequency, CoRR, vol. abs/1905.03761, 2019, http://arxiv.org/abs/1905.03761, 10 pages. [cited by applicant]
Li et al., Deep Learning for Direct Hybrid Precoding in Millimeter Wave Massive MIMO Systems, CoRR, vol. abs/1905.13212, 2019, http://arxiv.org/abs/1905.13212, 7 pages. [cited by applicant]
Alkhateeb et al., Deep learning coordinated beamforming for highly-mobile millimeter wave systems, IEEE Access, vol. 6, 2018, pp. 37 328-37 348. [cited by applicant]
Alkhateeb et al., Machine learning for reliable mmwave systems: Blockage prediction and proactive handoff, IEEE GlobalSIP, arXiv preprint arXiv:1807.02723, 2018, 5 pages. [cited by applicant]
Taha et al., Enabling large intelligent surfaces with compressive sensing and deep learning, arXiv preprint, arXiv:1904.10136, 2019, 33 pages. [cited by applicant]
Alrabeiah et al., Deep learning for mmwave beam and blockage prediction using Sub-6GHZ channels, submitted to IEEE Transactions on Communications, arXiv e-prints, p. arXiv:1910.02900, Oct. 2019, 31 pages. [cited by applicant]
Ecklemann et al., V2v-communication, lidar system and positioning sensors for future fusion algorithms in connected vehicles, Transportation Research Procedia, vol. 27, 2017, pp. 69-76. [cited by applicant]
He et al., Deep residual learning for image recognition, Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770-778. [cited by applicant]
Russakovsky et al., Imagenet large scale visual recognition challenge, International journal of computer vision, vol. 115, No. 3, 2015, pp. 211-252. [cited by applicant]
Goodfellow et al., Deep Learning, 2016, book in preparation for MIT Press, available online at: http://www.deeplearningbook.org/. [cited by applicant]
Alrabeiah et al., Viwi: a deep learning dataset framework for vision-aided wireless communications, submitted to IEEE Vehicular Technology Conference, Nov. 2019, https://www.viwi-dataset.net/, 5 pages. [cited by applicant]
Paszke et al., Automatic differentiation in PyTorch, NIPS Autodiff Workshop, 2017, 4 pages. [cited by applicant]
Alrabeiah et al., mmWave Base Stations with Cameras, available online at: https://github.com/malrabeiah/CameraPredictsBeams, 3 pages. [cited by applicant]
Parkvall et al., “NR: The new 5G radio access technology,” IEEE Communications Standards Magazine, vol. 1, No. 4, Dec. 2017. pp. 24-30. [cited by applicant]
Nakashima et al., “Impact of input data size on received power prediction using depth images for mm wave communications,” in 2018 IEEE 88th Vehicular Technology Conference (VTC-Fall), Aug. 2018, pp. 1-5. [cited by applicant]
International Search Report and Written Opinion mailed May 18, 2021 in corresponding International Application No. PCT/US2021/020353, 12 pages. [cited by applicant]