IP Library Granted Patent US 12,632,385
Granted Patent B2
US 12,632,385 · App. 18/754,079 · Granted May 19, 2026

Coherent communication between a processor core and an accelerator

Inventors: Simon Weishaupt (Kernen im Remstal, DE); Cedric Lichtenau (Stuttgart, DE); Simon Friedmann (Boeblingen, DE); Preetham M. Lobo (Bangalore, IN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F12/0831G06F12/084G06F2212/1016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,385
App. No.
18/754,079
Granted
May 19, 2026
Kind
B2
Abstract

One or more systems, devices, computer program products and/or computer-implemented methods of use provided herein relate to communication between a processor core and an accelerator. For example, a system can comprise a memory that can store computer executable components. The system can further comprise a processor that can execute the computer executable components stored in the memory, wherein the computer executable components can comprise a tracking component that can track a running state of an accelerator during execution of one or more functions by the accelerator. The computer executable components can further comprise an installation component that can install, via the accelerator, a message in a cache accessible to a processor core, wherein a cache line comprised within the cache can be updated based on installation of the message in the cache.

Claims (41)

1 . A system, comprising:

a memory that stores computer executable components;

a signaling bus that is directly connected between a processor core and an accelerator; and

a processor that executes the computer executable components stored in the memory, wherein the computer executable components comprise:

a selection component that selects the accelerator to execute one or more functions related to a program being executed by the processor core; and

a transmission component that transmits an address associated with a cache to the accelerator, wherein the address is transmitted within a start command from the processor core to the accelerator via the signaling bus, wherein the cache is accessible to both the processor core and the accelerator, and wherein the address is employable by the accelerator to install one or more messages in the cache via non-cached stores to communicate a running state of the accelerator to the processor core during execution of the one or more functions.

2 . The system of claim 1 , wherein the start command is transmitted from the processor core to the accelerator via a sideband signaling mechanism.

3 . The system of claim 2 , wherein the start command also comprises information about the one or more functions to be executed by the accelerator.

4 . The system of claim 1 , wherein the processor core is located on a first chip, and wherein the accelerator is located on the first chip, on the processor core, or on a second chip that is operatively coupled to the first chip.

5 . The system of claim 1 , wherein the accelerator is an artificial intelligence (AI) processing unit.

6 . The system of claim 1 , further comprising:

a monitoring component that:

accesses the one or more messages; and

monitors the running state of the accelerator at the cache, based on the one or more messages.

7 . The system of claim 6 , further comprising:

an inferencing component that:

determines, based on the monitoring, whether the running state indicates an updated state of the accelerator as compared to a previously recorded running state of the accelerator;

determines, based on the running state, a phase of the program being executed by the processor core; and

determines, based on the phase, whether the accelerator has completed execution of the one or more functions.

8 . The system of claim 7 , wherein the selection component acquires a lock for the accelerator after selecting the accelerator and releases the lock based on a determination that the accelerator has completed the execution of the one or more functions.

9 . A computer-implemented method, comprising:

selecting, by a system operatively coupled to a processor, an accelerator to execute one or more functions related to a program being executed by a processor core; and

transmitting, by the system, an address associated with a cache to the accelerator, wherein the address is transmitted within a start command from the processor core to the accelerator via a signaling bus that is directly connected between the processor core and the accelerator, wherein the cache is accessible to both the processor core and the accelerator, and wherein the address is employable by the accelerator to install one or more messages in the cache via non-cached stores to communicate a running state of the accelerator to the processor core during execution of the one or more functions.

10 . The computer-implemented method of claim 9 , wherein the start command is transmitted from the processor core to the accelerator via a sideband signaling mechanism.

11 . The computer-implemented method of claim 10 , wherein the start command also comprises information about the one or more functions to be executed by the accelerator.

12 . The computer-implemented method of claim 9 , wherein the processor core is located on a first chip, and wherein the accelerator is located on the first chip, on the processor core, or on a second chip that is operatively coupled to the first chip.

13 . The computer-implemented method of claim 9 , wherein the accelerator is an AI processing unit.

14 . The computer-implemented method of claim 9 , further comprising:

accessing, by the system, the one or more messages; and

monitoring, by the system, the running state of the accelerator at the cache, based on the one or more messages.

15 . The computer-implemented method of claim 14 , further comprising:

determining, by the system, based on the monitoring, whether the running state indicates an updated state of the accelerator as compared to a previously recorded running state of the accelerator;

determining, by the system, based on the running state, a phase of the program being executed by the processor core; and

determining, by the system, based on the phase, whether the accelerator has completed execution of the one or more functions.

16 . The computer-implemented method of claim 15 , further comprising:

acquiring, by the system, a lock for the accelerator after the selecting; and

releasing, by the system, the lock based on a determination that the accelerator has completed the execution of the one or more functions.

17 . A computer program product for communication between a processor core and an accelerator, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

select, by the processor, the accelerator to execute one or more functions related to a program being executed by the processor core; and

transmit, by the processor, an address associated with a cache to the accelerator, wherein the address is transmitted within a start command from the processor core to the accelerator via a signaling bus that is directly connected between the processor core and the accelerator, wherein the cache is accessible to both the processor core and the accelerator, and wherein the address is employable by the accelerator to install one or more messages in the cache via non-cached stores to communicate a running state of the accelerator to the processor core during execution of the one or more functions.

18 . The computer program product of claim 17 , wherein the start command is transmitted from the processor core to the accelerator via a sideband signaling mechanism.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2024
From: WEISHAUPT, SIMON; LICHTENAU, CEDRIC; FRIEDMANN, SIMON; LOBO, PREETHAM M.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 067836/0983 →
Continuity (1)
Related Publication 20250390433A1 · Dec 25, 2025
References Cited (27)
US 10169108B2 · Gou et al. · 2019 [cited by applicant]
US 11068397B2 · Gou et al. · 2021 [cited by applicant]
US 12026093B1 · John · 2024 [cited by examiner]
US 20100198997A1 · Archer et al. · 2010 [cited by applicant]
US 20140204099A1 · Ye · 2014 [cited by applicant]
US 20160283240A1 · Mishra et al. · 2016 [cited by applicant]
US 20190155754A1 · Kida · 2019 [cited by examiner]
US 20190188136A1 · Shah et al. · 2019 [cited by applicant]
US 20190347125A1 · Sankaran · 2019 [cited by examiner]
US 20200050490A1 · Schardt · 2020 [cited by examiner]
US 20200285942A1 · Yan · 2020 [cited by examiner]
US 20210312093A1 · Buyuktosunoglu et al. · 2021 [cited by applicant]
US 20220214903A1 · Zhao et al. · 2022 [cited by applicant]
US 20230153168A1 · Abali et al. · 2023 [cited by applicant]
US 20230409302A1 · Kodama · 2023 [cited by examiner]
US 20240143758A1 · Yim · 2024 [cited by examiner]
US 20240338327A1 · Syrivelis et al. · 2024 [cited by applicant]
US 20250013509A1 · Marchand · 2025 [cited by applicant]
US 20250173269A1 · Soltaniyeh et al. · 2025 [cited by applicant]
CN 116841835A · 2023 [cited by applicant]
Elliott Glenn et al., “An optimal k-exclusion real-time locking protocol motivated by multi-GPU systems”, Real Time Systems, Dec. 6, 2012, pp. 140-170, vol. 49, doi: https://doi.org/10.1007/s11241-012-9170-0. [cited by applicant]
International Searching Authority, “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or Declaration,” Patent Cooperation Treaty, Aug. 1, 20… [cited by applicant]
International Searching Authority, “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or Declaration,” Patent Cooperation Treaty, Aug. 4, 20… [cited by applicant]
M. Asri et al., “CASPHAr: Cache-Managed Accelerator Staging and Pipelining in Heterogeneous System Architecture”, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Nov. 2022, pp. 4325-4336, … [cited by applicant]
N. Agarwal et al., “Selective GPU caches to eliminate CPU-GPU HW cache coherence”, 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), Barcelona, Spain, 2016, pp. 494-506, doi: 10.1109/HP… [cited by applicant]
United States Non-Final Rejection dated Jun. 11, 2025, 14 pages, in U.S. Appl. No. 18/754,065. [cited by applicant]
United States Notice of Allowance dated Aug. 22, 2025, 12 pages, in U.S. Appl. No. 18/754,065. [cited by applicant]