IP Library › Granted Patent US 12,229,422
Granted Patent B2
US 12,229,422 · App. 18/523,335 · Granted Feb 18, 2025

On-chip atomic transaction engine

Inventors: Rishabh Jain (Austin, TX); Erik M. Schlanger (Austin, TX)
Assignee: Oracle International Corporation
G06F3/0631G06F3/0604G06F3/067G06F9/526G06F9/547G06F15/17331G06F15/7825
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,229,422
App. No.
18/523,335
Granted
Feb 18, 2025
Kind
B2
Abstract

A hardware-assisted Distributed Memory System may include software configurable shared memory regions in the local memory of each of multiple processor cores. Accesses to these shared memory regions may be made through a network of on-chip atomic transaction engine (ATE) instances, one per core, over a private interconnect matrix that connects them together. For example, each ATE instance may issue Remote Procedure Calls (RPCs), with or without responses, to an ATE instance associated with a remote processor core in order to perform operations that target memory locations controlled by the remote processor core. Each ATE instance may process RPCs (atomically) that are received from other ATE instances or that are generated locally. For some operation types, an ATE instance may execute the operations identified in the RPCs itself using dedicated hardware. For other operation types, the ATE instance may interrupt its local processor core to perform the operations.

Claims (49)

1. A method, comprising:

receiving, at a local atomic transaction engine, information describing an atomic transaction to be performed at a plurality of different memory addresses of a distributed shared memory, wherein individual portions of the distributed shared memory are controlled by a plurality of atomic transaction engines including the local atomic transaction engine and a plurality of remote atomic transaction engines; and

sending, responsive to the receiving, the information describing the atomic transaction to one or more of the plurality of remote atomic transaction engines to be performed at respective memory addresses of the plurality of different memory addresses controlled by the one or more of the plurality of remote atomic transaction engines.

2. The method of claim 1 , wherein the information describing the atomic transaction to be performed at the plurality of different memory addresses is received from a processor coupled to the local atomic transaction engine responsive to determining that an operation to be executed by the processor targets the plurality of different memory addresses in the distributed shared memory.

3. The method of claim 1 , further comprising:

causing performance of a portion of the atomic transaction by the local atomic transaction engine responsive to determining that a memory address of the plurality of different memory addresses of the distributed shared memory is controlled by the local atomic transaction engine.

4. The method of claim 3 , wherein causing performance of the portion of the atomic transaction by the local atomic transaction engine comprises:

performing the portion of the atomic transaction by the local atomic transaction engine responsive to determining that the portion of the atomic transaction is performable by circuitry within the local atomic transaction engine without intervention by a processor coupled to the local atomic transaction engine; and

initiating, by the local atomic transaction engine, performance of the portion of the atomic transaction by the processor responsive to determining that the atomic transaction is not performable by circuitry within the local atomic transaction engine without intervention by the processor.

5. The method of claim 4 , wherein initiating, by the local atomic transaction engine, performance of the portion of the atomic transaction by the processor comprises:

writing the information describing the portion of the atomic transaction into one or more storage locations that are accessible to the processor; and

issuing an interrupt to the processor indicating that the atomic transaction should be executed by the processor.

6. The method of claim 1 , further comprising:

receiving, by the local atomic transaction engine from the one or more of the plurality of remote atomic transaction engines, respective response data for the request, and in response to the receiving:

returning, by the local atomic transaction engine, the respective response data to a processor coupled to the local atomic transaction engine; or

writing, by the local atomic transaction engine, the respective response data to a location in memory from which the processor expects to retrieve it.

7. The method of claim 1 , wherein the local atomic transaction engine and the plurality of remote atomic transaction engines communicate other over a dedicated low-latency interconnect.

8. An apparatus, comprising:

a local transaction engine coupled to a processor and memory, wherein the memory implements a portion of a distributed shared memory, and wherein the local atomic transaction engine is configured to:

receive information describing an atomic transaction to be performed at a plurality of different memory addresses of a distributed shared memory, wherein individual portions of the distributed shared memory are controlled by a plurality of atomic transaction engines including the local atomic transaction engine and a plurality of remote atomic transaction engines; and

send, responsive to the receiving, the information describing the atomic transaction to one or more of the plurality of remote atomic transaction engines to be performed at respective memory addresses of the plurality of different memory addresses controlled by the one or more of the plurality of remote atomic transaction engines.

9. The apparatus of claim 8 , wherein the information describing the atomic transaction to be performed at the plurality of different memory addresses is received from a processor coupled to the local atomic transaction engine responsive to determining that an operation to be executed by the processor targets the plurality of different memory addresses in the distributed shared memory.

10. The apparatus of claim 8 , wherein the local atomic transaction engine is further configured to:

cause performance of a portion of the atomic transaction by the local atomic transaction engine responsive to determining that a memory address of the plurality of different memory addresses of the distributed shared memory is controlled by the local atomic transaction engine.

11. The apparatus of claim 10 , wherein to cause performance of the portion of the atomic transaction, the local atomic transaction engine is configured to:

perform the portion of the atomic transaction by the local atomic transaction engine responsive to determining that the portion of the atomic transaction is performable by circuitry within the local atomic transaction engine without intervention by a processor coupled to the local atomic transaction engine; and

initiate performance of the portion of the atomic transaction by the processor responsive to determining that the atomic transaction is not performable by circuitry within the local atomic transaction engine without intervention by the processor.

12. The apparatus of claim 11 , wherein to initiate performance of the portion of the atomic transaction by the processor, the local atomic transaction engine is configured to:

write the information describing the portion of the atomic transaction into one or more storage locations that are accessible to the processor; and

issue an interrupt to the processor indicating that the atomic transaction should be executed by the processor.

13. The apparatus of claim 12 , wherein the local atomic transaction engine and the plurality of remote atomic transaction engines communicate other over a dedicated low-latency interconnect.

14. A system, comprising:

a plurality of atomic transaction engines respectively coupled to respective processors and respective memories, wherein the respective memories collectively implement a distributed shared memory, and wherein a local atomic transaction engine of the plurality of atomic transaction engines is configured to:

receive information describing an atomic transaction to be performed at a plurality of different memory addresses of a distributed shared memory, wherein individual portions of the distributed shared memory are controlled by a plurality of atomic transaction engines including the local atomic transaction engine and a plurality of remote atomic transaction engines; and

send, responsive to the receiving, the information describing the atomic transaction to one or more of the plurality of remote atomic transaction engines to be performed at respective memory addresses of the plurality of different memory addresses controlled by the one or more of the plurality of remote atomic transaction engines.

15. The system of claim 14 , wherein the information describing the atomic transaction to be performed at the plurality of different memory addresses is received from a processor coupled to the local atomic transaction engine responsive to determining that an operation to be executed by the processor targets the plurality of different memory addresses in the distributed shared memory.

16. The system of claim 14 , wherein the local atomic transaction engine is further configured to:

cause performance of a portion of the atomic transaction by the local atomic transaction engine responsive to determining that a memory address of the plurality of different memory addresses of the distributed shared memory is controlled by the local atomic transaction engine.

17. The system of claim 16 , wherein to cause performance of the portion of the atomic transaction, the local atomic transaction engine is configured to:

perform the portion of the atomic transaction by the local atomic transaction engine responsive to determining that the portion of the atomic transaction is performable by circuitry within the local atomic transaction engine without intervention by a processor coupled to the local atomic transaction engine; and

initiate performance of the portion of the atomic transaction by the processor responsive to determining that the atomic transaction is not performable by circuitry within the local atomic transaction engine without intervention by the processor.

18. The system of claim 17 , wherein to initiate performance of the portion of the atomic transaction by the processor, the local atomic transaction engine is configured to:

write the information describing the portion of the atomic transaction into one or more storage locations that are accessible to the processor; and

issue an interrupt to the processor indicating that the atomic transaction should be executed by the processor.

19. The system of claim 18 , wherein the local atomic transaction engine is further configured to:

receive, from the one or more of the plurality of remote atomic transaction engines, respective response data for the request, and in response to the receiving:

return the respective response data to a processor coupled to the local atomic transaction engine; or

writing the respective response data to a location in memory from which the processor expects to retrieve it.

20. The system of claim 14 , wherein the local atomic transaction engine and the plurality of remote atomic transaction engines communicate other over a dedicated low-latency interconnect.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2023
From: JAIN, RISHABH; SCHLANGER, ERIK M.
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 065713/0035 →
Continuity (4)
Continuation 17663280 · May 13, 2022
Continuation 16945521 · Jul 31, 2020
Continuation 14863354 · Sep 23, 2015
Related Publication 20240111441A1 · Apr 4, 2024
References Cited (47)
US 5617570A · Russell et al. · 1997 [cited by applicant]
US 5987506A · Carter et al. · 1999 [cited by applicant]
US 6631439B2 · Saulsbury et al. · 2003 [cited by applicant]
US 10140149B1 · Hayes · 2018 [cited by examiner]
US 10732865B2 · Jain et al. · 2020 [cited by applicant]
US 11334262B2 · Jain et al. · 2022 [cited by applicant]
US 20020161848A1 · Willman et al. · 2002 [cited by applicant]
US 20030061395A1 · Hamdan et al. · 2003 [cited by applicant]
US 20030217104A1 · Hamdan et al. · 2003 [cited by applicant]
US 20070016830A1 · Hasha · 2007 [cited by applicant]
US 20110125974A1 · Anderson · 2011 [cited by applicant]
US 20120089762A1 · Zhu et al. · 2012 [cited by applicant]
US 20130054726A1 · Bugge · 2013 [cited by applicant]
US 20130212148A1 · Koponen et al. · 2013 [cited by applicant]
US 20140089572A1 · Koka et al. · 2014 [cited by applicant]
US 20140181421A1 · O'Connor et al. · 2014 [cited by applicant]
US 20140207871A1 · Miloushev et al. · 2014 [cited by applicant]
US 20140250276A1 · Blaner et al. · 2014 [cited by applicant]
US 20150215386A1 · Walsky et al. · 2015 [cited by applicant]
US 20160357702A1 · Shamis et al. · 2016 [cited by applicant]
US 20220276794A1 · Jain et al. · 2022 [cited by applicant]
CN 103493037 · 2014 [cited by applicant]
EP 0614139 · 1994 [cited by applicant]
International Search Report and Written Opinion from PCT/US2016/052968, Date of Mailing Dec. 2, 2016, Oracle International Corporation, pp. 1-16. [cited by applicant]
J. Silcock, et al., “Message Passing Remote Procedure Calls and Distributed Shared Memory as Communication Paradigms for Distributed Systems”, 2014, pp. 1-16. [cited by applicant]
“System Administrative Guide:Network Serivices Chapter 30 Monitoring Network Performance (Task)”, retrieved from https://docs.oracle.com/cd/E19253-01/816-4555/6maoquigp/index.html, pp. 1-7. [cited by applicant]
Taiyang Chen, “Remote Procedure Calls”, Retrieved from http://www.cs.cornell.edu/courses/cs6410/2009fa/lectures/11-rpc.pdf, Oct. 6, 2009, pp. 1-50. [cited by applicant]
David R. Cheriton, et al., “Optimized Memory-Based Messaging: Leveraging the Memory System for High-Performance Communication”, 1994, pp. 1-24. [cited by applicant]
Takashi Furkawa, et al., “A Hardware/Software Cosimulator with RTOS Supports for Multiprocessor Embedded Systems”, 2007, pp. 1-12. [cited by applicant]
“Editor, Know more about your laptop parts”, Retrieved from http://www.techadvisory.org/2014/02/know-more-about-your-laptop-parts, Feb. 20, 2014, Jul. 9, 2018 pp. 1-3. [cited by applicant]
Wolf, “Computers as Components: Principles of Embedded Computer System Design”, 2008, Elsevier, 2nd Edition, retrieved from http://www.bookspar.com/wp-content/uploads/vtu/notes/cs/7th-sem/ecs-72/computer-as-components-p… [cited by applicant]
“Switching—An Engineering Approach to Computer Networking”, Slides, Retrieved from http://cs.cornell/edu/skeshav/book/slides/index.html, Cornell University, 1999, pp. 1-63. [cited by applicant]
“What is a network?”, Retrieved from http://www.bbc.co.uk/schools/gesebitesize/ict/datacomm/2networksrev1.shtml, Jul. 16, 2018, pp. 1-10. [cited by applicant]
Snoeren, “Network File System/RPC”, 2006, University of California Sand Diego, Slides, Retrieved form https://cseweb.ucsd.edu/classes/fa06/cse120/lectures/120-fa06-l15.pdf. [cited by applicant]
Thorvaldsson, “Atomic Transfer for Distributed Systems”, 2009, WUSL, Retrieved from http://citeseerx.ist.psu.edu/viewdoc/download;jsessionid=41A7E344411C590069B1A92FB1E590C3?doi-10.1.1.466.2829&rep=rep1&type=pdf, pp. 25… [cited by applicant]
Incognito, “FAQ: TR-069”, Retrieved from https://www.incognito.com/tips-and-tutorials/faq-tr-069/, Aug. 29, 2013, pp. 1-7. [cited by applicant]
Martin Bond, et al., “Using RPC-Style Web Services with J2EE”, http://www.informit.com/artices/printerfriendly/29416, Sep. 2, 2002, pp. 1-34. [cited by applicant]
Renesse, et al., “Connecting RPC-Based Distributed System Using Wide-Area Networks”, Retrieved from https://pdfs.semanticsscholar.org/0360/65a536609fdc36adf1b3d6acaaa23b12419.pdf?ga=2.62631596.1823017162.1531731074-2012… [cited by applicant]
Pautasso, “Remote Procedure Call”, ISCRG, Retrieved from https://disco.ethz.ch/courses/ss06/vs/material/chapter 7/RPC_4.pdf, 2006, pp. 1-7. [cited by applicant]
Plusquellic, “Distributed Shared-Memory Architectures”, CSEE, UMBC, Retrieved from http://ece-research.unm.edu/jimp/611/slides/chap8-3.html, 1999, pp. 1-11. [cited by applicant]
Krishna Kavi, et al., “Shared Memory and Distributed Shared Memory Systems: A Survey”, 2000, Elsevier, Retrieved from http://web.engr.oregonstate.edu/ benl/Publications/Book_Chapters/Advances_in_Computers_DSM00.pdf, pp.… [cited by applicant]
Jelica Protic, et al, “Distributed Shared Memory: Concepts and Systems”, IEEE, Retrieved from https://www.cc.gatech.edu/classes/AY2009/cs4210_fall/papers/DSM_protic.pdf, 1996, pp. 63-79. [cited by applicant]
Angela Demke Brown, “Lecture 16/17: Distributed Shared Memory”, Retrieved from http:www.cs.toronto.edu/˜demke/469f.06/Lectures/Lecture16.PDF, 2006, pp. 1-22. [cited by applicant]
Michele Mazzucco et al., “Engineering Distributed Shared Memory Middleware for Java”, Retrieved from https://pdf/semanticsscholar.org/9323/fba3989edd3975ec2b1aa91d6003b05cd2e3/pdf, 2009, pp. 531-548. [cited by applicant]
Rene W. Schmidt, et al., “Using Shared Memory for Read-Mostly RPC Services”, Proceedings of the 29th Annual Hawaii International Conference Sciences, IEEE, Retrieved from https://ieeexplore.ieee.org/stam/stamp.jsp?tp=&a… [cited by applicant]
Office Action in European Application No. 16779232.4 mailed Jun. 22, 2021, Oracle International Corporation, pp. 1-8. [cited by applicant]
Office Action mailed Jul. 5, 2021 in Chinese Patent Application No. 201680055397.1, Oracle International Corporation, including translation. [cited by applicant]