IP Library Granted Patent US 12,373,212
Granted Patent B2
US 12,373,212 · App. 18/045,332 · Granted Jul 29, 2025

In-memory computing with cache coherent protocol

Inventors: Krishna T. Malladi (San Jose, CA); Andrew Chang (Los Altos, CA)
Assignee: Samsung Electronics Co., Ltd.
G06F9/30047G06F9/30029G06F9/3887G06F12/0828G06F13/1694G06F13/4221G06F2213/0026
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,212
App. No.
18/045,332
Granted
Jul 29, 2025
Kind
B2
Abstract

A system for computing. In some embodiments, the system includes: a memory, the memory including one or more function-in-memory circuits; and a cache coherent protocol interface circuit having a first interface and a second interface. A function-in-memory circuit of the one or more function-in-memory circuits may be configured to perform an operation on operands including a first operand retrieved from the memory, to form a result. The first interface of the cache coherent protocol interface circuit may be connected to the memory, and the second interface of the cache coherent protocol interface circuit may be configured as a cache coherent protocol interface on a bus interface.

Claims (43)

1. A system for computing, the system comprising:

a memory comprising a function-in-memory circuit, the function-in-memory circuit to perform an operation on data retrieved from the memory to form a result in the memory; and

an interface circuit having a first interface and a second interface, and comprising a processing circuit to perform a computing task comprising a matrix-based operation, the processing circuit being located in a component that is external from the function-in-memory circuit, and being located between the first interface and the second interface, such that a signal for performing the computing task is received by the processing circuit from the second interface, and a result of the computing task is sent from the processing circuit to the first interface,

the first interface being connected to the memory, and

the second interface being configured to receive communications via a bus interface in compliance with a cache-coherent protocol.

2. The system of claim 1 , wherein:

the function-in-memory circuit is arranged in a parallel processor configuration; or

the function-in-memory circuit is arranged in a network of data processing circuits configuration.

3. The system of claim 1 , wherein the cache-coherent protocol comprises compute express link (CXL).

4. The system of claim 1 , wherein a function-in-memory circuit is on a semiconductor chip with a dynamic random-access memory.

5. The system of claim 1 , wherein the first interface is configured to operate according to a protocol comprising a double data rate memory.

6. The system of claim 1 , wherein a function-in-memory circuit comprises:

a register,

a multiplexer, and

an arithmetic logic unit.

7. The system of claim 1 , wherein a function-in-memory circuit is configured to perform an arithmetic operation comprising addition, subtraction, multiplication, or division.

8. The system of claim 1 , wherein a function-in-memory circuit is configured to perform an arithmetic operation comprising floating-point addition, floating-point subtraction, floating-point multiplication, or floating-point division.

9. The system of claim 1 , wherein a function-in-memory circuit is configured to perform a logical operation comprising bitwise AND, bitwise OR, bitwise exclusive OR, and bitwise ones complement.

10. The system of claim 1 , wherein:

the data comprises an operand, and

the function-in-memory circuit is configured to:

in a first state, store the result in the memory, and,

in a second state, send the result to the interface circuit.

11. The system of claim 1 , further comprising a host processing circuit connected to the second interface.

12. The system of claim 11 , wherein the host processing circuit comprises a root complex connected to the second interface.

13. A system for computing, the system comprising:

a memory comprising a function-in-memory circuit, the function-in-memory circuit to perform an operation on data retrieved from the memory to form a result in the memory; and

an interface circuit having a first interface and a second interface, and comprising a processing circuit to perform a computing task comprising a matrix-based operation, the processing circuit being located in a component that is external from the function-in-memory circuit, and being located between the first interface and the second interface, such that a signal for performing the computing task is received by the processing circuit from the second interface, and a result of the computing task is sent from the processing circuit to the first interface,

the first interface being connected to the memory, and

the second interface being configured to receive communications via a bus interface in compliance with a cache-coherent protocol.

14. The system of claim 13 , wherein the data comprises an operand.

15. The system of claim 14 , wherein:

the function-in-memory circuit is arranged in a parallel processor configuration; or

the function-in-memory circuit is arranged in a network of data processing circuits configuration.

16. The system of claim 14 , wherein the cache-coherent protocol comprises compute express link (CXL).

17. The system of claim 14 , wherein a function-in-memory circuit is on a semiconductor chip with a dynamic random-access memory.

18. The system of claim 14 , wherein the first interface is configured to operate according to a protocol comprising a double data rate memory.

19. The system of claim 14 , wherein a function-in-memory circuit is configured, in a first state, to store the result in the memory, and, in a second state, to send the result to the interface circuit.

20. A method for computing, the method comprising:

receiving, by a cache coherent protocol interface circuit, a packet from a host processing circuit; and

sending, by the cache coherent protocol interface circuit, based on receiving the packet, to a function-in-memory circuit in a memory connected to the cache coherent protocol interface circuit, an instruction,

wherein the cache coherent protocol interface circuit is configured to receive communications via a bus interface in compliance with a cache-coherent protocol, and comprises a processing circuit to perform a computing task comprising a matrix-based operation, the processing circuit being located in a component that is external from the function-in-memory circuit, and being located between a first interface, connected to the function-in-memory circuit, and the bus interface, such that a signal for performing the computing task is received by the processing circuit from the bus interface, and a result of the computing task is sent from the processing circuit to the first interface, and

wherein the function-in-memory circuit is configured to perform an operation on data retrieved from the memory to form a result in the memory.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: MALLADI, KRISHNA T.; CHANG, ANDREW ZHENWEN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062083/0963 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2022
From: MALLADI, KRISHNA; CHANG, ANDREW
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061429/0268 →
Continuity (3)
Continuation 16914129 · Jun 26, 2020
Provisional Application 63003701 · Apr 1, 2020
Related Publication 20230069786A1 · Mar 2, 2023
References Cited (43)
US 7971003B2 · Cousin et al. · 2011 [cited by applicant]
US 8365016B2 · Gray et al. · 2013 [cited by applicant]
US 9281026B2 · Felch et al. · 2016 [cited by applicant]
US 9747105B2 · Gopal et al. · 2017 [cited by applicant]
US 11126548B1 · Yudanov · 2021 [cited by applicant]
US 20030115402A1 · Dahlgren et al. · 2003 [cited by applicant]
US 20080005484A1 · Joshi · 2008 [cited by applicant]
US 20090083471A1 · Frey et al. · 2009 [cited by applicant]
US 20130311753A1 · Kandadai · 2013 [cited by applicant]
US 20140310232A1 · Plattner et al. · 2014 [cited by applicant]
US 20150032968A1 · Heidelberger et al. · 2015 [cited by applicant]
US 20160378465A1 · Venkatesh et al. · 2016 [cited by applicant]
US 20170344479A1 · Boyer et al. · 2017 [cited by applicant]
US 20170344546A1 · Nam · 2017 [cited by applicant]
US 20170358327A1 · Oh · 2017 [cited by examiner]
US 20180089093A1 · Heidelberger et al. · 2018 [cited by applicant]
US 20180098869A1 · Reis et al. · 2018 [cited by applicant]
US 20180191374A1 · Wu et al. · 2018 [cited by applicant]
US 20190227981A1 · Tomishima · 2019 [cited by examiner]
US 20190310911A1 · Sundaram et al. · 2019 [cited by applicant]
US 20200026669A1 · Kim · 2020 [cited by examiner]
US 20200065290A1 · Natu · 2020 [cited by applicant]
US 20200117400A1 · Golov · 2020 [cited by applicant]
US 20200125503A1 · Graniello · 2020 [cited by examiner]
US 20200201932A1 · Gradstein et al. · 2020 [cited by applicant]
US 20200294558A1 · Yu · 2020 [cited by examiner]
US 20200328879A1 · Makaram et al. · 2020 [cited by applicant]
US 20200371973A1 · Shan · 2020 [cited by examiner]
US 20210019147A1 · Studer · 2021 [cited by examiner]
US 20210248094A1 · Norman · 2021 [cited by examiner]
US 20210271597A1 · Verma et al. · 2021 [cited by applicant]
US 20210271680A1 · Lee · 2021 [cited by examiner]
US 20210279008A1 · Lea et al. · 2021 [cited by applicant]
US 20210303265A1 · Yudanov · 2021 [cited by applicant]
US 20210335393A1 · Zhao et al. · 2021 [cited by applicant]
US 20210343334A1 · Grover et al. · 2021 [cited by applicant]
TW I610235B · 2018 [cited by applicant]
TW 201937490A · 2019 [cited by applicant]
Lenjani, Marzieh, et al., Fulcrum: A Simplified Control and Access Mechanism Toward Flexible and Practical in-Situ Accelerators, 2020 IEEE International Symposium on High Performance Computer Architecture, 14 pages. [cited by applicant]
Sano, Kentaro, et al., Systolic Computational Memory Approach to High-Speed Codebook Design, 2005 IEEE International Symposium on Signal Processing and Information Technology, 6 pages. [cited by applicant]
European Summons to attend oral proceedings action for Application No. 21166087.3, dated Dec. 20, 2022, 11 pages. [cited by applicant]
Singh, G. et al., “Near-Memory Computing: Past, Present, and Future”, Aug. 7, 2019, pp. 1-16, arXiv:1908.02640v1. [cited by applicant]
Taiwanese Office Action dated Jul. 9, 2024, issued in corresponding Taiwanese Patent Application No. 110111352 (19 pages). [cited by applicant]