IP Library Granted Patent US 12,639,003
Granted Patent B2
US 12,639,003 · App. 17/627,478 · Granted May 26, 2026

Compute accelerated stacked memory

Inventors: Mark D. Kellam (Siler City, NC); Steven C. Woo (Saratoga, CA); Thomas Vogelsang (Mountain View, CA); John Eric Linstadt (Palo Alto, CA)
Assignee: Rambus Inc.
G06F3/0656G06F3/0604G06F3/0673G06F15/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,003
App. No.
17/627,478
Granted
May 26, 2026
Kind
B2
Abstract

An integrated circuit that includes a set of one or more logic layers that are, when the integrated circuit is stacked in an assembly with the set of stacked memory devices, electrically coupled to a set of stacked memory devices. The set of one or more logic layers include a coupled chain of processing elements. The processing elements in the coupled chain may independently compute partial results as functions of data received, store partial results, and pass partial results directly to a next processing element in the coupled chain of processing elements. The processing elements in the chains may include interfaces that allow direct access to memory banks on one or more DRAMs in the stack. These interfaces may access DRAM memory banks via TSVs that are not used for global I/O. These interfaces allow the processing elements to have more direct access to the data in the DRAM.

Claims (39)

1 . An integrated circuit, comprising:

a set of one or more logic layers to interface to a set of stacked memory devices when the integrated circuit is stacked with the set of stacked memory devices; and

the set of one or more logic layers comprising:

a serially coupled chain of processing elements, wherein:

respective outputs of processing elements are directly coupled to respective inputs of next processing elements in the serially coupled chain of processing elements; and,

processing elements in the serially coupled chain are to independently of other processing elements in the serially coupled chain of processing elements compute partial results as functions of data received, store partial results, and pass partial results directly to a next processing element in the serially coupled chain of processing elements.

2 . The integrated circuit of claim 1 , wherein the serially coupled chain of processing elements includes an input processing element to receive data from an input interface to the serially coupled chain of processing elements.

3 . The integrated circuit of claim 2 , wherein the serially coupled chain of processing elements includes an output processing element to pass results to an output interface of the serially coupled chain of processing elements.

4 . The integrated circuit of claim 3 , wherein, a processing system is formed when the integrated circuit is stacked with the set of stacked memory devices.

5 . The integrated circuit of claim 4 , wherein the set of one or more logic layers further comprises:

a centrally located region of the integrated circuit that includes global input and output circuitry to interface the processing system and an external processing system.

6 . The integrated circuit of claim 5 , wherein the set of one or more logic layers further comprises:

first staging buffers coupled between the global input and output circuitry and the serially coupled chain of processing elements to communicate data with at least one of the input processing element and the output processing element.

7 . The integrated circuit of claim 6 , wherein the set of one or more logic layers further comprises:

a plurality of serially coupled chains of processing elements and a plurality of staging buffers, respective ones of the plurality of staging buffers coupled between the global input and output circuitry and corresponding ones of the plurality of serially coupled chains of processing elements to communicate data with at least one of a respective input processing element and a respective output processing element of the corresponding one of the plurality of serially coupled chains of processing elements.

8 . An integrated circuit configured to be attached to, and interface with, a stack of memory devices, the integrated circuit comprising:

a first set of processing elements that are connected in a first serially connected chain topology having respective outputs of the first set of processing elements directly coupled with respective inputs of next processing elements in the serially connected chain topology, where processing elements in the first serially connected chain topology are to, independently of other processing elements in the first serially connected chain topology, compute partial results using received data, to store partial results, and to directly pass partial results to a next processing element in the first serially connected chain topology.

9 . The integrated circuit of claim 8 , wherein the first serially connected chain topology includes a first input processing element to receive data from a first input interface of the first serially connected chain topology.

10 . The integrated circuit of claim 9 , wherein the first serially connected chain topology includes a first output processing element to pass results to a first output interface of the first serially connected chain topology.

11 . The integrated circuit of claim 10 , wherein the first input processing element and the first output processing element are the same processing element.

12 . The integrated circuit of claim 10 , further comprising:

a centrally located region of the integrated circuit that includes global input and output circuitry to interface the stack of memory devices and the integrated circuit with an external processing system.

13 . The integrated circuit of claim 12 , further comprising:

first staging buffers coupled between the first input interface, the first output interface, and the global input and output circuitry.

14 . The integrated circuit of claim 13 , further comprising:

a second set of processing elements that are connected in a second serially connected chain topology, where processing elements in the second serially connected chain topology are to, independently of other processing elements of the second serially connected chain topology, compute partial results using received data, to store partial results, and to directly pass partial results to a next element in the second serially connected chain topology, wherein the second serially connected chain topology includes a second input processing element to receive data from a second input interface of the second serially connected chain topology and a second output processing element to pass results to a second output interface of the second serially connected chain topology; and,

second staging buffers coupled between the second input interface, the second output interface, and the global input and output circuitry.

15 . A system, comprising:

a set of stacked memory devices comprising memory cell circuitry;

a set of one or more processing devices electrically coupled to the set of stacked memory devices, the set of processing devices comprising:

a first set of at least two processing elements that are connected in a serial chain topology having respective outputs of processing elements in the first set directly coupled with respective inputs of next processing elements in the serial chain topology, where processing elements in the first set are to, independently of other of the first set of at least two processing elements, compute partial results using received data, to store partial results, and to directly pass partial results to a next processing element in the serial chain topology, wherein the first set includes a first input processing element to receive data from a first input interface to the first set and a first output processing element to pass results to a first output interface of the first set.

16 . The system of claim 15 , wherein the set of one or more processing devices further comprise:

a second set of at least two processing elements that are connected in the serial chain topology, where processing elements in the second set are to, independently of other of the second set of at least two processing elements, compute partial results using received data, to store partial results, and to directly pass partial results to a next element in the serial chain topology, wherein the second set includes a second input processing element to receive data from a second input interface to the second set and a second output processing element to pass results to a second output interface of the second set.

17 . The system of claim 16 , wherein the set of processing devices further comprise:

a set of staging buffers connected in a ring topology, wherein a first at least one of the set of staging buffers is coupled to the first input interface to supply data to the first input processing element, a second at least one of the set of staging buffers is coupled to the second input interface to supply data to the second input processing element.

18 . The system of claim 16 , wherein a third at least one of the set of staging buffers is coupled to the first output interface to receive data from the first output processing element, a fourth at least one of the set of staging buffers is coupled to the second output interface to receive data from the second output processing element.

19 . The system of claim 18 , wherein the set of processing devices further comprise:

a memory interface coupled to the set of staging buffers and coupleable to an external device that is external to the system, the memory interface to perform operations that access, for the external device, the set of stacked memory devices.

20 . The system of claim 19 , wherein the memory interface is to perform operations that access, for the external device, the set of staging buffers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2022
From: KELLAM, MARK D; WOO, STEVEN C; VOGELSANG, THOMAS; LINSTADT, JOHN ERIC
To: RAMBUS INC.
Reel/Frame 058664/0457 →
Continuity (3)
Provisional Application 62923289 · Oct 18, 2019
Provisional Application 62876488 · Jul 19, 2019
Related Publication 20220269436A1 · Aug 25, 2022
References Cited (54)
US 8258619B2 · Foster, Sr. et al. · 2012 [cited by applicant]
US 8310841B2 · Foster, Sr. et al. · 2012 [cited by applicant]
US 8315068B2 · Foster, Sr. et al. · 2012 [cited by applicant]
US 8432027B2 · Foster, Sr. et al. · 2013 [cited by applicant]
US 8437163B2 · Nakanishi et al. · 2013 [cited by applicant]
US 8493089B2 · Cordero et al. · 2013 [cited by applicant]
US 8516409B2 · Coteus et al. · 2013 [cited by applicant]
US 8546955B1 · Wu · 2013 [cited by applicant]
US 8593170B2 · Van der Plas et al. · 2013 [cited by applicant]
US 8612687B2 · Li · 2013 [cited by examiner]
US 8778734B2 · Metsis · 2014 [cited by applicant]
US 8977809B2 · Johnson · 2015 [cited by applicant]
US 9100006B2 · Ch · 2015 [cited by applicant]
US 9559086B2 · Mei et al. · 2017 [cited by applicant]
US 10002653B2 · Pelley et al. · 2018 [cited by applicant]
US 10079044B2 · Jayasena et al. · 2018 [cited by applicant]
US 20020019843A1 · Killian et al. · 2002 [cited by applicant]
US 20060184923A1 · Pires Dos Reis Moreira · 2006 [cited by examiner]
US 20060271765A1 · Tell et al. · 2006 [cited by applicant]
US 20080195848A1 · Fayad et al. · 2008 [cited by applicant]
US 20110119467A1 · Cadambi et al. · 2011 [cited by applicant]
US 20110296107A1 · Li · 2011 [cited by examiner]
US 20120256653A1 · Cordero et al. · 2012 [cited by applicant]
US 20140009992A1 · Poulton · 2014 [cited by applicant]
US 20140040532A1 · Watanabe et al. · 2014 [cited by applicant]
US 20140085959A1 · Saraswat et al. · 2014 [cited by applicant]
US 20140289445A1 · Savich · 2014 [cited by examiner]
US 20150106560A1 · Perego et al. · 2015 [cited by applicant]
US 20150279431A1 · Li et al. · 2015 [cited by applicant]
US 20160328179A1 · Quinn · 2016 [cited by examiner]
US 20180188702A1 · Jiang · 2018 [cited by examiner]
US 20180210671A1 · Kowles · 2018 [cited by examiner]
US 20180322387A1 · Sridharan et al. · 2018 [cited by applicant]
US 20190088339A1 · Miyashita et al. · 2019 [cited by applicant]
US 20190187930A1 · Chinnakkonda Vidyapoornachary · 2019 [cited by examiner]
US 20200081651A1 · Aga · 2020 [cited by examiner]
US 20200210296A1 · Kang · 2020 [cited by examiner]
US 20200387427A1 · Mekhanik · 2020 [cited by examiner]
US 20210397348A1 · Agarwal · 2021 [cited by examiner]
US 20220269436A1 · Kellam · 2022 [cited by examiner]
CN 108805795A · 2018 [cited by applicant]
Gabriel H. Loh, “3-D-Stacked Memory Architectures for Multi-Core Processors”, 2008, IEEE (Year: 2008). [cited by examiner]
EP Response filed Jan. 2, 2024 in Response to the Extended European Search Report dated Jun. 19, 2023 and the Official Communication Pursuant to Rules 70(2) and 70a(2) EPC dated Jul. 6, 2023. 21 pages. [cited by applicant]
EP Extended European Search Report with Mail Date Jun. 19, 2023 re: EP Appln. No. 20844273.1. 10 pages. [cited by applicant]
Schuiki, Fabian et al., “A Scalable Near-Memory Architecture for Training Deep Neural Networks on Large In-Memory Datasets”, IEEE Transactions on Computers, vol. 68, No. 4, Apr. 2019. 14 pages. [cited by applicant]
Debabani Choudhury “3D Integration Technologies for Emerging Microsystems” 978-1-4244-7732-6/101 © 2010 IEEE. 4 pages. [cited by applicant]
Duckhwan Kim et al. “Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory” 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture. 14 pages. [cited by applicant]
Gentry, “Computing arbitrary functions of encrypted data”, In: Communication of the ACM. Jul. 29, 2010 (Jul. 29, 2010) Retrieved on Oct. 26, 2020 (Oct. 26, 2020) from <http://www.hit.bme.hu/-buttyan/courses/BMEVIHM219/2… [cited by applicant]
Mingyu Gao, “Tetris: Scalable and Efficient Neural Network Acceleration with 3D Memory” 2017 Copyright Publication rights licensed to ACM. ISBN 978-1-4503-4465-4/17/04. 14 pages. [cited by applicant]
Notification of Transmittal of the International Search Report and the Written Opinion of the International Search Authority, or the Declaration with Mail Date Nov. 24, 2020 re: Int'l Appln. No. PCT/US2020/040884. 21 pa… [cited by applicant]
EP Response filed Nov. 7, 2024 in Response to the Official Communication Pursuant to Art. 94(3) EPC dated Jul. 15, 2024. 20 pages. [cited by applicant]
CN Office Action with Mail Date May 6, 2024 re: CN Appln. No. 202080052156.8. 18 pages. (W/translation). [cited by applicant]
EP Communication Pursuant to Article 94(3) EPC with Mail Date Jul. 15, 2024 re: EP Appln. No. 20844273.1. 5 pages. [cited by applicant]
EP communication Pursuant to Article 94(3) EPC with mail date Feb. 5, 2026 re: EP Appln. No. 20844273.1. 6 pages. [cited by applicant]