IP Library Granted Patent US 12,554,656
Granted Patent B2
US 12,554,656 · App. 17/214,778 · Granted Feb 17, 2026

Systems, methods, and devices for near data processing

Inventors: Wenqin Huangfu (Goleta, CA); Krishna T. Malladi (San Jose, CA); Dongyan Jiang (San Jose, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F12/1433G06F9/5016G06F13/1673G06F13/4234G06F2209/508
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,656
App. No.
17/214,778
Granted
Feb 17, 2026
Kind
B2
Abstract

A memory module may include one or more memory devices, and a near-memory computing module coupled to the one or more memory devices, the near-memory computing module including one or more processing elements configured to process data from the one or more memory devices, and a memory controller configured to coordinate access of the one or more memory devices from a host and the one or more processing elements. A method of processing a dataset may include distributing a first portion of the dataset to a first memory module, distributing a second portion of the dataset to a second memory module, constructing a first local data structure at the first memory module based on the first portion of the dataset, constructing a second local data structure at the second memory module based on the second portion of the dataset, and merging the first and second local data structures.

Claims (73)

1 . An apparatus comprising:

a memory module comprising:

a first rank of one or more first memory devices;

a near-memory computing circuit coupled to one of the one or more first memory devices, the near-memory computing circuit comprising:

a first processing element configured to process data from the one or more first memory devices; and

a first memory controller configured to coordinate access of the one or more first memory devices by a host and the first processing element;

a second rank of one or more second memory devices;

a bus structure configured to transfer data between the first rank and the second rank; and

wherein the memory module is configured to:

send a first portion of a dataset to the one or more first memory devices, wherein the one or more first memory devices are coupled to the first processing element and a first connector to transfer the first portion of the dataset;

send a second portion of the dataset to the one or more second memory devices, wherein the one or more second memory devices are coupled to a second processing element and a second connector to transfer the second portion of the dataset;

construct, using the first processing element, a first data structure at the one or more first memory devices based on the first portion of the dataset;

construct, using the second processing element, a second data structure at the one or more second memory devices based on the second portion of the dataset;

merge the first data structure and the second data structure to generate a merged data structure;

send at least a first portion of the merged data structure to the one or more first memory devices; and

send a second portion of the merged data structure to the one or more second memory devices.

2 . The apparatus of claim 1 , wherein the first processing element is configured to process data from the one or more first memory devices by performing a counting operation on the data.

3 . The apparatus of claim 1 , wherein:

the near-memory computing circuit comprises a first near-memory computing circuit; and

the second rank comprises:

one or more second memory devices; and

a second near-memory computing circuit coupled to the one or more second memory devices, the second near-memory computing circuit comprising:

the second processing element configured to process data from the one or more second memory devices; and

a second memory controller configured to coordinate access of the one or more second memory devices from a host and the second processing element.

4 . The apparatus of claim 1 , wherein the near-memory computing circuit further comprises a workload monitor configured to modify a first workload of a first one of the one or more processing elements based on a second workload of a second one of the one or more processing elements.

5 . The apparatus of claim 1 , wherein the signal comprises a control signal or an address signal.

6 . The apparatus of claim 1 , wherein the signal comprises a data signal.

7 . A method comprising:

sending a first portion of a dataset to a first memory module comprising a first processing element and a first connector to transfer the first portion of the dataset;

sending a second portion of the dataset to a second memory module comprising a second processing element and a second connector to transfer the second portion of the dataset;

constructing, using the first processing element, a first data structure at the first memory module based on the first portion of the dataset;

constructing, using the second processing element, a second data structure at the second memory module based on the second portion of the dataset;

merging the first data structure and the second data structure to generate a merged data structure;

sending at least a first portion of the merged data structure to the first memory module; and

sending a second portion of the merged data structure to the second memory module.

8 . The method of claim 7 , wherein:

the method further comprises performing a counting operation on the merged data structure at the first memory module and the second memory module.

9 . The method of claim 7 , wherein:

the merging the first data structure and the second data structure comprises reducing the first data structure and the second data structure.

10 . The method of claim 9 , further comprising sending the first portion of the dataset to two or more ranks at the first memory module.

11 . The method of claim 7 , further comprising modifying a first workload of the first processing element based on a second workload of the second processing element.

12 . The method of claim 7 , further comprising:

accessing first data of the first data structure;

performing, using the first data, a task;

determining a result of the task; and

accessing, based on the result of the task, second data of the first data structure.

13 . The method of claim 7 , further comprising:

performing a first task based on a first memory access of the first portion of the dataset; and

performing a second task based on a second memory access of the first portion of the dataset.

14 . The method of claim 7 , wherein the merged data structure is a first merged data structure and the portion of the merged data structure is a first portion of the merged data structure, the method further comprising:

sending a second portion of the first merged data structure to the second memory module;

constructing a third data structure at the first memory module based on the first merged data structure;

constructing a fourth data structure at the first memory module based on the first merged data structure;

merging the third data structure and the fourth data structure to form a second merged data structure; and

performing a counting operation on the second merged data structure at the first memory module and the second memory module.

15 . The method of claim 7 , wherein:

the dataset comprises a genetic sequence;

the first data structure comprises a Bloom filter; and

the Bloom filter comprises one or more k mers of the genetic sequence.

16 . The method of claim 7 , wherein the first memory module comprises a first memory device and a second memory device, the method further comprising sending the first portion of the dataset to the first memory device and the second memory device based on a location of device information in an address for the first portion of the dataset.

17 . The method of claim 7 , wherein:

the first data structure comprises a first counting filter; and

the second data structure comprises a second counting filter.

18 . A system comprising:

a first memory module configured to construct a first local data structure based on a first portion of a dataset, the first memory module comprising a first connector to transfer the first portion of the dataset;

a second memory module configured to construct a second local data structure based on a second portion of the dataset, the second memory module comprising a second connector to transfer the second portion of the dataset; and

a host coupled to the first memory module and the second memory module using one or more memory channels, wherein the host is configured to:

send the first portion of the dataset to the first memory module;

send the second portion of the dataset to the second memory module;

merge the first local data structure and the second local data structure to generate a merged data structure;

send a first portion of the merged data structure to the first memory module; and

send a second portion of the merged data structure to the second memory module.

19 . The system of claim 18 , wherein the first memory module is configured to perform a counting operation on the first portion of the merged data structure.

Continuity (2)
Provisional Application 63021675 · May 7, 2020
Related Publication 20210349837A1 · Nov 11, 2021
References Cited (51)
US 9703925B1 · Aerni et al. · 2017 [cited by applicant]
US 10185499B1 · Wang · 2019 [cited by examiner]
US 10303662B1 · Menezes · 2019 [cited by examiner]
US 10678717B2 · Li et al. · 2020 [cited by applicant]
US 10776308B2 · Ong et al. · 2020 [cited by applicant]
US 10846233B2 · Kang et al. · 2020 [cited by applicant]
US 10884672B2 · Song et al. · 2021 [cited by applicant]
US 10915470B2 · Kim et al. · 2021 [cited by applicant]
US 20040078544A1 · Lee · 2004 [cited by examiner]
US 20040153447A1 · Hiratsuka · 2004 [cited by examiner]
US 20080005516A1 · Meinschein · 2008 [cited by examiner]
US 20090228684A1 · Ramesh · 2009 [cited by examiner]
US 20130138923A1 · Barber · 2013 [cited by examiner]
US 20140195725A1 · Bennett · 2014 [cited by examiner]
US 20160132640A1 · Layer · 2016 [cited by examiner]
US 20160299874A1 · Liao · 2016 [cited by examiner]
US 20180081583A1 · Breternitz et al. · 2018 [cited by applicant]
US 20180122434A1 · Lee · 2018 [cited by examiner]
US 20180253254A1 · Jain · 2018 [cited by examiner]
US 20190171566A1 · Jung · 2019 [cited by examiner]
US 20190212918A1 · Wang et al. · 2019 [cited by applicant]
US 20190266159A1 · Chainani et al. · 2019 [cited by applicant]
US 20190384845A1 · Saxena · 2019 [cited by examiner]
US 20200341775A1 · Mathew et al. · 2020 [cited by applicant]
EP 3396533A2 · 2018 [cited by applicant]
TW 201937374A · 2019 [cited by applicant]
WO 2019114979A1 · 2019 [cited by applicant]
US 10,853,278 B2, 12/2020, Kim et al. (withdrawn) [cited by applicant]
Farmahini-Farahani, Amin, et al., “NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,” 2015 IEEE 21st International Symposium on High Performance Computer Architectur… [cited by applicant]
Huangfu, Wenqin, et al., “MEDAL: Scalable DIMM based Near Data Processing Accelerator for DNA Seeding Algorithm,” MICRO '19: Proceedings of the 52nd Annual IEEE, Acm International Symposium on Microarchitecture, Oct. 20… [cited by applicant]
Bloom, Burton H., “Space/time trade-offs in hash coding with allowable errors”, Communications of the ACM, vol. 13, Issue 7, pp. 422-426, Jul. 1970. [cited by applicant]
Bose, Simmi M. et al., “k-Core: Hardware Accelerator for k-Mer Generation and Counting used in Computational Genomics”, 2019 32nd International Conference on VLSI Design and 2019 18th International Conference on Embedde… [cited by applicant]
Cadenelli, Nicola et al., “Considerations in using OpenCL on GPUs and FPGAs for throughput-oriented genomics workloads”, Future Generation Computer Systems, vol. 94, May 2019, pp. 148-159. [cited by applicant]
Chandrasekar, Karthik et al., “DRAMPower: Open-source DRAM Power & Energy Estimation Tool”, 2012, URL: http://www.drampower.info. [cited by applicant]
Chikhi, Rayan et al., “Space-efficient and exact de Bruijn graph representation based on a Bloom filter”, Algorithms for Molecular Biology, vol. 8, Article No. 22, Sep. 16, 2013. [cited by applicant]
Cong, Jason et al., “AIM: accelerating computational genomics through scalable and noninvasive accelerator-interposed memory”, MEMSYS '17: Proceedings of the International Symposium on Memory Systems, Oct. 2017, pp. 3-1… [cited by applicant]
Deorowicz, Sebastian et al., “KMC 2: fast and resource-frugal k-mer counting”, Bioinformatics, vol. 31, Issue 10, May 2015, pp. 1569-1576. [cited by applicant]
Erbert, Marius et al., “Gerbil: a fast and memory-efficient k-mer counter with GPU-support”, Algorithms for Molecular Biology 12, Article No. 9 (2017). [cited by applicant]
Guan, Wei-jie et al., “Clinical Characteristics of Coronavirus Disease 2019 in China”, New England Journal of Medicine, Apr. 30, 2020. [cited by applicant]
Joardar, Biresh Kumar et al., “NoC-enabled software/hardware co-design framework for accelerating k-mer counting”, NOCS '19: Proceedings of the 13th IEEE/ACM International Symposium on Networks-on-Chip, Oct. 2019, Artic… [cited by applicant]
Jouppi, Norman P. et al., “CACTI-IO: CACTI With OFF-Chip Power-Area-Timing Models”, IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 23, No. 7, pp. 1254-1267, Jul. 2015. [cited by applicant]
Kim, Yoongu et al., “Ramulator: A Fast and Extensible DRAM Simulator”, IEEE Computer Architecture Letters, vol. 15, Issue: 1, Jan.-Jun. 2016, pp. 45-49. [cited by applicant]
Lu, Roujian et al., “Genomic characterisation and epidemiology of 2019 novel coronavirus: implications for virus origins and receptor binding”, The Lancet, Feb. 22-28, 2020, 395(10224), pp. 565-574. [cited by applicant]
Mcvicar, Nathaniel et al., “K-Mer Counting Using Bloom Filters with an FPGA-Attached HMC”, 2017 IEEE 25th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), Napa, CA, USA, 2017, pp. 2… [cited by applicant]
Melsted, Pall et al., “Efficient counting of k-mers in DNA sequences using a bloom filter”, BMC bioinformatics 12, Article 333, 2011. [cited by applicant]
Merker, Jason D. et al., “Long-read genome sequencing identifies causal structural variation in a Mendelian disease”, Genetics in Medicine 20, 1 (2018), 159-163. [cited by applicant]
Pan, Tony C. et al., “Optimizing High Performance Distributed Memory Parallel Hash Tables for DNA k-mer Counting”, SC18: International Conference for High Performance Computing, Networking, Storage and Analysis, Dallas,… [cited by applicant]
Pugsley, Seth H. et al., “Comparing Implementations of Near-Data Computing with In-Memory MapReduce Workloads”, IEEE Micro, vol. 34, No. 4, pp. 44-52, Jul.-Aug. 2014. [cited by applicant]
Ramachandran, Anand et al., “FPGA accelerated DNA error correction”, 2015 Design, Automation & Test in Europe Conference & Exhibition (Date), Grenoble, France, 2015, pp. 1371-1376. [cited by applicant]
Shendure, Jay et al., “Next-generation DNA sequencing”, Nature biotechnology 26, pp. 1135-1145, 2008. [cited by applicant]
SYNOPSYS 2018. Design Complier. https://www.synopsys.com/support/training/rtl-synthesis/design-compiler-rtl-synthesis.html. [cited by applicant]