IP Library › Granted Patent US 12,333,304
Granted Patent B2
US 12,333,304 · App. 18/582,520 · Granted Jun 17, 2025

Methods for performing processing-in-memory operations, and related systems

Inventors: Dmitri Yudanov (Cordova, CA); Sean S. Eilert (Penryn, CA); Sivagnanam Parthasarathy (Carlsbad, CA); Shivasankar Gunasekaran (Folsom, CA); Ameen D. Akel (Rancho Cordova, CA)
Assignee: Micron Technology, Inc.
G06F9/3001G06F7/5443G06F9/30032G06F9/30043
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,304
App. No.
18/582,520
Granted
Jun 17, 2025
Kind
B2
Abstract

Methods, apparatuses, and systems for in-or near-memory processing are described. Strings of bits (e.g., vectors) may be fetched and processed in logic of a memory device without involving a separate processing unit. Operations (e.g., arithmetic operations) may be performed on numbers stored in a bit-parallel way during a single sequence of clock cycles. Arithmetic may thus be performed in a single pass as numbers are bits of two or more strings of bits are fetched and without intermediate storage of the numbers. Vectors may be fetched (e.g., identified, transmitted, received) from one or more bit lines. Registers of a memory array may be used to write (e.g., store or temporarily store) results or ancillary bits (e.g., carry bits or carry flags) that facilitate arithmetic operations. Circuitry near, adjacent, or under the memory array may employ XOR or AND (or other) logic to fetch, organize, or operate on the data.

Claims (46)

1. A system, comprising:

logic configured to:

multiply each bit of a number of bits from an array by a first bit of an input vector from a sequencer to generate a first row of bits;

multiply each bit of the number of bits by one or more additional bits of the input vector to generate one or more additional rows of bits; and

generate an output row of bits based on at least the first row of bits and the one or more additional rows of bits.

2. The system of claim 1 , the logic comprising one or more fused-multiply-add (FMA) units to generate the one or more additional rows of bits.

3. The system of claim 1 , the logic comprising:

the sequencer configured to receive the input vector; and

a sense amplifier configured to receive the number of bits in a bit-parallel manner.

4. The system of claim 1 , wherein the logic is further configured to:

multiply each bit of a second number of bits from the array by a first bit of a second input vector from the sequencer to generate a third row of bits;

multiply each bit of the second number of bits by one or more additional bits of the second input vector to generate at least one additional row of bits;

generate an additional output row of bits based on at least the third row of bits and the at least one additional row of bits; and

sum, along columns, the output row and the additional output row to generate a row of an output matrix.

5. The system of claim 1 , further comprising a memory device including the logic, wherein the memory device includes the sequencer.

6. The system of claim 1 , further comprising a memory device including the logic, wherein the sequencer is external to the memory device.

7. The system of claim 1 , wherein the logic comprises a number of groups of sense amplifiers, each group of sense amplifiers of the number of groups of sense amplifiers configured to receive a tile associated with a portion of the number of bits.

8. The system of claim 1 , wherein the logic is configured to sum, along associated bit positions, the first row of bits and the one or more additional rows of bits.

9. The system of claim 1 , further comprising:

at least one input device;

at least one output device;

at least one processor device coupled to the input device and the output device; and

at least one memory device coupled to the at least one processor device and comprising the logic.

10. A method, comprising:

multiplying each bit of a number of bits from an array by a first bit of an operand from a sequencer to generate a first row of bits;

multiplying each bit of the number of bits by one or more additional bits of the operand to generate one or more additional rows of bits; and

generating an output row based on at least the first row of bits and the one or more additional rows of bits.

11. The method of claim 10 , further comprising:

multiplying each bit of a second number of bits from the array by a second bit of the operand to generate a third row of bits;

multiplying each bit of the second number of bits by at least one additional bit of the operand to generate at least one additional row of bits; and

generating another output row based on at least the third row of bits and the at least one additional rows of bits.

12. The method of claim 10 , further comprising loading a plurality of bits of the operand from a second array into the sequencer.

13. The method of claim 10 , further comprising loading the number of bits from the array into a second array.

14. A memory system, comprising:

at least one memory array; and

logic coupled to the at least one memory array and configured to:

multiply each bit of a group of bits by each bit of an input to generate a number of rows, each row of the number of rows including a number of columns and each row of each of the number of rows shifted at least one column position relative to an adjacent row within the number of rows; and

generate an output row based on the number of rows.

15. The memory system of claim 14 , wherein the logic is configured to sum bits across columns within the number of rows to generate the output row.

16. The memory system of claim 14 , the logic further configured to:

multiply each bit of a second group of bits by each bit of the input to generate a number of additional rows, each row of the number of additional rows including a number of additional columns and each row of each of the number of additional rows shifted at least one column position relative to an adjacent row within the number of additional rows; and

generate an additional output row based on the number of additional rows.

17. The memory system of claim 16 , wherein the logic is configured to sum bits across columns within the number of additional rows to generate the additional output row.

18. The memory system of claim 14 , the logic further configured to receive the group of bits from a memory array of the at least one memory array, wherein each bit of the group of bits is received at the logic substantially simultaneously.

19. The memory system of claim 14 , wherein the logic comprises a sequencer including the input.

20. The memory system of claim 14 , wherein the logic comprises a sense amplifier configured to receive the group of bits in a bit-parallel manner.

Continuity (3)
Continuation 16841222 · Apr 6, 2020
Provisional Application 62896228 · Sep 5, 2019
Related Publication 20240192953A1 · Jun 13, 2024
References Cited (74)
US 6484231B1 · Kim · 2002 [cited by applicant]
US 8762445B2 · Duncan · 2014 [cited by examiner]
US 9430735B1 · Vali et al. · 2016 [cited by applicant]
US 9496023B2 · Wheeler et al. · 2016 [cited by applicant]
US 9704540B2 · Manning et al. · 2017 [cited by applicant]
US 10074407B2 · Manning et al. · 2018 [cited by applicant]
US 10249350B2 · Manning et al. · 2019 [cited by applicant]
US 10416927B2 · Lea et al. · 2019 [cited by applicant]
US 10453499B2 · Manning et al. · 2019 [cited by applicant]
US 10490257B2 · Wheeler et al. · 2019 [cited by applicant]
US 10497442B1 · Kumar et al. · 2019 [cited by applicant]
US 10635398B2 · Lin et al. · 2020 [cited by applicant]
US 10642922B2 · Knag et al. · 2020 [cited by applicant]
US 11537861B2 · Yudanov et al. · 2022 [cited by applicant]
US 20110040821A1 · Eichenberger et al. · 2011 [cited by applicant]
US 20110055517A1 · Eichenberger et al. · 2011 [cited by applicant]
US 20110153707A1 · Ginzburg et al. · 2011 [cited by applicant]
US 20120182799A1 · Bauer · 2012 [cited by applicant]
US 20130173888A1 · Hansen et al. · 2013 [cited by applicant]
US 20150006810A1 · Busta et al. · 2015 [cited by applicant]
US 20150357024A1 · Hush et al. · 2015 [cited by applicant]
US 20160004508A1 · Elmer · 2016 [cited by applicant]
US 20160064045A1 · La Fratta · 2016 [cited by applicant]
US 20160099073A1 · Cernea · 2016 [cited by applicant]
US 20160321031A1 · Hancock · 2016 [cited by applicant]
US 20170277659A1 · Akerib et al. · 2017 [cited by applicant]
US 20170345481A1 · Hush · 2017 [cited by applicant]
US 20180129474A1 · Ahmed · 2018 [cited by applicant]
US 20180210994A1 · Wuu et al. · 2018 [cited by applicant]
US 20180246855A1 · Redfern et al. · 2018 [cited by applicant]
US 20180286468A1 · Willcock · 2018 [cited by applicant]
US 20190019538A1 · Li et al. · 2019 [cited by applicant]
US 20190035449A1 · Saida et al. · 2019 [cited by applicant]
US 20190042199A1 · Sumbul et al. · 2019 [cited by applicant]
US 20190043560A1 · Sumbul et al. · 2019 [cited by applicant]
US 20190080230A1 · Hatcher et al. · 2019 [cited by applicant]
US 20190102170A1 · Chen et al. · 2019 [cited by applicant]
US 20190102358A1 · Asnaashari et al. · 2019 [cited by applicant]
US 20190221243A1 · Manning et al. · 2019 [cited by applicant]
US 20190228301A1 · Thorson et al. · 2019 [cited by applicant]
US 20200020393A1 · Al-Shamma · 2020 [cited by applicant]
US 20200035305A1 · Choi et al. · 2020 [cited by applicant]
US 20200082871A1 · Wheeler et al. · 2020 [cited by applicant]
US 20200210369A1 · Song · 2020 [cited by applicant]
US 20200342938A1 · Tran et al. · 2020 [cited by applicant]
US 20200357459A1 · Zidan et al. · 2020 [cited by applicant]
US 20210104551A1 · Lin et al. · 2021 [cited by applicant]
US 20210110235A1 · Hoang et al. · 2021 [cited by applicant]
CN 101331554A · 2008 [cited by applicant]
CN 101789234A · 2010 [cited by applicant]
CN 105027212A · 2015 [cited by applicant]
CN 105556607A · 2016 [cited by applicant]
CN 106445468A · 2017 [cited by applicant]
CN 106485321A · 2017 [cited by applicant]
CN 108595239A · 2018 [cited by applicant]
CN 109427384A · 2019 [cited by applicant]
EP 0290111A2 · 1988 [cited by applicant]
KR 1020140033937A · 2014 [cited by applicant]
KR 1020170138143A · 2017 [cited by applicant]
KR 1020180042111A · 2018 [cited by applicant]
KR 1020190051766A · 2019 [cited by applicant]
WO 2018154268A1 · 2018 [cited by applicant]
Chinese First Office Action for Chinese Application No. 202080061754.1, dated Jun. 28, 2024, 20 pages with translation. [cited by applicant]
Korean Notice of Reasons for Rejection for Korean Application No. 10-2022-7010454, dated Jun. 11, 2024, 16 pages with English translation. [cited by applicant]
Chiu et al, “A Binarized Neural Network Accelerator with Differential Crosspoint Memristor Array for Energy-Efficient MAC Operations,” 2019 IEEE International Symposium on Circuits and Systems (ISCAS), 2019, pp. 1-5, do… [cited by applicant]
European Extended Search Report and Opinion for European Application No. 20861928.8, dated Jun. 27, 2023, 8 pages. [cited by applicant]
International Search Report for Application No. PCT/US2020/070418, mailed Nov. 27, 2020, 5 pages. [cited by applicant]
International Search Report for Application No. PCT/US2020/070450, mailed Nov. 25, 2020, 3 pages. [cited by applicant]
Patterson et al., “Computer Organization and Design: The Hardware/Software Interface”, Fifth Edition, Chapter 3.2 pp. 178-182, 2014. Retrieved from <https://ict.iitk.ac.in/wp-content/uploads/CS422-Computer-Architecture-… [cited by applicant]
Written Opinion of the International Searching Authority for Application No. PCT/US2020/070418, mailed Nov. 27, 2020, 7 pages. [cited by applicant]
Written Opinion of the International Searching Authority for Application No. PCT/US2020/070450, mailed Nov. 25, 2020, 5 pages. [cited by applicant]
Yudanov et al., U.S. Appl. No. 16/717,890 titled Methods for Performing Processing-in-Memory Operations on Serially Allocated Data, and Related Memory Devices and Systems filed Dec. 17, 2019. [cited by applicant]
Chinese Notice of Allowance and Search Report for Chinese Application No. 202080061754.1, dated Nov. 6, 2024, 4 pages. [cited by applicant]
Zhang et al., “Design of High Performance Memory Controller Based on SOC”, Microelectronics & Computer, No. 5, (2015), 5 pages, English Translation of Abstract. [cited by applicant]