IP Library › Granted Patent US 12,664,412
Granted Patent B2
US 12,664,412 · App. 18/743,059 · Granted Jun 23, 2026

System and method for in-memory image processing

Inventors: Sarma Vrudhula (Chandler, AZ); Gian Singh (Tempe, AZ); Ayushi Dube (Tempe, AZ)
Assignee: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,412
App. No.
18/743,059
Granted
Jun 23, 2026
Kind
B2
Abstract

A system and method for in-memory image processing. In some embodiments, a system includes a memory, a first neuron processing circuit, and a second neuron processing circuit. The first neuron processing circuit may be connected to a first plurality of bit lines of the memory, and the second neuron processing circuit may be connected to a second plurality of bit lines of the memory. The first neuron processing circuit may include a plurality of configurable processing circuits, each of the configurable processing circuits including: an artificial neuron for calculating the sign of a weighted sum of a plurality of single-bit digital input signals at respective inputs of the artificial neuron, and a plurality of multiplexers, each having an output connected to a respective input of the inputs of the artificial neuron.

Claims (111)

1 . A system comprising:

a memory comprising a Dynamic Random Access Memory (DRAM);

a first neuron processing circuit; and

a second neuron processing circuit,

the first neuron processing circuit being connected to a first plurality of bit lines of the memory,

the second neuron processing circuit being connected to a second plurality of bit lines of the memory,

the first neuron processing circuit comprising a plurality of configurable processing circuits,

each of the configurable processing circuits comprising:

an artificial neuron for calculating the sign of a weighted sum of a plurality of single-bit digital input signals at respective inputs of the artificial neuron, and

a plurality of multiplexers, each having an output connected to a respective input of the inputs of the artificial neuron,

wherein:

the artificial neuron comprises:

a left input network (LIN) to calculate a first contribution to the weighted sum, the first contribution comprising a sum of terms with positive weights;

a right input network (RIN) to calculate a second contribution to the weighted sum, the second contribution comprising a sum of terms with negative weights;

a sense amplifier connected to the LIN and to the RIN; and

a set-reset latch connected to the sense amplifier,

the LIN comprises a first set of switchable current sources,

each switchable current source of the first set of switchable current sources is configured:

to be switched on or off according to whether the value at a respective corresponding input of the LIN is a one or a zero; and

to drive an output of the LIN,

the RIN comprises a second set of switchable current sources,

each switchable current source of the second set of switchable current sources is configured:

to be switched on or off according to whether the value at a respective corresponding input of the RIN is a one or a zero; and

to drive an output of the RIN, and

the DRAM and the first and second neuron processing circuits are in a same semiconductor chip.

2 . The system of claim 1 , wherein each of the configurable processing circuits of the first neuron processing circuit is connected to each of the other configurable processing circuits of the first neuron processing circuit.

3 . The system of claim 1 , wherein each of the configurable processing circuits of the first neuron processing circuit further comprises a register for storing one or more bits.

4 . The system of claim 1 , wherein the weighted sum is a weighted sum of four terms.

5 . The system of claim 4 , wherein the weights corresponding to the four terms are 1, 1, 1, and −2.

6 . The system of claim 1 , wherein the memory, the first neuron processing circuit, and the second neuron processing circuit are fabricated on a single semiconductor chip.

7 . A system comprising:

a memory comprising a Dynamic Random Access Memory (DRAM);

a first neuron processing circuit; and

a second neuron processing circuit,

the first neuron processing circuit being connected to a first plurality of bit lines of the memory,

the second neuron processing circuit being connected to a second plurality of bit lines of the memory,

the first neuron processing circuit comprising:

a plurality of multiplexers and being configured to perform:

a multi-bit addition when the multiplexers receive a first set of control signals, and

a multi-bit comparison when the multiplexers receive a second set of control signals,

a plurality of configurable processing circuits, a first configurable processing circuit of the configurable processing circuits comprising:

an artificial neuron for calculating the sign of a weighted sum of a plurality of single-bit digital input signals at respective inputs of the artificial neuron, the artificial neuron comprising:

a left input network (LIN) to calculate a first contribution to the weighted sum, the first contribution comprising a sum of terms with positive weights;

a right input network (RIN) to calculate a second contribution to the weighted sum, the second contribution comprising a sum of terms with negative weights;

a sense amplifier connected to the LIN and to the RIN; and

a set-reset latch connected to the sense amplifier,

wherein the LIN comprises a first set of switchable current sources,

wherein each switchable current source of the first set of switchable current sources is configured:

to be switched on or off according to whether the value at a respective corresponding input of the LIN is a one or a zero; and

to drive an output of the LIN,

wherein the RIN comprises a second set of switchable current sources,

wherein each switchable current source of the second set of switchable current sources is configured:

to be switched on or off according to whether the value at a respective corresponding input of the RIN is a one or a zero; and

to drive an output of the RIN, and

wherein the DRAM and the first and second neuron processing circuits are in a same semiconductor chip.

8 . The system of claim 7 , wherein the first configurable processing circuit of the configurable processing circuits further comprising:

a plurality of the multiplexers of the first neuron processing circuit, each of the plurality of multiplexers of the first configurable processing circuit having an output connected to a respective input of the inputs of the artificial neuron.

9 . The system of claim 8 , wherein each of the configurable processing circuits of the first neuron processing circuit is connected to each of the other configurable processing circuits of the first neuron processing circuit.

10 . The system of claim 8 , wherein each of the configurable processing circuits of the first neuron processing circuit further comprises a register for storing one or more bits.

11 . The system of claim 8 , wherein the weighted sum is a weighted sum of four terms.

12 . The system of claim 11 , wherein the weights corresponding to the four terms are 1, 1, 1, and −2.

13 . The system of claim 7 , wherein the memory, the first neuron processing circuit, and the second neuron processing circuit are fabricated on a single semiconductor chip.

14 . A method for computing, by a system,

the system comprising:

a memory comprising a Dynamic Random Access Memory (DRAM);

a first neuron processing circuit; and

a second neuron processing circuit,

the first neuron processing circuit being connected to a first plurality of bit lines of the memory,

the second neuron processing circuit being connected to a second plurality of bit lines of the memory,

the first neuron processing circuit comprising a plurality of configurable processing circuits,

a first configurable processing circuit of the configurable processing circuits comprising:

an artificial neuron for calculating the sign of a weighted sum of a plurality of single-bit digital input signals at respective inputs of the artificial neuron, and

a plurality of multiplexers, each having an output connected to a respective input of the inputs of the artificial neuron,

wherein:

the artificial neuron comprises:

a left input network (LIN) to calculate a first contribution to the weighted sum, the first contribution comprising a sum of terms with positive weights;

a right input network (RIN) to calculate a second contribution to the weighted sum, the second contribution comprising a sum of terms with negative weights;

a sense amplifier connected to the LIN and to the RIN; and

a set-reset latch connected to the sense amplifier,

the LIN comprises a first set of switchable current sources,

each switchable current source of the first set of switchable current sources is configured:

to be switched on or off according to whether the value at a respective corresponding input of the LIN is a one or a zero; and

to drive an output of the LIN,

the RIN comprises a second set of switchable current sources,

each switchable current source of the second set of switchable current sources is configured:

to be switched on or off according to whether the value at a respective corresponding input of the RIN is a one or a zero; and

to drive an output of the RIN, and

the method comprising:

applying a first sequence of control signals to the multiplexers of the first configurable processing circuit; and

in response to the first sequence of control signals, calculating a sum, by the first configurable processing circuit,

wherein the DRAM and the first and second neuron processing circuits are in a same semiconductor chip.

15 . The method of claim 14 , further comprising:

applying a second sequence of control signals to the multiplexers of the first configurable processing circuit; and

in response to the first sequence of control signals, performing a comparison, by the first configurable processing circuit.

16 . The method of claim 14 , wherein the sum is a sum of:

a first binary integer, left-shifted by a first number of bit positions, and

a second binary integer, left-shifted by a second number of bit positions,

the second number of bit positions being different from the first number of bit positions.

17 . The method of claim 16 , further comprising forming the second binary integer as a subset of bits of a third binary integer.

18 . The method of claim 14 , further comprising calculating an approximate product of a first integer and a second integer,

the calculating comprising:

subtracting a third integer from the first integer to form a difference, and

left-shifting the difference by a first number of bit positions,

wherein:

the third integer is determined based on the first integer;

the first number of bit positions is based on the second integer; and

the left-shifting of the difference corresponds to multiplying the difference by a fourth integer, the fourth integer being the closest power of 2 to the second integer.

19 . The method of claim 18 , wherein the third integer is determined by:

determining that the first integer is within a first range of values, the first range of values being based on a previously performed comparison, and

setting the third integer equal to a value associated with the first range of values.

20 . The method of claim 14 , wherein the memory, the first neuron processing circuit, and the second neuron processing circuit are fabricated on a single semiconductor chip.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2024
From: VRUDHULA, SARMA; SINGH, GIAN; DUBE, AYUSHI
To: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
Reel/Frame 068838/0117 →
Continuity (2)
Provisional Application 63508253 · Jun 14, 2023
Related Publication 20240419955A1 · Dec 19, 2024
References Cited (48)
US 11823037B1 · Maslov · 2023 [cited by examiner]
US 20120259804A1 · Brezzo · 2012 [cited by examiner]
US 20180075344A1 · Ma · 2018 [cited by examiner]
US 20180373977A1 · Carbon · 2018 [cited by examiner]
US 20210295134A1 · Stevens · 2021 [cited by examiner]
US 20220239924A1 · Kirchhoffer · 2022 [cited by examiner]
US 20230280976A1 · Khwa · 2023 [cited by examiner]
US 20230325645A1 · Tran · 2023 [cited by examiner]
“File: Original lena512.jpg”, MediaWiki, 2016, 2 pages, Retrieved from website: http://boofcv.org/index.php?title=File:Original_lena512.jpg&oldid=2014. [cited by applicant]
Al-Maaitah, K. et al., “Configurable-Accuracy Approximate Adder Design with Light-Weight Fast Convergence Error Recovery Circuit”, 2017 IEEE Jordan Conference on Applied Electrical Engineering and Computing Technologies… [cited by applicant]
Anusha, G. et al., “Design of approximate adders and multipliers for error tolerant image processing”, Microprocessors and Microsystems, Nov. 11, 2019, pp. 1-7. [cited by applicant]
Boroumand, A. et al., “Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks”, Session 4A: Memory 1, Mar. 24-28, 2018, pp. 316-331. [cited by applicant]
Celia, D. et al., “Probabilistic Error Modeling for Two-part Segmented Approximate Adders”, IEEE Xplore, 2018, 5 pages. [cited by applicant]
Chakraborty, I. et al., “Resistive Crossbars as Approximate Hardware Building Blocks for Machine Learning: Opportunities and Challenges”, Proceedings of the IEEE, Jul. 15, 2020, pp. 2276-2310, vol. 108, No. 12. [cited by applicant]
Deng, Q. et al., “LAcc: Exploiting Lookup Table-based Fast and Accurate Vector Multiplication in DRAM-based CNN Accelerator”, Association for Computing Machinery, 2019, 6 pages. [cited by applicant]
Dube, A., et al., “Tunable Precision Control for Approximate Image Filtering in an In-Memory Architecture with Embedded Neurons”, ICCAD '22, Oct. 30-Nov. 3, 2022, 9 pages, San Diego, CA, USA. [cited by applicant]
Giacomin, E. et al., “A Robust Digital RRAM-Based Convolutional Block for Low-Power Image Processing and Learning Applications”, IEEE Transactions on Circuits and Systems—I: Regular Papers, Oct. 9, 2018, pp. 643-654, vo… [cited by applicant]
Gu, P. et al., “iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank Architecture”, 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), 2020, pp. 804-817. [cited by applicant]
Gupta, V. et al., “Low-Power Digital Signal Processing Using Approximate Adders”, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Dec. 19, 2012, pp. 124-137, vol. 32, No. 1. [cited by applicant]
Haj-Ali, A. et al., “IMAGING: In-Memory AlGorithms for Image processING”, IEEE Transactions on Circuits and Systems—I: Regular Papers, Jun. 27, 2018, pp. 4258-4271, vol. 65, No. 12. [cited by applicant]
He, M. et al., “Newton: A DRAM-maker's Accelerator-in-Memory (AiM) Architecture for Machine Learning”, 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2020, pp. 372-385. [cited by applicant]
Jha, C. et al., “Quality Tunable Approximate Adder for Low Energy Image Processing Applications”, IEEE Xplore, 2019, pp. 642-645. [cited by applicant]
Jung, M. et al., “A New Bank Sensitive DRAMPower Model for Efficient Design Space Exploration”, IEEE Xplore, 2016, pp. 283-288. [cited by applicant]
Kan, Y. et al., “A Multi-Grained Reconfigurable Accelerator for Approximate Computing”, 2020 IEEE Computer Society Annual Symposium on VLSI (ISVLSI), 2020, pp. 90-95. [cited by applicant]
Khaveen Investments, “Micron, Samsung, And SK Hynix: The DRAMOligopoly”, May 12, 2020, 12 pages. [cited by applicant]
Kim, Y. et al., “Ramulator: A Fast and Extensible DRAM Simulator”, IEEE Computer Architecture Letters, Mar. 17, 2015, pp. 45-49, vol. 15, No. 1. [cited by applicant]
Kim, Y-B. et al., “Assessing Merged DRAM/Logic ‘Technology’”, IEEE Xplore, 1996, pp. 133-136. [cited by applicant]
Kwon, Y-C. et al., “A 20nm 6GB Function-In-Memory DRAM, Based on HBM2 with a 1.2TFLOPS Programmable Computing Unit Using Bank-Level Parallelism, for Machine Learning Applications”, 2021 IEEE International Solid-State Ci… [cited by applicant]
Lee, J. et al., “A Novel Approximate Adder Design Using Error Reduced Carry Prediction and Constant Truncation”, IEEE Access, Aug. 27, 2021, pp. 119939-119953, vol. 9. [cited by applicant]
Lei, L. et al., “Joint Computation Offloading and Multiuser Scheduling Using Approximate Dynamic Programming in NB-IoT Edge Computing System”, IEEE Internet Of Things Journal, Feb. 21, 2019, pp. 5345-5362, vol. 6, No. 3. [cited by applicant]
Liu, B. et al.,“Binarized Weight Neural-network Inspired Ultra-low Power Speech Recognition Processor with Time-domain based Digital-analog Mixed Approximate Computing”, IEEE, 2020, 5 pages. [cited by applicant]
Liu, B., et al., “EERA-ASR: An Energy-Efficient Reconfigurable Architecture for Automatic Speech Recognition With Hybrid DNN and Approximate Computing”, IEEE Access, 2018, pp. 52227-52237, vol. 6. [cited by applicant]
Mazahir, S., et al., “Adaptive Approximate Computing in Arithmetic Datapaths”, IEEE Design & Test, Jul./Aug. 2018, pp. 65-74, IEEE CEDA, IEEE CASS, IEEE SSCS and TTTC. [cited by applicant]
Prabakaran, B.S., et al., “ApproxFPGAs: Embracing ASIC-Based Approximate Arithmetic Components for FPGA-Based Systems”, 2020, 6 pages, IEEE Xplore. [cited by applicant]
Salavati, S., et al., “Ultra-Efficient Nonvolatile Approximate Full-Adder With Spin-Hall-Assisted MTJ Cells for In-Memory Computing Applications”, IEEE Transactions on Magnetics, May 2021, vol. 57, No. 5. [cited by applicant]
San Miguel, J., et al., “Load Value Approximation”, 2014 47 [cited by applicant]
Seshadri, V., et al., “Ambit: In-Memory Accelerator for Bulk Bitwise Operations Using Commodity DRAM Technology”, MICRO-50, Oct. 14-18, 2017, pp. 273-287, Association for Computing Machinery, Cambridge, MA, USA. [cited by applicant]
Singh, G., et al., “CIDAN-XE: Computing in DRAM with Artificial Neurons”, Frontiers in Electronics, Feb. 2022, pp. 1-17, vol. 3, Article 834146. [cited by applicant]
Soares, L.B., et al., “Design Methodology to Explore Hybrid Approximate Adders for Energy-Efficient Image and Video Processing Accelerators”, IEEE Transactions on Circuits and Systems—I: Regular Papers, Jun. 2019, pp. 2… [cited by applicant]
Sutradhar, P.R., et al., “pPIM: A Programmable Processor-in-Memory Architecture With Precision-Scaling for Deep Learning”, IEEE Computer Architecture Letters, Jul.-Dec. 2020, pp. 118-121, vol. 19, No. 2. [cited by applicant]
Wagle, A., et al., “A Configurable BNN ASIC using a Network of Programmable Threshold Logic Standard Cells”, ICCD, 2020, pp. 433-440, IEEE. [cited by applicant]
Wagle, A., et al., “Threshold Logic in a Flash”, ICCD, 2019, pp. 550-558, IEEE. [cited by applicant]
Wulf, W. A., et al., “Hitting the Memory Wall: Implications of the Obvious”, Dec. 1994, pp. 20-24. [cited by applicant]
Xu, W., et al., “A Simple Yet Efficient Accuracy-Configurable Adder Design”, IEEE Transactions on Very large Scale Integration (VLSI) Systems, Jun. 2018, pp. 1112-1125, vol. 26, No. 6, IEEE. [cited by applicant]
Yantir, H.E., et al., “A Hybrid Approximate Computing Approach for Associative In-Memory Processors”, IEEE Journal on Emerging and Selected Topics in Circuits and Systems, Dec. 2018, pp. 758-769, IEEE. [cited by applicant]
Yin, S., et al., “Vesti: Energy-Efficient In-Memory Computing Accelerator for Deep Neural Networks”, IEEE Transactions on Very large Scale Integration (VLSI) Systems, Jan. 2020, pp. 48-61, vol. 28, No. 1, IEEE. [cited by applicant]
Zayer, F., et al., “RRAM Crossbar-Based In-Memory Computation of Anisotropic Filters for Image Preprocessing”, IEEE Access, 2020, pp. 127569-127580, vol. 8, Creative Commons Attribution. [cited by applicant]
Zhu, N. et al., “Enhanced Low-Power High-Speed Adder For Error-Tolerant Application”, IEEE Xplore, 2010, pp. 323-327. [cited by applicant]