IP Library › Granted Patent US 12,248,696
Granted Patent B2
US 12,248,696 · App. 17/340,866 · Granted Mar 11, 2025

Techniques to repurpose static random access memory rows to store a look-up-table for processor-in-memory operations

Inventors: Saurabh Jain (Bengaluru, IN); Srivatsa Rangachar Srinivasa (Hillsboro, OR); Akshay Krishna Ramanathan (State College, PA); Gurpreet Singh Kalsi (Bangalore, IN); Kamlesh R. Pillai (Bangalore, IN); Sreenivas Subramoney (Bangalore, IN)
Assignee: Intel Corporation
G06F3/0655G06F3/0604G06F3/0673G06F7/523
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,696
App. No.
17/340,866
Granted
Mar 11, 2025
Kind
B2
Abstract

Example compute-in-memory (CIM) or processor-in-memory (PIM) techniques using repurposed or dedicated static random access memory (SRAM) rows of an SRAM sub-array to store look-up-table (LUT) entries for use in a multiply and accumulate (MAC) operation.

Claims (47)

1. A memory device comprising:

a plurality of static random access memory (SRAM) sub-arrays; and

a plurality of compute circuitry, individual compute circuitry from among the plurality of compute circuitry arranged to access data maintained in a corresponding SRAM sub-array from among the plurality of SRAM sub-arrays to enable the individual compute circuitry to execute a multiply and accumulate (MAC) operation, the plurality of SRAM sub-arrays to each include:

a first precharge circuitry to precharge bitlines coupled with all bit cell rows; and

a second precharge circuitry to precharge bitlines coupled with a portion of the bit cell rows,

wherein, responsive to a first signal received from a corresponding individual compute circuitry, the corresponding SRAM sub-array causes switches separately included in the corresponding SRAM sub-array to open respective bitline circuits to isolate the bitlines coupled with all bit cell rows from the bitlines coupled with the portion of bit cell rows, the portion of bit cell rows to store look up table (LUT) entries that include multiplication results and to store a configuration block that includes instructions for the corresponding individual compute circuitry to execute a MAC operation using the multiplication results in the LUT entries.

2. The memory device of claim 1 , comprising the corresponding SRAM sub-array to activate the second precharge circuitry following the opening of the respective bitline circuits to provide access to selected bit cells included in the portion of bit cell rows responsive to a read or a write request from the corresponding individual compute circuitry.

3. The memory device of claim 1 , further comprising the corresponding SRAM sub-array to activate the second precharge circuitry following the opening of the respective bitline circuits to provide access to selected bit cells included in the portion of bit cell rows responsive to a read or a write request from the corresponding individual compute circuitry.

4. The memory device of claim 1 , wherein the corresponding SRAM sub-array, responsive to a second signal received from the corresponding individual compute circuitry, causes the switches to close the respective bitline circuits, the corresponding SRAM sub-array to activate the first precharge circuitry to provide access to selected bit cells included in all bit cell rows responsive to a read or a write request from the corresponding individual compute circuitry.

5. The memory device of claim 1 , comprising the plurality of SRAM sub-arrays and the plurality of compute circuitry included in a sub-bank of multiple sub-banks, the multiple sub-banks included in a bank of multiple banks, the multiple banks included in a slice of multiple slices of a cache for a processor, the cache to include a cache controller to configure a slice controller to control the slice, the slice controller to cause the configuration block to be stored to the portion of bit cell rows coupled with the second precharge circuitry, wherein the corresponding individual compute circuitry is to send the first signal responsive to an enable signal received from the slice controller.

6. The memory device of claim 5 , the slice controller to send the enable signal to the plurality of compute circuitry responsive to the cache for the processor being switched to an accelerator mode to execute a processor-in-memory operation associated with a deep neural network workload.

7. The memory device of claim 1 , the LUT entries that include multiplication results comprises multiplication results based on a product of multiplying 7 odd valued, 4-bit operands with the same 7 odd valued, 4-bit operands, the 7 odd valued, 4-bit operands to include odd values of 3, 5, 7, 9, 11, 13 and 15, wherein the LUT entries include 49, 8-bit values.

8. The memory device of claim 7 , the configuration block that includes instructions for the corresponding compute circuitry to execute the MAC operation comprises the instructions to include an indication to execute a matrix multiplication MAC operation with 4-bit operands, a start address for the 4-bit operands, and an end address for the 4-bit operands, the start and the end addresses located in bit cell rows of the corresponding SRAM sub-array that are separate from the portion of bit cell rows to store the LUT entries.

9. The memory device of claim 8 , wherein if both 4-bit operands to be multiplied as part of the matrix multiplication MAC operation have odd values, the corresponding compute circuitry to access product results of the multiplied, odd 4-bit operands of the matrix multiplication from the LUT entries.

10. The memory device of claim 8 , wherein if both 4-bit operands to be multiplied as part of the matrix multiplication MAC operation have even values that are non-powers of two, the corresponding compute circuitry is to decompose both 4-bit operands into multiples of an odd number value and a power of two, and shift a partial product obtained from the LUT entries.

11. The memory device of claim 8 , wherein if either of the 4-bit operands to be multiplied as part of the matrix multiplication MAC operation are powers of two, the corresponding compute circuitry is to shift a non-power of two operand's value by a value of a power of two operand.

12. A system comprising:

a processor; and

a static random access memory (SRAM) cache for the processor that includes:

a cache controller; and

a plurality of slices, each slice to include a slice controller, each slice to also include a plurality of banks, each bank to include a plurality of sub-banks, each sub-bank to include a plurality of sub-arrays coupled with a plurality of compute circuitry, individual compute circuitry from among the plurality of compute circuitry arranged to access data maintained in a corresponding sub-array from among the plurality of sub-arrays to enable the individual compute circuitry to execute a multiply and accumulate (MAC) operation, the plurality of sub-arrays to each include:

a first precharge circuitry to precharge bitlines coupled with all bit cell rows; and

a second precharge circuitry to precharge bitlines coupled with a portion of the bit cell rows,

wherein, responsive to a first signal received from a corresponding individual compute circuitry, the corresponding sub-array causes switches separately included in the corresponding sub-array to open respective bitline circuits to isolate the bitlines coupled with all bit cell rows from the bitlines coupled with the portion of bit cell rows, the portion of bit cell rows to store look up table (LUT) entries that include multiplication results and to store a configuration block that includes instructions for the corresponding individual compute circuitry to execute a MAC operation using the multiplication results in the LUT entries.

13. The system of claim 12 , comprising the corresponding sub-array to activate the second precharge circuitry following the opening of the respective bitline circuits to provide access to selected bit cells included in the portion of bit cell rows responsive to a read or a write request from the corresponding individual compute circuitry.

14. The system of claim 12 , wherein the corresponding sub-array, responsive to a second signal received from the corresponding individual compute circuitry, causes the switches to close the respective bitline circuits, the corresponding sub-array to activate the first precharge circuitry to provide access to selected bit cells included in all bit cell rows responsive to a read or a write request from the corresponding individual compute circuitry.

15. The system of claim 12 , comprising the corresponding individual compute circuitry to send the first signal responsive to an enable signal received from the slice controller.

16. The system of claim 15 , the slice controller to send the enable signal to the plurality of compute circuitry responsive to the SRAM cache for the processor being switched to an accelerator mode to execute a processor-in-memory operation associated with a deep neural network workload.

17. The system of claim 12 , the LUT entries that include multiplication results comprises multiplication results based on a product of multiplying 7 odd valued, 4-bit operands with the same 7 odd valued, 4-bit operands, the 7 odd valued, 4-bit operands to include odd values of 3, 5, 7, 9, 11, 13 and 15, wherein the LUT entries include 49, 8-bit values.

18. The system of claim 17 , the configuration block that includes instructions for the corresponding compute circuitry to execute the MAC operation comprises the instructions to include an indication to execute a matrix multiplication MAC operation with 4-bit operands, a start address for the 4-bit operands, and an end address for the 4-bit operands, the start and the end addresses located in bit cell rows of the corresponding sub-array that are separate from the portion of bit cell rows to store the LUT entries.

19. The system of claim 12 , the processor comprises a central processing unit (CPU) having a plurality of cores, the plurality of slices separately assigned for direct access by a corresponding core from among the plurality of cores.

20. The system of claim 12 , further comprising:

a display communicatively coupled to the processor;

a network interface communicatively coupled to the processor; or

a battery to power the processor and the SRAM cache.

21. A method, comprising:

receiving, at a compute circuitry coupled with a static random access memory (SRAM) sub-array, an enable signal from a controller of a memory device;

sending a first signal to the SRAM sub-array to cause first bitlines coupled to a first portion of bit cell rows of the sub-array to be isolated from second bitlines coupled to a second portion of bit cell rows of the sub-array, the first portion of bit cell rows to store look up table (LUT) entries that include multiplication results and to store a configuration block that includes instructions for the compute circuitry to execute a multiply and accumulate (MAC) operation using the multiplication results in the LUT entries;

causing a first precharge circuitry that is dedicated to the first bitlines to precharge the first bitlines to enable access to selected bit cells included in the first portion of bit cell rows; and

sending a read or write request to access the selected bit cells in the first portion of bit cell rows.

22. The method of claim 21 , the first signal to cause first bitlines coupled to a first portion of bit cell rows of the SRAM sub-array to be isolated from second bitlines coupled to a second portion of bit cell rows of the SRAM sub-array comprises the first signal to cause switches included in the SRAM sub-array to open respective bitline circuits to isolate the first bitlines coupled with the first portion of bit cell rows from the second bitlines coupled with the second portion of bit cell rows.

23. The method of claim 22 , further comprising:

sending a second signal to the SRAM sub-array to cause the switches to close the respective bitline circuits,

causing a second precharge circuitry to precharge bitlines coupled to the first and second bit cell rows to enable access to selected bit cells included in the first and second bit cell rows; and

sending a read or a write request to access the selected bit cells included in the first and second bit cell rows.

24. The method of claim 21 , comprising the SRAM sub-array and the compute circuitry included in a sub-bank of multiple sub-banks, the multiple sub-banks included in a bank of multiple banks, the multiple banks included in a slice of multiple slices of a cache for a processor, the cache to include a cache controller to configure a slice controller to control the slice, the slice controller to cause the configuration block to be stored to the first portion of bit cell rows coupled with the first precharge circuitry dedicated to the first portion of bit cell rows, wherein the compute circuitry sends the first signal responsive to receiving an enable signal from the slice controller.

25. The method of claim 24 , the slice controller to send the enable signal to the compute circuitry responsive to the cache for the processor being switched to an accelerator mode to execute a processor-in-memory operation associated with a deep neural network workload.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2021
From: JAIN, SAURABH; RANGACHAR SRINIVASA, SRIVATSA; RAMANATHAN, AKSHAY KRISHNA; KALSI, GURPREET SINGH; PILLAI, KAMLESH R.; SUBRAMONEY, SREENIVAS
To: INTEL CORPORATION
Reel/Frame 056466/0836 →
Continuity (1)
Related Publication 20220391128A1 · Dec 8, 2022
References Cited (16)
US 20190042160A1 · Kumar · 2019 [cited by examiner]
US 20210111722A1 · Kalsi · 2021 [cited by examiner]
“Micro 2020: Main Program,” 53rd IEEE/ACM International Symposium on Microarchitecture, Oct. 17-21, 2020, Available online: https://www.microarch.org/micro53/program/ (21 pages). [cited by applicant]
A. K. Ramanathan et al., “Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration,” 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (Micro), Athens, Greece… [cited by applicant]
Ali Shafiee et al., “Isaac: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,” In Proceedings of the 43rd International Symposium on Computer Architecture, ISCA '16, pp. 14-26, Pisc… [cited by applicant]
Charles Eckert et al., “Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks,” In Proceedings of the 45th Annual International Symposium on Computer Architecture, ISCA '18, pp. 383-396, Piscataway, NJ,… [cited by applicant]
Daichi Fujiki et al., “Duality Cache for Data Parallel Acceleration,” In Proceedings of the 46th International Symposium on Computer Architecture, ISCA '19, pp. 397-410, New York, NY, USA, Jun. 22-26, 2019. ACM (14 page… [cited by applicant]
J. Wang et al., “A 28-nm Compute SRAM With Bit-Serial Logic/Arithmetic Operations for Programmable In-Memory Vector Computing,” IEEE Journal of Solid-Slate Circuits, 55(1):76-86, Jan. 2020 (11 pages). [cited by applicant]
Jingyang Zhang, “Exploring Bit-Slice Sparsity in Deep Neural Networks for Efficient ReRAM-Based Deployment,” arXiv:1909.08496v2 [cs.LG], Nov. 19, 2019 (5 pages). [cited by applicant]
P. Hung et al., “Fast Division Algorithm with a Small Lookup Table,” In Conference Record of the Thirty-Third Asilomar Conference on Signals, Systems, and Computers (Cal. No. CH37020), vol. 2, pp. 1465-1468 vol. 2, Oct.… [cited by applicant]
Ping Chi et al., “PRIME: A Novel Processing-in-memory Architecture for Neural Network Computation in ReRam-based Main Memory,” In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), pp. 27… [cited by applicant]
Shaizeen Aga et al., “Compute Caches,” In 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA), pp. 481-492, Feb. 2017 (12 pages). [cited by applicant]
Song Han et al., “EIE: Efficient Inference Engine on Compressed Deep Neural Network,” In Proceedings of the 43rd International Symposium on Computer Architecture, ISCA '16, p. 243-254. IEEE Press, Jun. 2016 (12 pages). [cited by applicant]
Supreet Jeloka et al., “A 28 nm Configurable Memory (TCAM/BCAM/SRAM) Using Push-Rule 6T Bit Cell Enabling Logic-in-Memory,” IEEE Journal of Solid-Slate Circuits, 51(4):1009-1021, Apr. 2016 (13 pages). [cited by applicant]
Y. Bengio and J. Senecal, “Adaptive Importance Sampling to Accelerate Training of a Neural Probabilistic Language Model,” IEEE Transactions on Neural Networks, 19(4):713-722, Apr. 2008 (10 pages). [cited by applicant]
Yoshua Bengio, Nicholas Leonard, and Aaron Courville, “Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation,” arXiv:1308.3432v1 [cs.LG], Aug. 15, 2013 (12 pages). [cited by applicant]