IP Library › Granted Patent US 12,461,712
Granted Patent B2
US 12,461,712 · App. 17/285,409 · Granted Nov 4, 2025

In-memory near-data approximate acceleration

Inventors: Nam Sung Kim (Champaign, IL); Hadi Esmaeilzadeh (Atlanta, GA); Amir Yazdanbakhsh (Atlanta, GA)
Assignees: The Board of Trustees of the University of Illinois; Georgia Tech Research Corporation
G06F7/5443G06F1/03G06F7/527G06N3/04G11C7/1012G06F17/17
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,461,712
App. No.
17/285,409
Granted
Nov 4, 2025
Kind
B2
Abstract

A random access memory may include memory banks and arithmetic approximation units. Each arithmetic approximation unit may be dedicated to one or more of the memory banks and include a respective multiply-and-accumulate unit and a respective lookup-table unit. The respective multiply-and-accumulate unit is configured to iteratively perform shift and add operations with two inputs and to provide a result of the shift and add operations to the respective lookup-table unit. The result approximates or is a product of the two inputs. The respective lookup-table unit is configured produce an output by applying a pre-defined function to the result. The arithmetic approximation units are configured for parallel operation. The random access memory may also include a memory controller configured to receive instructions, from a processor, regarding locations within the memory banks from which to obtain the two inputs and in which to write the output.

Claims (46)

1 . A random access memory comprising:

a plurality of memory banks, wherein the plurality of memory banks comprise four bank-groups each with four of the memory banks;

a plurality of arithmetic units, each associated with one or more of the memory banks, wherein the arithmetic units are configured to perform operations on respective inputs and to provide respective results of the operations, and wherein the arithmetic units are configured for parallel operation, wherein each of the memory banks comprises two half-banks, and wherein a group of four arithmetic units from the plurality of arithmetic units is dedicated to two pairs of the two half-banks; and

a memory controller configured to receive instructions, from a processor, regarding locations within the memory banks from which to obtain the respective inputs or in which to write respective outputs based on the respective results.

2 . The random access memory of claim 1 , wherein a plurality of lookup tables are associated pairwise with the plurality of arithmetic units and are each configured to perform a pre-defined function on the respective results to produce the respective outputs.

3 . The random access memory of claim 2 , wherein the pre-defined function maps each of the respective results to a range between 0 and 1 inclusive or −1 and 1 inclusive, wherein the pre-defined function is monotonically increasing, or wherein the pre-defined function is a sigmoid function.

4 . The random access memory of claim 2 , wherein respective registers are configured to store the respective results from the arithmetic units and to provide the respective results to the associated lookup tables.

5 . The random access memory of claim 1 , wherein 128-bit datalines are shared between one of the pairs of the two half-banks.

6 . The random access memory of claim 5 , wherein the respective inputs for each of the four arithmetic units is configured to be 32 bits read from the 128-bit datalines, and wherein the respective outputs for each group of four of the arithmetic units is configured to be 32 bits written to the 128-bit datalines.

7 . The random access memory of claim 1 , wherein the respective inputs each include two or more inputs, the random access memory further comprising:

a weight register configured to store one of the two or more inputs for each of the four arithmetic units.

8 . The random access memory of claim 1 , wherein the respective inputs each include two or more inputs, wherein a first input of the two or more inputs is represented as a series of shift amounts based on a jth leading one of the first input, and wherein a second input of the two or more inputs is a multiplicand, and wherein performing the operations comprises:

initializing a cumulative sum register to be zero; and

for each respective shift amount of the shift amounts, (i) bitwise shifting left the multiplicand by the respective shift amount, and (ii) adding the multiplicand as shifted to the cumulative sum register.

9 . The random access memory of claim 1 , wherein the respective inputs each include two or more inputs, wherein the instructions include a first instruction that causes the memory controller to read a first memory region containing configuration data for the arithmetic units, wherein the configuration data includes a first input of the two or more inputs, wherein the instructions include a second instruction that causes the memory controller to read a second memory region containing operands for the arithmetic units as configured by the configuration data, wherein the operands include a second input of the two or more inputs, and wherein the instructions include a third instruction that causes the memory controller to write the respective outputs to a third memory region.

10 . The random access memory of claim 9 , wherein the configuration data defines a topology and weights of a neural network, and wherein the operands define input for the neural network.

11 . A method comprising:

receiving, by an arithmetic unit and from a memory bank of a plurality of memory banks, one or more inputs, wherein the arithmetic unit and the memory bank are both disposed within a random access memory configured for parallel operation of multiple arithmetic units, wherein the plurality of memory banks comprise four bank-groups each with four of the memory banks, wherein each of the memory banks comprises two half-banks, and wherein a group of four arithmetic units from a plurality of arithmetic units is dedicated to two pairs of the two half-banks;

performing, by the arithmetic unit, operations on the one or more inputs;

providing, by the arithmetic unit, a result of the operations; and

storing, in the memory bank, an output based on the result.

12 . The method of claim 11 , wherein the one or more inputs include two inputs, wherein a first input of the two inputs is a represented as a series of shift amounts based on a jth leading one of the first input, and wherein a second input of the two inputs is a multiplicand, and wherein performing the operations comprises:

initializing a cumulative sum register to be zero; and

for each respective shift amount of the shift amounts, (i) bitwise shifting left the multiplicand by the respective shift amount, and (ii) adding the multiplicand as shifted to the cumulative sum register.

13 . The method of claim 11 , wherein the one or more inputs include two inputs, wherein receiving the one or more inputs comprises:

receiving, from a processor, a first instruction and a second instruction, wherein the first instruction causes a memory controller disposed within the random access memory to read a first memory region containing configuration data for the arithmetic unit, wherein the configuration data includes a first input of the two inputs, wherein the second instruction causes the memory controller to read a second memory region containing operands for the arithmetic unit as configured by the configuration data, wherein the operands include a second input of the two inputs.

14 . The method of claim 13 , wherein the configuration data defines a topology and weights of a neural network, and wherein the operands define input for the neural network.

15 . A system comprising:

a processor; and

a random access memory including: (i) a memory bank of a plurality of memory banks, wherein the plurality of memory banks comprise four bank-groups each with four of the memory banks, and (ii) a plurality of arithmetic units, each associated with one or more of the memory banks, wherein the arithmetic units are configured to perform operations on respective inputs and to provide respective results of the operations, wherein each of the memory banks comprises two half-banks, wherein a group of four arithmetic units from the plurality of arithmetic units is dedicated to two pairs of the two half-banks, and wherein the arithmetic units are configured for parallel operation; and (iii) a memory controller configured to receive instructions, from the processor, regarding locations within the memory bank from which to obtain the respective inputs and in which to write respective outputs based on the respective results.

16 . The system of claim 15 , wherein the respective inputs each include two or more inputs, wherein a first input of the two or more inputs is a represented as a series of shift amounts based on a jth leading one of the first input, and wherein a second input of the two or more inputs is a multiplicand, and wherein performing the operations comprises:

initializing a cumulative sum register to be zero; and

for each respective shift amount of the shift amounts, (i) bitwise shifting left the multiplicand by the respective shift amount, and (ii) adding the multiplicand as shifted to the cumulative sum register.

17 . The system of claim 15 , wherein the respective inputs each include two or more inputs, wherein the instructions include a first instruction that causes the memory controller to read a first memory region containing configuration data for the arithmetic units, wherein the configuration data includes a first input of the two or more inputs, wherein the instructions include a second instruction that causes the memory controller to read a second memory region containing operands for the arithmetic units as configured by the configuration data, wherein the operands include a second input of the two or more inputs, and wherein the instructions include a third instruction that causes the memory controller to write the respective outputs to a third memory region.

18 . The system of claim 15 , wherein the respective inputs each include two or more inputs, the system further comprising:

a weight register configured to store one of the two or more inputs for two or more of the arithmetic units.

19 . A random access memory comprising:

a plurality of memory banks;

a plurality of arithmetic units, each associated with one or more of the memory banks, wherein the arithmetic units are configured to perform operations on respective inputs and to provide respective results of the operations, and wherein the arithmetic units are configured for parallel operation; and

a memory controller configured to receive instructions, from a processor, regarding locations within the memory banks from which to obtain the respective inputs or in which to write respective outputs based on the respective results, wherein the respective inputs each include two or more inputs, wherein a first input of the two or more inputs is represented as a series of shift amounts based on a jth leading one of the first input, and wherein a second input of the two or more inputs is a multiplicand, and wherein performing the operations comprises:

initializing a cumulative sum register to be zero; and

for each respective shift amount of the shift amounts, (i) bitwise shifting left the multiplicand by the respective shift amount, and (ii) adding the multiplicand as shifted to the cumulative sum register.

20 . A random access memory comprising:

a plurality of memory banks;

a plurality of arithmetic units, each associated with one or more of the memory banks, wherein the arithmetic units are configured to perform operations on respective inputs and to provide respective results of the operations, and wherein the arithmetic units are configured for parallel operation; and

a memory controller configured to receive instructions, from a processor, regarding locations within the memory banks from which to obtain the respective inputs or in which to write respective outputs based on the respective results, wherein the respective inputs each include two or more inputs, wherein the instructions include a first instruction that causes the memory controller to read a first memory region containing configuration data for the arithmetic units, wherein the configuration data includes a first input of the two or more inputs, wherein the instructions include a second instruction that causes the memory controller to read a second memory region containing operands for the arithmetic units as configured by the configuration data, wherein the operands include a second input of the two or more inputs, and wherein the instructions include a third instruction that causes the memory controller to write the respective outputs to a third memory region.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2022
From: ESMAEILZADEH, HADI; YAZDANBAKHSH, AMIR
To: GEORGIA TECH RESEARCH CORPORATION
Reel/Frame 061027/0333 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2022
From: KIM, NAM SUNG
To: THE BOARD OF TRUSTEES OF THE UNIVERSITY OF ILLINOIS
Reel/Frame 061028/0014 →
Continuity (2)
Provisional Application 62745688 · Oct 15, 2018
Related Publication 20210382691A1 · Dec 9, 2021
References Cited (124)
US 5530661A · Garbe · 1996 [cited by examiner]
US 6724684B2 · Kim · 2004 [cited by applicant]
US 7623134B1 · Danilak · 2009 [cited by applicant]
US 9710265B1 · Temam et al. · 2017 [cited by applicant]
US 20080028181A1 · Tong et al. · 2008 [cited by applicant]
US 20190303749A1 · Appuswamy · 2019 [cited by examiner]
US 20200201633A1 · Madduri · 2020 [cited by examiner]
Mohsen Imani et al., “Ultra-Efficient processing In-Memory for Data Intensive Applications”, Jun. 18, 2017, pp. 1-6. [cited by applicant]
Gao Di et al., “A Design Framework for Processing-In-Memory Accelerator”, 2018 ACM/IEEE International Workshop on System Level Interconnect Prediction (SLIP), ACM, Jun. 23, 2018, pp. 1-6. [cited by applicant]
Jeon Dong-Ik et al., “HMC-MAC: Processing-in Memory Architecture for Multiply-Accumulate Operations with Hybrid memory Cube”, IEEE Computer Architecture Letters, IEEE, US, vol. 17, No. 1, Jan. 1, 2018, pp. 5-8. [cited by applicant]
International Search Report and Written Opinion dated Jan. 22, 2020 for International Application No. PCT/US2019/056068, 15 pages. [cited by applicant]
X. Chen, et al., “Adaptive Cache Management for Energy-Efficient GPU Computing,” in MICRO, 2014. [cited by applicant]
Y. Tian, et al., “Adaptive GPU Cache Bypassing,” in Proceedings of the 8th Workshop on General Purpose Processing using GPUs, 2015. [cited by applicant]
M. Carbin, et al., “Verifying Quantitative Reliability for Programs that Execute on Unreliable Hardware,” in OOPSLA, 2013. [cited by applicant]
J. Park, et al., “FlexJava: Language Support for Safe and Modular Approximate Programming,” in FSE, 2015. [cited by applicant]
I. Singh, et al., “Cache Coherence for GPU Architectures,” in HPCA, 2013. [cited by applicant]
Hynix. Hynix GDDR5 SGRAM Part H5GQ1H24AFR Revision 1.0., 2009. [cited by applicant]
“Nvidia Corporation. CUDA Programming Guide.” http://docs.nvidia.com/cuda/cuda-c-programmingguide, 2015. [cited by applicant]
Jedec, “High Bandwidth Memory DRAM.” http://www.jedec.org/standardsdocuments/docs/jesd235, Oct. 2013. [cited by applicant]
S. W. Keckler, et al., “GPUs and the Future of Parallel Computing,” IEEE Micro, vol. 31, No. 5, pp. 7-17, 2011. [cited by applicant]
J. Mukundan, et al., “Understanding and Mitigating Refresh Overheads in High-density DDR4 DRAM Systems,” in ISCA, 2013. [cited by applicant]
Q. Guo, et al., “AC-DIMM: Associative Computing with STT-MRAM,” in ISCA, 2013. [cited by applicant]
S. M. Hassan, et al., “Near Data Processing: Impact and Optimization of 3D Memory System Architecture on the Uncore,” in MEMSYS, 2015. [cited by applicant]
K. Hsieh, et al., “Transparent Offloading and Mapping (TOM): Enabling Programmer-Transparent Near-Data Processing in GPU Systems,” in ISCA, 2016. [cited by applicant]
B. Akin, et al., “Data Reorganization in Memory using 3D-stacked DRAM,” in ISCA, 2015. [cited by applicant]
Y.-H. Chen, et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” in ISCA, 2016. [cited by applicant]
M. Gao, et al., “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,” 2017. [cited by applicant]
N. Chatterjee, et al., “Managing DRAM Latency Divergence in Irregular GPGPU Applications,” in SC, 2014. [cited by applicant]
P. Boudier and G. Sellers, “Memory System on Fusion APUs,” AMD Fusion developer summit, 2011. [cited by applicant]
S. Kato, et al., “Gdev: First-Class GPU Resource Management in the Operating System,” in USENIX, 2012. [cited by applicant]
Y. Fujii, et al., “Data Transfer Matters for GPU Computing,” in ICPADS, 2013. [cited by applicant]
J. Power, et al., “Supporting x86-64 Address Translation for 100s of GPU Lanes,” in HPCA, 2014. [cited by applicant]
H. Wong, et al., “Demystifying GPU Microarchitecture through Microbenchmarking,” in ISPASS, 2010. [cited by applicant]
X. Mei and X. Chu, “Dissecting GPU Memory Hierarchy through Microbenchmarking,” IEEE Transactions on Parallel and Distributed Systems, No. 99, 2016. [cited by applicant]
B. Pichai, et al., “Architectural Support for Address Translation on GPUs: Designing Memory Management Units for CPU/GPUs with Unified Address Spaces,” in ASPLOS, 2014. [cited by applicant]
N. Wilt, The CUDA Handbook: A Comprehensive Guide to GPU Programming. Pearson Education, 2013. [cited by applicant]
T. Beri, et al., “A Scheduling and Runtime Framework for a Cluster of Heterogeneous Machines with Multiple Accelerators,” in IPDPS, 2015. [cited by applicant]
M. Harris, “Inside Pascal: Nvidia's Newest Computing Platform.”https://devblogs.nvidia.com/parallelforall/inside-pascal/, 2016. [cited by applicant]
A. Yazdanbakhsh, et al., “AxBench: A Multi-Platform Benchmark Suite for Approximate Computing,” IEEE Design and Test, 2016. [cited by applicant]
“FreePDK45 and the Nangate Open Cell Library,” https://mflowgen.readthedocs.io/en/latest/stdlib-freepdk45.html, 2020. [cited by applicant]
T. Johnson and D. Shasha, “2Q: A Low Overhead High Performance Buffer Management Replacement Algorithm,” in VLDB, 1994. [cited by applicant]
S. Rixner, “Memory Controller Optimizations for Web Servers,” in MICRO, 2004. [cited by applicant]
S. Rixner, et al., “Memory Access Scheduling,” ISCA, 2000. [cited by applicant]
T. Vogelsang, “Understanding the Energy Consumption of Dynamic Random Access Memories,” in MICRO, 2010. [cited by applicant]
M. O'Connor, “Highlights of the High-Bandwidth Memory (HBM) Standards,” in The Memory Forum co-located with ISCA, 2014. [cited by applicant]
S. Galal, Energy Efficient Floating-Point Unit Design. PhD thesis, The Department of Electrical Engineering of Stanford University, 2012. [cited by applicant]
J. Leng, et al., “GPUWattch: Enabling Energy Optimizations in GPGPUs,” in ISCA, 2013. [cited by applicant]
S. Li, et al., “McPAT: An Integrated Power, Area, and Timing Modeling Framework for Multicore and Manycore Architectures,” in MICRO, 2009. [cited by applicant]
N. Muralimanohar, et al., “Optimizing NUCA Organizations and Wiring Alternatives for Large Caches with CACTI 6.0,” in MICRO, 2007. [cited by applicant]
Y. H. Son, et al., “Reducing Memory Access Latency with Asymmetric DRAM Bank Organizations,” in ISCA, 2013. [cited by applicant]
K. K. Chang, et al., “Low-Cost Inter-Linked Subarrays (LISA): Enabling fast inter-subarray data movement in DRAM,” in HPCA, 2016. [cited by applicant]
Z. Du, et al., “Leveraging the Error Resilience of Machine-Learning Applications for Designing Highly Energy Efficient Accelerators,” in ASP-DAC, 2014. [cited by applicant]
B. Belhadj, et al., “Continuous Real-World Inputs Can Open Up Alternative Accelerator Designs,” in ISCA, 2013. [cited by applicant]
B. Grigorian and G. Reinman, “Accelerating Divergent Applications on SIMD Architectures using Neural Networks,” in ICCD, 2014. [cited by applicant]
S. Eldridge, et al., “Towards General-Purpose Neural Network Computing,” in PACT, 2015. [cited by applicant]
M. F. Deering, et al., “FBRAM: A New Form of Memory Optimized for 3D Graphics,” in SIGGRAPH, 1994. [cited by applicant]
J. Draper, et al., “The Architecture of the DIVA Processing-in-Memory Chip,” in Supercomputing, 2002. [cited by applicant]
D. Patterson, et al., “A Case for Intelligent RAM,” Micro, IEEE, vol. 17, No. 2, 1997. [cited by applicant]
D. G. Elliott, et al., “Computational RAM: A Memory-SIMD Hybrid and its Application to DSP,” in Custom Integrated Circuits Conference, vol. 30, 1992. [cited by applicant]
Y. Kang, et al., “FlexRAM: Toward an Advanced Intelligent Memory System,” in ICCD, 2012. [cited by applicant]
K. Mai, et al., “Smart Memories: A Modular Reconfigurable Architecture,” in ISCA, 2000. [cited by applicant]
M. Oskin, et al., “Active Pages: a Computation Model for Intelligent Memory,” in ISCA, 1998. [cited by applicant]
L. Nai, et al., “GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks,” in HPCA, 2017. [cited by applicant]
D. Zhang, et al., “TOP-PIM: Throughput-Oriented Programmable Processing in Memory,” in HPDC, 2014. [cited by applicant]
J. Ahn, et al., “A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing,” in ISCA, 2015. [cited by applicant]
S. Pugsley, et al., “Comparing Implementations of Near-Data Computing with In-Memory MapReduce Workloads,” Micro, IEEE, vol. 34, No. 4, 2014. [cited by applicant]
R. Nair, et al., “Active Memory Cube: A Processing-in-Memory Architecture for Exascale Systems,” IBM Journal of Research and Development, vol. 59, No. 2/3, 2015. [cited by applicant]
Q. Guo, et al., “3D-Stacked Memory-Side Acceleration: Accelerator and System Design,” in WoNDP, 2014. [cited by applicant]
C. Shelor, et al., “Dataflow based Near Data Processing using Coarse Grain Reconfigurable Logic,” in WoNDP, 2015. [cited by applicant]
Q. Zhu, et al., “Accelerating Sparse Matrix-Matrix Multiplication with 3D-Stacked Logic-in-Memory Hardware,” in HPEC, 2013. [cited by applicant]
M. Gao and C. Kozyrakis, “HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing,” in HPCA, 2016. [cited by applicant]
A. Farmahini-Farahani, et al., “DRAMA: An Architecture for Accelerated Processing Near Memory,” CAL, vol. 14, No. 1, 2015. [cited by applicant]
D. Kim, et al., “NeuroCube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory,” in ISCA, 2016. [cited by applicant]
T. P. Morgan, “Accelerating Compute by Cramming it into DRAM Memory,” TheNextPlatform, Oct. 3, 2019. [cited by applicant]
B. M. Rogers, et. al., “Scaling the Bandwidth Wall: Challenges in and Avenues for CMP Scaling,” in ISCA, 2009. [cited by applicant]
A. Yazdanbakhsh, et. al., “RFVP: Rollback-Free Value Prediction with Safe to Approximate Loads,” in TACO, 2015. [cited by applicant]
N. Vijaykumar, et. al., “A Case for Core-Assisted Bottleneck Acceleration in GPUs: Enabling Efficient Data Compression,” in ISCA, 2015. [cited by applicant]
B. Casper, “Energy Efficient Multi-GB/s I/O: Circuit and System Design Techniques,” in IEEE Workshop on Microelectronics and Electron Devices, 2011. [cited by applicant]
M. Horowitz, “CS 598sm Probabilistic & Approximate Computing,” http://misailo.web.engr.Illinois.edu/courses/cs598. [cited by applicant]
S. Han, et. al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization, and Huffman Coding,” in ICLR, 2016. [cited by applicant]
S. Han, et. al., “Learning both Weights and Connections for Efficient Neural Network,” in NIPS, 2015. [cited by applicant]
A. Farmahini-Farahani, et. al., “NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,” in HPCA, 2015. [cited by applicant]
R. Sampson, et. al., “Sonic Millip3De: A Massively Parallel 3D-stacked Accelerator for 3D Ultrasound,” in HPCA, 2013. [cited by applicant]
R. Hou, et. al., “Efficient Data Streaming with On-chip Accelerators: Opportunities and Challenges,” in HPCA, 2011. [cited by applicant]
H. Asghari-Moghaddam, et. al., “Chameleon: Versatile and Practical Near-DRAM Acceleration Architecture for Large Memory Systems,” in MICRO, 2016. [cited by applicant]
D. U. Lee, et. al., “25.2 A 1.2V 8Gb 8-Channel 128GB/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I/O Test Methods using 29nm Process and TSV,” in ISSCC, 2014. [cited by applicant]
S. Liu, et. al., “Hardware/Software Techniques for DRAM Thermal Management,” in HPCA, 2011. [cited by applicant]
J. Iyer, et. al., “System Memory Power and Thermal Management in Platforms Build on Intel Centrino Duo Technology,” Intel Technology Journal, vol. 10, No. 2, 2006. [cited by applicant]
D. Kanter, A Preview of Intel's Bensley Platform (Part I), Real World Tech, Nov. 8, 2005. [cited by applicant]
J. Lin, et. al., “Thermal Modeling and Management of DRAM Memory Systems,” in ISCA, 2007. [cited by applicant]
J. Lin, et. al., “Software Thermal Management of DRAM Memory for Multicore Systems,” SIGMETRICS, vol. 36, No. 1, pp. 337-348, 2008. [cited by applicant]
J. Lin, et. al., “Thermal Modeling and Management of DRAM Systems,” IEEE Transactions on Computers, vol. 62, No. 10, pp. 2069-2082, 2013. [cited by applicant]
K. Koo, et. al., “A 1.2V 38nm 2.4Gb/s/pin 2Gb DDR4 SDRAM with Bank Group and x4 Half-Page Architecture,” in ISSCC, pp. 40-41, 2012. [cited by applicant]
T. Y. Oh, et. al., “A 7 Gb/s/pin 1 Gbit GDDR5 SDRAM With 2.5 ns Bank to Bank Active Time and No Bank Group Restriction,” JSSC, vol. 46, No. 1, pp. 107-118, 2011. [cited by applicant]
T. Y. Oh, et. al., “A 7Gb/s/pin GDDR5 SDRAM with 2.5ns Bank-to-Bank Active Time and no Bank-group Restriction,” in ISSCC'10. [cited by applicant]
H. Esmaeilzadeh, et. al. , “Neural Acceleration for General-Purpose Approximate Programs,” in MICRO, 2012. [cited by applicant]
R. S. Amant, et. al., “General-Purpose Code Acceleration with Limited-Precision Analog Computation,” in ISCA, 2014. [cited by applicant]
A. Yazdanbakhsh, et. al., “Neural Acceleration for GPU Throughput Processors,” in MICRO, 2015. [cited by applicant]
B. Grigorian, et. al., “BRAINIAC: Bringing Reliable Accuracy Into Neurally-Implemented Approximate Computing,” in HPCA, 2015. [cited by applicant]
T. Moreau, et. al., “SNNAP: Approximate Computing on Programmable SoCs via Neural Acceleration,” in HPCA, 2015. [cited by applicant]
V. Govindaraju, et. al., “DySER: Unifying Functionality and Parallelism Specialization for Energy-Efficient Computing,” EEE Micro, vol. 32, No. 5, pp. 38-51, 2012. [cited by applicant]
W. Baek and T. M. Chilimbi, “Green: A Framework for Supporting Energy-Conscious Programming using Controlled Approximation,” in PLDI, 2010. [cited by applicant]
M. Samadi, et. al., “SAGE: Self-Tuning Approximation for Graphics Engines,” in MICRO, 2013. [cited by applicant]
M. Samadi, et. al., “Paraprox: Pattern-based Approximation for Data Parallel Applications,” in ASPLOS, 2014. [cited by applicant]
S. Sidiroglou-Douskos, et. al., “Managing Performance vs. Accuracy Trade-offs with Loop Perforation,” in FSE, 2011. [cited by applicant]
S. Misailovic, et. al., “Quality of Service Profiling,” in ICSE, 2010. [cited by applicant]
M. Rinard, et. al., “Patterns and Statistical Analysis for Understanding Reduced Resource Computing,” in Onward!, 2010. [cited by applicant]
L. McAfee and K. Olukotun, “EMEURO: A Framework for Generating Multi-Purpose Accelerators via Deep Learning,” in CGO, 2015. [cited by applicant]
B. Grigorian and G. Reinman, “Accelerating Divergent Applications on SIMD Architectures Using Neural Networks,” ACM Trans. Archit. Code Optim., vol. 12, pp. 2:1-2:23, Mar. 2015. [cited by applicant]
H. Esmaeilzadeh, et. al., “Architecture Support for Disciplined Approximate Programming,” in ASPLOS, 2012. [cited by applicant]
L. N. Chakrapani, et. al., “Ultra-efficient (Embedded) SOC Architectures based on Probabilistic CMOS (PCMOS) Technology,” in Date, 2006. [cited by applicant]
L. Leem, et. al., “ERSA: Error Resilient System Architecture for Probabilistic Applications,” in Date, 2010. [cited by applicant]
J. Sartori and R. Kumar, “Branch and Data Herding: Reducing Control and Memory Divergence for Error-Tolerant GPU Applications,” Multimedia, IEEE Transactions on, vol. 15, No. 2, 2013. [cited by applicant]
A. Sampson, et. al., “EnerJ: Approximate Data Types for Safe and General Low-Power Computation,” in PLDI, 2011. [cited by applicant]
A. Yazdanbakhsh, et. al., “Axilog: Language Support for Approximate Hardware Design,” in Date, 2015. [cited by applicant]
A. Ranjan, et. al., “ASLAN: Synthesis of Approximate Sequential Circuits,” in Date, 2014. [cited by applicant]
S. Venkataramani, et. al., “SALSA: Systematic Logic Synthesis of Approximate Circuits,” in DAC, 2012. [cited by applicant]
K. Nepal, et. al., “ABACUS: A Technique for Automated Behavioral Synthesis of Approximate Computing Circuits,” in Data, 2014. [cited by applicant]
A. Lingamneni, et. al., “Synthesizing Parsimonious Inexact Circuits Through Probabilistic Design Techniques,” ACM Trans. Embed. Comput. Syst., vol. 12, No. 2s, 2013. [cited by applicant]
A. Lingamneni, et. al., “Algorithmic Methodologies for Ultra-efficient Inexact Architectures for Sustaining Technology Scaling,” in CF, 2012. [cited by applicant]
T. G. Rogers, et. al., “Cache-Conscious Wavefront Scheduling,” in MICRO, 2012. [cited by applicant]
A. Bakhoda, et. al., “Analyzing CUDA Workloads using a Detailed GPU Simulator,” in ISPASS, 2009. [cited by applicant]
B. Pichai, et. a., “Architectural Support for Address Translation on GPUs: Designing Memory Management Units for CPU/GPUs with Unified Address Spaces,” in ACM SIGARCH Computer Architecture News, vol. 42, pp. 743-757, 20… [cited by applicant]
H. Jooybar, et. al., “GPUDet: A Deterministic GPU Architecture,” in ASPLOS, 2013. [cited by applicant]