IP Library Granted Patent US 12,292,838
Granted Patent B2
US 12,292,838 · App. 18/066,161 · Granted May 6, 2025

Host device performing near data processing function and accelerator system including the same

Inventors: Hyungkyu Ham (Pohang, KR); Hyunuk Cho (Incheon, KR); Hyojin Sung (Pohang, KR); Eunhyeok Park (Pohang, KR); Gwangsun Kim (Pohang, KR)
Assignees: SK Hynix Inc.; POSTECH ACADEMY-INDUSTRY FOUNDATION
G06F12/1081G06F12/125
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,292,838
App. No.
18/066,161
Granted
May 6, 2025
Kind
B2
Abstract

A host device includes a unit processor configured to generate a near data processing (NDP) request, a host expansion control circuit configured to receive the NDP request; and a local memory device configured to store data corresponding to the NDP request according to control by the expansion control circuit. In response to receiving the NDP request, the host expansion control circuit performs a request processing operation to perform a read or a write operation corresponding to the NDP request on the local memory device and performs a computation operation using the requested data corresponding to the NDP request.

Claims (65)

1. A host device comprising:

a unit processor configured to generate a near data processing (NDP) request;

a host expansion control circuit configured to receive the NDP request; and

a local memory device configured to store data corresponding to the NDP request according to control by the host expansion control circuit,

wherein in response to the NDP request, the host expansion control circuit performs:

a request processing operation to perform a memory operation corresponding to the NDP request on the local memory device, the memory operation including a read operation or a write operation, and

a computation operation using the data corresponding to the NDP request,

wherein the host expansion control circuit comprises:

one or more host NDP request control circuits; and

an interface circuit configured to receive the NDP request, select a host NDP request control circuit from among the one or more host NDP request control circuits according to an address of the NDP request, and provide the NDP request to the selected host NDP control circuit,

wherein the selected host NDP request control circuit is configured to control the request processing operation and the computation operation corresponding to the NDP request,

wherein the selected host NDP request control circuit comprises:

a filter circuit configured to identify the NDP request;

an NDP circuit configured to produce a request for the request processing operation and to perform the computation operation according to the NDP request identified at the filter circuit; and

a memory controller configured to control the local memory device according to the request for the request processing operation produced by the NDP circuit,

wherein the host NDP request control circuit further includes a cache memory connected between the filter circuit and the memory controller,

wherein the host expansion control circuit is further configured to receive a normal request that does not require a computation operation, and

wherein the filter circuit is further configured to identify the normal request and to bypass the identified normal request to the memory controller via the cache memory.

2. The host device of claim 1 , wherein the filter circuit stores a table including address information, and wherein the filter circuit identifies the NDP request and the normal request with reference to the address information.

3. The host device of claim 1 , wherein the NDP circuit comprise:

a computation circuit configured to perform a computation operation corresponding to the NDP request;

an instruction storage circuit configured to store an instruction for the computation operation and a request for the request processing operation; and

a register file including a plurality of registers for the computation operation.

4. The host device of claim 3 , wherein the NDP circuit further comprise an instruction cache to store a plurality of instructions, and wherein the instruction storage circuit stores an instruction received from the instruction cache corresponding to the NDP request.

5. The host device of claim 4 , wherein the NDP circuit further comprises a request decoder to perform a decoding operation by using information included in the NDP request, and

wherein the request decoder includes an NDP kernel table that associatively stores the NDP request and instruction cache address corresponding to the NDP request.

6. The host device of claim 5 , wherein the NDP circuit further comprises:

a micro-context storage circuit configured to associatively store the NDP request and a base address for one or more registers allocated for use in the computation operation; and

a register address translation circuit configured to generate register address used for the computation operation with reference to the base address.

7. The host device of claim 1 , wherein the host expansion control circuit further includes a direct memory access (DMA) circuit connected to the interface circuit and configured to generate an NDP request, and

wherein the interface circuit is configured to provide the NDP request generated by the DMA circuit to the NDP request control circuit or to a device external to the host expansion control circuit.

8. An accelerator system comprising:

a host device including a unit processor;

a memory expansion device; and

an interconnect circuit configured to connect the host device and the memory expansion device,

wherein the host device among the plurality of memory expansion devices includes:

a host expansion control circuit configured to receive a near data processing (NDP) request provided to the unit processor; and

a local memory device configured to store data corresponding to the NDP request according to control by the host expansion control circuit, and

wherein in response to the NDP request, the host expansion control circuit performs:

a request processing operation to perform a memory operation corresponding to the NDP request on the local memory device, the memory operation including a read operation or a write operation, and

a computation operation using the data corresponding to the NDP request,

wherein the host expansion control circuit comprises:

one or more host NDP request control circuits; and

an interface circuit configured to receive the NDP request, select a host NDP request control circuit from among the one or more host NDP request control circuits according to an address of the NDP request, and provide the NDP request to the selected host NDP control circuit, and

wherein the selected host NDP request control circuit configured to control the request processing operation and the computation operation corresponding to the NDP request,

wherein the selected host NDP request control circuit comprises:

a filter circuit configured to identify the NDP request;

an NDP circuit configured to produce a request for the request processing operation and to perform the computation operation according to the NDP request identified at the filter circuit; and

a memory controller configured to control the local memory device according to the request for the request processing operation produced by the NDP circuit,

wherein the host NDP request control circuit includes a cache memory connected between the filter circuit and the memory controller,

wherein the host expansion control circuit is further configured to receive a normal request that does not require a computation operation, and

wherein the filter circuit is further configured to identify the normal request and to bypass the identified normal request to the memory controller via the cache memory controller.

9. The accelerator system of claim 8 , wherein the filter circuit stores a table including address information, and wherein the filter circuit identifies the NDP request and the normal request with reference to the address information.

10. The accelerator system of claim 8 , wherein the NDP circuit comprise:

a computation circuit configured to perform a computation operation corresponding to the NDP request;

an instruction storage circuit configured to store an instruction for the computation operation and a request for the request processing operation; and

a register file including a plurality of registers for the computation operation.

11. The accelerator system of claim 10 , wherein the NDP circuit further comprise an instruction cache to store a plurality of instructions, and wherein the instruction storage circuit stores an instruction received from the instruction cache corresponding to the NDP request.

12. The accelerator system of claim 11 , wherein the NDP circuit further comprises a request decoder to perform a decoding operation by using information included in the NDP request, and

wherein the request decoder includes an NDP kernel table that associatively stores the NDP request and instruction cache address corresponding to the NDP request.

13. The accelerator system of claim 12 , wherein the NDP circuit further comprises:

a micro-context storage circuit configured to associatively store the NDP request and a base address for one or more registers allocated for use in the computation operation; and

a register address translation circuit configured to generate register address used for the computation operation with reference to the base address.

14. The accelerator system of claim 8 , wherein the host expansion control circuit further includes a direct memory access (DMA) circuit connected to the interface circuit and configured to generate an NDP request, and

wherein the interface circuit is configured to provide the NDP request generated at by DMA circuit to the NDP request control circuit or to a device external to the host expansion control circuit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: HAM, HYUNGKYU; CHO, HYUNUK; SUNG, HYOJIN; PARK, EUNHYEOK; KIM, GWANGSUN
To: SK HYNIX INC.; POSTECH ACADEMY-INDUSTRY FOUNDATION
Reel/Frame 062096/0878 →
Priority Claims (2)
KR 10-2021-0184439 · Dec 22, 2021 · national
KR 10-2022-0137080 · Oct 24, 2022 · national
Continuity (1)
Related Publication 20230195651A1 · Jun 22, 2023
References Cited (168)
US 7623134B1 · Danilak · 2009 [cited by applicant]
US 11468001B1 · Hassaan · 2022 [cited by examiner]
US 20210117131A1 · Kim et al. · 2021 [cited by applicant]
US 20210311739A1 · Malladi et al. · 2021 [cited by applicant]
US 20210349837A1 · Huangfu et al. · 2021 [cited by applicant]
US 20230026505A1 · Lee · 2023 [cited by examiner]
US 20230195459A1 · Puthoor · 2023 [cited by examiner]
KR 1020190018888A · 2019 [cited by applicant]
KR 1020200018188A · 2020 [cited by applicant]
Ham et al. “Near-Data Processing in Memory Expander for DNN Acceleration on GPUs,” IEEE Computer Architecture Letters, vol. 20, No. 2, Jul.-Dec. 2021 (Year: 2021). [cited by examiner]
D. Amodei et al., “Ai and compute”, https://openai.com/blog/ai-and-compute, 2018. [cited by applicant]
S. Iofee et al., “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning, ser. Proceedings of Machine Learn… [cited by applicant]
“Compute Express Link Specification 2.0,” CXL Consortium, 2020, https://www.computeexpresslink.org/download-the-specification, Oct. 2020. [cited by applicant]
“An Introduction to CCIX White Paper,” CCIX Consortium Inc, https://www.ccixconsortium.com/wp-content/uploads/2019/11/CCIX-White-Paper-Rev111219.pdf, 2019. [cited by applicant]
M. Krause et al., “Gen-Z DRAM and Persistent Memory Theory of Operation,” Gen-Z Consortium, 2019. [cited by applicant]
K. He et al., “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. [cited by applicant]
M. Sandler et al., “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2018. [cited by applicant]
Y. Kwon et al., “TensorDIMM: A practical near memory processing architecture for embeddings and tensor operations in deep learning,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitectur… [cited by applicant]
L. Ke et al., “RecNMPp: Accelerating personalized recommendation with near-memory processing,” in Proceedings of the ACM/IEEE 47th Annual International Symposium on Computer Architecture, ser. ISCA '20, 2020, p. 790-803. [cited by applicant]
M. He et al., “Newton: A dram-maker's accelerator-in-memory (aim) architecture for machine learning,” in 53rd Annual IEEE/ACM International Symposium on Microarchitecture, 2020. [cited by applicant]
S. Lee et al., “Hardware architecture and software stack for PIM based on commercial dram technology,” in Proc. ACM/IEEE 48th Annu. Int. Symp. Comput. Archit., 2021, pp. 43-56. [cited by applicant]
Nvidia, “Convolutional layers user guide,” Nvidia Docs, https://docs.nvidia.com/deeplearning/performance/dl-performanceconvolutional/index.html, 2021. [cited by applicant]
M. Khairy, et al., “Accel-sim: An extensible simulation framework for validated gpu modeling,” in ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), 2020, pp. 473-486. [cited by applicant]
Y. Kim et al., “Ramulator: A fast and extensible dram simulator,” IEEE Computer Architecture Letters, vol. 15, No. 1,pp. 45-49, 2016. [cited by applicant]
K. Simonyan et al., “Very deep convolutional networks for large-scale image recognition,” ICLR, 2015. [cited by applicant]
N. Muralimanohar et al., “Cacti 6.0: A tool to model large caches,” HP laboratories, vol. 27, Apr. 2009. [cited by applicant]
M. Hibben, “TSMC, not intel, has the lead in semiconductor processes,” https://seekingalpha.com/article/4151376-tsmc-not-intel-lead-in-semiconductor-processes, 2018. [cited by applicant]
S. Mach et al., “FPnew: An open-source multiformat floating-point unit architecture for energy proportional transprecision computing,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 29, No. 04, p… [cited by applicant]
C. Sun et al., “DSENT—A tool connecting emerging photonics with electronics for opto-electronic networks-on-chip modeling,” IEEE/ACM Sixth International Symposium on Networks-on-Chip, 2012. [cited by applicant]
R. Hwang et al., “Centaur: A chiplet-based, hybrid sparse-dense accelerator for personalized recommendations,” in Proc. ACM/IEEE 47th Annu. Int. Symp. Comput. Archit., 2020, pp. 968-981. [cited by applicant]
Y. Kwon et al., “Beyond the memory wall: A case for memory centric hpc system for deep learning,” in Proceedings of the 51stAnnual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO-51. IEEE Press, 2018, … [cited by applicant]
E. Choukse et al., “Buddy compression: Enabling larger memory for deep learning and hpc workloads on gpus,” ACM/IEEE47th Annual International Symposium on Computer Architecture, 2020. [cited by applicant]
“NCCL Tests”, https://github.com/NVIDIA/nccl-tests, 2016. [cited by applicant]
“XLA: Optimizing compiler for machine learning”, https://www.tensorflow.org/xla, Nov. 2020. [cited by applicant]
“Nvidia nvlink high-speed interconnect: Application performance”, Nvidia Whitepaper, 2014. [cited by applicant]
“OpenCAPI overview,” OpenCAPI Consortium, 2016. [cited by applicant]
N. Gebara et al., “In-network aggregation for shared machine learning clusters,” in Proceedings of Machine Learning and Systems, vol. 3, 2021, pp. 829-844. [cited by applicant]
“Introducing AMD CDNA architecture,” AMD whitepaper, 2020. [cited by applicant]
“Nvidia DGX A100 System Architecture,” Nvidia Technical WhitePaper, 2020. [cited by applicant]
“Nvidia data center deep learning product performance”, https://developer.nvidia.com/deep-learning-performance-training-inference, Dec. 2021. [cited by applicant]
M. Abadi et al., “TensorFlow: A system for large-scale machine learning,” in 12thUSENIX Symposium on Operating Systems Design and Implementation (OSDI 16). Savannah, GA: USENIX Association, Nov. 2016, pp. 265-283. [cited by applicant]
M. Andersch et al., “Tensor Core DL Performance Guide,” Nvidia GPU Technology Conference, 2019. [cited by applicant]
B. Asgari et al., FAFNIR: Accelerating sparse gathering by using efficient near-memory intelligent reduction, in IEEE International Symposium on High Performance Computer Architecture (HPCA), 2021, pp. 908-920. [cited by applicant]
L. J. Ba et al.,“Layer normalization,” CoRR, vol. abs/1607.06450, 2016. [cited by applicant]
T. Brown et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems, vol. 33.Curran Associates, Inc., 2020, pp. 1877-1901. [cited by applicant]
M. Caron et al., “Unsupervised learning of visual features by contrasting cluster assignments,” in Advances in Neural Information Processing Systems, vol. 33. Curran Associates, Inc., 2020, pp. 9912-9924. [cited by applicant]
T. Chen et al., “TVM: An automated end-to-end optimizing compiler for deep learning,” in13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), Oct. 2018, pp. 578-594. [cited by applicant]
T. Chen et al., “DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,” in Proceedings of the 19th International Conference on Architectural Support for Programming Languages and Operat… [cited by applicant]
T. Chen et al., “A simple framework for contrastive learning of visual representations,” in Proceedings of the 37th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, H. D. III … [cited by applicant]
T. Chen et al., “Big self-supervised models are strong semi-supervised learners,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33.Curr… [cited by applicant]
Y.-H. Chen et al., “Eyeriss: An energy efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE Journal of Solid-State Circuits, vol. 52, No. 1, pp. 127-138, 2017. [cited by applicant]
S. Cho et al., “McDRAM v2: In dynamic random access memory systolic array accelerator to address the large model problem in deep neural networks on the edge,” IEEE Access, vol. 8, pp. 135 223-135 243, 2020. [cited by applicant]
B. Dally, “GTC China 2020 keynote,” https://investor.nvidia.com/events-and-presentations/events-and-presentations/event-details/2020/GTC-China-2020-Keynote-Bill-Dally/default.aspx, 2020. [cited by applicant]
Q. Deng et al., “DrAcc: a dram based accelerator for accurate cnn inference,” in 55thACM/ESDA/IEEE Design Automation Conference (DAC), 2018. [cited by applicant]
F. Devaux, “True Processing In Memory with DRAM accelerator,” HotChips, 2019. [cited by applicant]
J. Devlin et al., “BERT: Pretraining of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: … [cited by applicant]
[4V. Elango et al., “Diesel: Dsl for linear algebra and neural net computations on gpus,” in Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, 2018. [cited by applicant]
M. Emani et al., “Accelerating scientific applications with sambanova reconfigurable dataflow architecture,” Computing in Science Engineering, vol. 23, No. 2, pp. 114-119, 2021. [cited by applicant]
M. Gao et al., “Tetris: Scalable and efficient neural network acceleration with 3d memory,” in Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating S… [cited by applicant]
M. Tan et al., “Efficientdet: Scalable and efficient object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020. [cited by applicant]
G. Urban et al., “Do deep convolutional nets really need to be deep and convolutional?” in 5thInternational Conference on Learning Representations, ICLR 2017. [cited by applicant]
A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, 2017. [cited by applicant]
O. Villa, et al., “Nvbit: A dynamic binary instrumentation framework for nvidia gpus,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO '52, 2019, p. 372-383. [cited by applicant]
G. Wang et al., “Blink: Fast and generic collectives for distributed ml,” in Proceedings of Machine Learning and Systems, vol. 2, 2020, pp. 172-186. [cited by applicant]
M. Wilkening et al., “RecSSD: Near data processing for solid state drive based recommendation inference,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Op… [cited by applicant]
Y. Wu et al., “Tuning applications for efficient gpu offloading to in-memory processing,” in Proceedings of the 34th ACM International Conference on Supercomputing, ser. ICS'20, 2020. [cited by applicant]
Y. Wu et al., “Group normalization,” in Proceedings of the European Conference on Computer Vision (ECCV), Sep. 2018. [cited by applicant]
C. Xie, S. L. Song, J. Wang, W. Zhang, and X. Fu, “Processing-in-memory enabled graphics processors for 3d rendering,” in 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2017, pp. 637… [cited by applicant]
D. Zhang et al., “TOP-PIM: Throughput-oriented programmable processing in memory,” in Proceedings of the 23rd International Symposium on High-Performance Parallel and Distributed Computing, ser. HPDC '14, 2014, p. 85-98. [cited by applicant]
H. Zhang et al., “Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters,” in Proceedings of the 2017 USENIX Conference on Usenix Annual Technical Conference, ser. USENIX ATC '17… [cited by applicant]
“Batchnormalization layer,” https://keras.io/api/layers/normalization_layers/batch_normalization, 2021. [cited by applicant]
“Nvidia a100 tensor core GPU architecture,” https://images.nvidia.com/aem-dam/en-zz/Solutions/datacenter/nvidia-ampere-architecture-whitepaper.pdf, 2020. [cited by applicant]
“Nvidia a100 tensor core GPU,” https://www.nvid/cia.com/content/dam/en-zz/Solutions/DataCenter/a100/pdf/a100-80gb-datasheet-update-nvidia-us-1521051-r2-web.pdf, Jan. 2021. [cited by applicant]
P. Brown, “Graphcore sets new ai performance standards with mk2ipu systems.” [Online]. Available:https://www.graphcore.ai/posts/graphcore-sets-new-aiperformance-standards-with-mk2-ipu-systems. [cited by applicant]
D. Foley et al., “Ultra-performance pascal gpu and nvlink interconnect,” IEEE Micro, vol. 37, No. 2, pp. 7-17, 2017. [cited by applicant]
N. P. Jouppi et al., “A domain-specific supercomputer for training deep neural networks,” Commun. ACM, vol. 63, No. 7, p. 67-78, Jun. 2020. [cited by applicant]
W. Jung et al., “Deepcuts: A deep learning optimization framework for versatile gpu workloads,” in Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, ser. PLD… [cited by applicant]
S. Knowles, “Graphcore Colossus Mk2 IPU,” in 2021 IEEE Hot Chips 33 Symposium(HCS), 2021, pp. 1-25. [cited by applicant]
G. Koo et al., “Access pattern-aware cache management for improving data utilization in GPU,” in Proceedings of the 44th Annual International Symposium on Computer Architecture, ser. ISCA '17, 2017, p. 307-319. [cited by applicant]
K. Lakhotia et al., “In-network reductions on multi-dimensional hyperx,” in 2021 IEEE Symposiumon High-Performance Interconnects (HOTI), 2021, pp. 1-8. [cited by applicant]
S. Lee, “A 1ynm 1.25v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications,” in 2022 IEEE International Solid-State Circuits … [cited by applicant]
S. Lie, “Multi-Million Core, Multi-Wafer AI Cluster,” HotChips,2021. [cited by applicant]
J. Liu et al., “Processing-in-memory for energy-efficient neural network training: A heterogeneous approach,” in 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2018, pp. 655-668. [cited by applicant]
P. Micikevicius et al., “Mixed precision training,” in 6th International Conference on Learning Representations, ICLR 2018. [cited by applicant]
W. Niu et al., “DNNfusion: Accelerating deep neural networks execution with advanced operator fusion,” in Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, s… [cited by applicant]
P. M. Phothilimthana et al., “A flexible approach to autotuning multi-pass machine learning compilers,” in 2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT). [cited by applicant]
S. Rajbhandari et al., “ZeRO: Memory optimization towards training A trillion parameter models,” CoRR, vol. abs/1910.02054, 2019. http://arxiv.org/abs/1910.02054. [cited by applicant]
J. Ren et al., “ZeRO-Offload: Democratizing Billion Scale model training,” in 2021 USENIX Annual Technical Conference. [cited by applicant]
F. Schuiki et al., “A scalable near-memory architecture for training deep neural networks on large in-memory datasets,” IEEE Transactions on Computers, vol. 68, No. 4,pp. 484-497, 2019. [cited by applicant]
N. Vijaykumar et al., “The locality descriptor: A holistic cross-layer abstraction to express data locality in gpus,” in 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA), 2018, pp. 829-8… [cited by applicant]
Z. Wang et al., “Enabling efficient large-scale deep learning training with cache coherent disaggregated memory systems,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2022. [cited by applicant]
H. Zhang et al., “Context encoding for semantic segmentation,” in 2018IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Los Alamitos, CA, USA: IEEE Computer Society, Jun. 2018, pp. 7151-7160. [cited by applicant]
Z. Zheng et al., “Astitch: Enabling anew multi-dimensional optimization space for memory-intensive ml training and inference on modern simt architectures,” in Proceedings of the 27th ACM International Conference on Arch… [cited by applicant]
D. Abts, “Think fast: A tensor streaming processor (TSP) for accelerating deep learning workloads,” in Proceedings of the ACM/IEEE 47th Annual International Symposium on Computer Architecture. IEEE Press, 2020, p. 145-1… [cited by applicant]
J. Ahn, “PIM-enabled instructions: A low-overhead, locality-aware processing-in-memory architecture,” in Proceedings of the 42nd Annual International Symposium on Computer Architecture, New York, NY, USA, 2015, p. 336-3… [cited by applicant]
A. Boroumand, “Google workloads for consumer devices: Mitigating data movement bottlenecks,” in Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating S… [cited by applicant]
G. Chen, “An 8-core RISC-V processor with compute near last level cache in intel 4 cmos,” in 2022 Symposium on VLSI Circuits, 2022. [cited by applicant]
S. Chetlur et al., “cuDNN: Efficient primitives for deep learning,” CoRR, vol. abs/1410.0759, 2014 http://arxiv.org/abs/1410.0759. [cited by applicant]
G. Kim et al., “Memory-centric system interconnect design with Hybrid Memory Cubes,” Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques, 2013, pp. 145-155, doi: 10.1109… [cited by applicant]
M. Imani et al., “FloatPIM: In-memory acceleration of deep neural network training with high precision,” in2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA), 2019, pp. 802-815. [cited by applicant]
J. Park et al., “TRiM: Enhancing processor-memory interfaces with scalable tensor reduction in memory,” in MICRO-54: 54th Annual IEEE/ACM InternationalSymposium on Microarchitecture, ser. MICRO '21, 2021, p. 268-281. [cited by applicant]
A. Pattnaik et al., “Opportunistic computing in GPU architectures,” in Proceedings of the 46th International Symposium on Computer Architecture, ser. ISCA '19, 2019, p. 210-223. [cited by applicant]
D. Wu et al., “SECO: A scalable accuracy approximate exponential function via cross-layer optimization,” in 2019 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), 2019,pp. 1-6. [cited by applicant]
“Nvidia tesla v100 GPU architecture,” Santa Clara, CA, USA, Nvidia, WhitePaper, 2017. [Online]. Available: https://images.nvidia.com/content/voltaarchitecture/pdf/volta-architecture-whitepaper.pdf. [cited by applicant]
“AMD instinct™ MI100 accelerator,” https://www.amd.com/en/products/server-accelerators/instinct-mi100, 2020. [cited by applicant]
“AMD instinct™ mi250x accelerator,” [Online]. Available: https://www.amd.com/en/products/server-accelerators/instinct-mi250x, 2021. [cited by applicant]
“AMD radeon instinct™ mi50 accelerator (16gb),” https://www.amd.com/en/products/professional-graphics/instinct-mi50,2018. [cited by applicant]
“CuBLAS toolkit documentation,” https://docs.nvidia.com/cuda/cublas/index.html, 2012. [cited by applicant]
G. Kim et al., “Multi-GPU System Design with Memory Networks,” 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, 2014, pp. 484-495, doi: 10.1109/MICRO.2014.55. [cited by applicant]
G. Kim et al., “FlexiBuffer: Reducing leakage power in on-chip network routers,” 2011 48th ACM/EDAC/IEEE Design Automation Conference (DAC), 2011, pp. 936-941. [cited by applicant]
G. Kim et al., “Contention-based congestion management in large-scale networks,” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2016, pp. 1-13, doi: 10.1109/MICRO.2016.7783733. [cited by applicant]
G. Kim et al., “TCEP: Traffic Consolidation for Energy-Proportional High-Radix Networks,” 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA), 2018, pp. 712-725, doi: 10.1109/ISCA.2018.0006… [cited by applicant]
G. Kim et al., “Automatically exploiting implicit Pipeline Parallelism from multiple dependent kernels for GPUs,” 2016 International Conference on Parallel Architecture and Compilation Techniques (PACT), 2016, pp. 339-3… [cited by applicant]
S. Xie et al., “Aggregated Residual Transformations for Deep Neural Networks,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 5987-5995, doi: 10.1109/CVPR.2017.634. [cited by applicant]
G. Huang et al., “Densely Connected Convolutional Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2261-2269, doi: 10.1109/CVPR.2017.243. [cited by applicant]
G.-S. Xia et al., “DOTA: A Large-Scale Dataset for Object Detection in Aerial Images,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 3974-3983, doi: 10.1109/CVPR.2018.00418. [cited by applicant]
J. Alammar, “How GPT3 Works—Visualizations and Animations”, http://jalammar.github.io/how-gpt3-works-visualizations-animations/, 2020. [cited by applicant]
A. Chaudhary, “The Illustrated SimCLR Framework”, https://amitness.com/2020/03/illustrated-simclr/, 2020. [cited by applicant]
N. Luehr, “NCCL: Accelerated collective communications for GPUS”, https://on-demand.gputechconf.com/gtc/2016/presentation/s6616-nathan-luehr-nccl.pdf, 2016. [cited by applicant]
A. Gholami et al., “Ai and memory wall,” RiseLab Medium Post, https://medium.com/riselab/ai-and-memory-wall-2cb4265cb0b8, 2021. [cited by applicant]
R. L. Graham et al., “Scalable hierarchical aggregation protocol (sharp): A hardware architecture for efficient data reduction,” First International Workshop on Communication Optimizations in HPC (COMHPC), 2016. [cited by applicant]
J.-B. Grill et al., “Bootstrap your own latent—a new approach to self-supervised learning,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vo… [cited by applicant]
B. Hong et al., “Multi-dimensional parallel training of winograd layer on memory-centric architecture,” in Proceedings of the 51st Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO-51. IEEE Press,… [cited by applicant]
K. Hsieh et al., “Transparent offloading and mapping (tom): Enabling programmer-transparent near-data processing in gpu systems,” in Proceedings of the 43rd International Symposium on Computer Architecture, 2016. [cited by applicant]
J. Hu et al., “Squeeze-and-excitation networks,” in2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 7132-7141. [cited by applicant]
A. Ishii et al., “NVSwitch and DGX-2,” HotChips,2018. [cited by applicant]
A. Ivanov et al., “Data Movement is all you need: A case study on optimizing transformers,” in Proceedings of Machine Learning and Systems, vol. 3, 2020. [cited by applicant]
S. Jeaugey, “S21107-Distributed Training and Fast Inter-GPU communication with NCCL,” Nvidia GPU Technology Conference, 2020. [cited by applicant]
Z. Jia et al., “Dissecting the graphcore IPU architecture via microbenchmarking,” 2019. [cited by applicant]
L. Jiang et al., “XNOR-POP: A processing in-memory architecture for binary convolutional neural networks inwide-io2 drams,” in 2017 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), 2017, pp… [cited by applicant]
N. Jiang et al., “A detailed and flexible cycle accurate network-on-chip simulator,” in 2013 IEEE International Symposium on Performance Analysis of Systems and Software, 2013. [cited by applicant]
N. P. Jouppi et al., “Ten lessons from three generations shaped google's tpuv4i : Industrial product,” in2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), 2021, pp. 1-14. [cited by applicant]
W. Jung et al., “Restructuring batch normalization to accelerate cnn training,” in Proceedings of Machine Learning and Systems, vol. 1, 2019, pp. 14-26. [cited by applicant]
V. Kandiah et al., “Accelwattch: A power modeling framework for modern gpus,” in 54th Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO '21. [cited by applicant]
B. Kim et al., “TRIM: Tensor reduction in memory,” IEEE Computer Architecture Letters, vol. 20,No. 1, pp. 5-8, 2021. [cited by applicant]
D. Kim et al., “Neurocube: A programmable digital neuromorphic architecture with high-density 3d memory,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016. [cited by applicant]
G. Kim et al., “Toward standardized near-data processing with unrestricted data placement for gpus,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. … [cited by applicant]
H. Kim et al., “GradPIM: A practical processing-in-dram architecture for gradient descent,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2021. [cited by applicant]
B. Klenk et al., “An in-network architecture for accelerating shared-memory multiprocessor collectives,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), 2020, pp. 996-1009. [cited by applicant]
D. P. Kingma et al., “ADAM: A method for stochastic optimization,” in 3rd International Conference on Learning Representations, ICLR 2015. [cited by applicant]
Y. Kwon et al., “Tensor casting: Co-designing algorithm-architecture for personalized recommendation training,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2021, pp. 235-248. [cited by applicant]
C. Szegedy et al., “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR), Jun. 2016. [cited by applicant]
S. Li et al., “DRISA: A DRAM-based reconfigurable in-situ accelerator,” in 2017 50thAnnual IEEE/ACM International Symposium on Microarchitecture(MICRO), 2017, pp. 288-301. [cited by applicant]
Y. Li et al., “Accelerating distributed reinforcement learning with in-switch computing,” in Proceedings of the 46th International Symposium on Computer Architecture, ser. ISCA '19, 2019, p. 279-291. [cited by applicant]
K. Kim et al., “Disaggregated memory for expansion and sharing in blade servers,” in Proceedings of the 36th Annual International Symposium on Computer Architecture, ser. ISCA '09, 2009, p. 267-278. [cited by applicant]
T. Lin et al., “Feature pyramid networks for object detection,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Los Alamitos, CA, USA: IEEE Computer Society, Jul. 2017, pp. 936-944. [cited by applicant]
S. Liu et al., “Cambricon: An instruction set architecture for neural networks,” in2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016, pp. 393-405. [cited by applicant]
S. A. Mojumder et al., “MGPU-TSM: A multi gpu system with truly shared memory,” CoRR, vol. abs/2008.02300,2020. [cited by applicant]
R. Nair et al., “Active memory cube: A processing-in-memory architecture for exascale systems,” IBM Journal of Research and Development, vol. 59, No. 2/3, pp. 17:1-17:14, 2015. [cited by applicant]
A. V. Nori et al., “Reduct: Keep it close, keep it cool! : Efficient scaling of dnn inference on multicore cpus with near-cache compute,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (IS… [cited by applicant]
T. Park et al., “Semantic image synthesis with spatially-adaptive normalization,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),2019, pp. 2332-2341. [cited by applicant]
A. Paszke et al., “PyTorch: An imperative style, high performance deep learning library,” in Advances in Neural Information Processing Systems 32. Curran Associates, Inc., 2019. [cited by applicant]
P. Patarasuk et al., “Bandwidth optimal all-reduce algorithms for clusters of workstations,” J. Parallel Distrib. Comput., vol. 69,No. 2, p. 117-124, Feb. 2009. [cited by applicant]
A. Pattnaik et al., “Scheduling techniques for gpu architectures with processing-in-memory capabilities,” in Proceedings of the 2016 International Conference on Parallel Architectures and Compilation, ser. PACT '16, 201… [cited by applicant]
C. Peng et al., “MegDet: A large mini-batch object detector,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 6181-6189. [cited by applicant]
S. Rashidi et al., “Enabling compute-communication overlap in distributed deep learning training platforms,” in ACM/IEEE 48thAnnual International Symposium on Computer Architecture (ISCA), 2021. [cited by applicant]
O. Ronneberger et al., “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015.Cham: Springer International Publishing, 2015, pp. 234-… [cited by applicant]
N. Rotem et al., “Glow: Graph lowering compiler techniques for neural networks,” arXiv:1805.00907v3, 2018. [cited by applicant]
O. Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision, vol. 115, No. 3, 2015. [cited by applicant]
A. Sapio et al., “Scaling distributed machine learning with In-Network aggregation,” in 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI21), Apr. 2021, pp. 785-808. [cited by applicant]
H. Shin et al., “McDRAM: Low latency and energy-efficient matrix computations in dram,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 37, No. 11, pp. 2613-2622, 2018. [cited by applicant]
M. Shoeybi et al., “Megatron-Im: Training multi-billion parameter language models using model parallelism,” arXiv:1909.08053v4, 2020. [cited by applicant]
G. Singh et al., “FPGA-based near-memory acceleration of modern data-intensive applications,” IEEE Micro, vol. 41, No. 4,pp. 39-48, 2021. [cited by applicant]
G. Singh et al., “NERO: A near high-bandwidth memory stencil accelerator for weather prediction modeling,” in 202030th International Conference on Field-Programmable Logic and Applications (FPL), 2020, pp. 9-17. [cited by applicant]
D. Stosic, “Introduction to Mixed Precision Training,” ICCV'19 Tutorial on Accelerating Computer Vision with Mixed Precision, 2019. [cited by applicant]
I. Sutskever et al., “On the importance of initialization and momentum in deep learning,” in Proceedings of the 30th International Conference on International Conference on Machine Learning—vol. 28, ser. ICML'13. JMLR.o… [cited by applicant]
M. Tan et al., “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proceedings of the 36th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 97.… [cited by applicant]
M. Tan et al., “Efficientnetv2: Smaller models and faster training,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 139.PMLR, 2021, pp. 10 09… [cited by applicant]