KR 1020190018888A
· 2019
[cited by applicant]
KR 1020200018188A
· 2020
[cited by applicant]
R. Balasubramonian et al., “Near-Data Processing: Insights from a MICRO-46 Workshop,” in IEEE Micro, vol. 34, No. 4, pp. 36-42, Jul.-Aug. 2014, doi: 10.1109/MM.2014.55. (Year: 2014).
[cited by examiner]
P. Kogge et al., “Processor-In-Memory (PIM) Based Architectures for PetaFlops Potential Massively Parallel Processing”; NASA Grant NAG 5-2998; Jul. 15, 1996; [online] retrieved from https://ntrs.nasa.gov/archive/nasa/ca…
[cited by examiner]
Nvidia, “Nvidia nvlink high-speed interconnect: Application performance”, Nvidia Whitepaper, 2014.
[cited by applicant]
OpenCAPI, “OpenCAPI overview,” OpenCAPI Consortium, 2016.
[cited by applicant]
CCIX, “An Introduction to CCIX,” CCIX Consortium Inc, 2019.
[cited by applicant]
CXL, “Compute Express Link Specification 2.0,” CXL Consortium, 2020, https://www.computeexpresslink.org/download-the-specification, Oct. 2020.
[cited by applicant]
AMD, “Introducing amd cdna architecture,” AMD whitepaper, 2020.
[cited by applicant]
Nvidia, “Nvidia DGX A100 System Architecture,” Nvidia Technical WhitePaper, 2020.
[cited by applicant]
Nvidia. Developer, “Nvidia data center deep learning product performance”, https://developer.nvidia.com/deep-learning-performance-training-inference, Dec. 2021.
[cited by applicant]
M. Abadi et al., “TensorFlow: A system for large-scale machine learning,” in 12thUSENIX Symposium on Operating Systems Design and Implementation (OSDI 16). Savannah, GA: USENIX Association, Nov. 2016, pp. 265-283.
[cited by applicant]
M. Andersch et al., “Tensor Core DL Performance Guide,” Nvidia GPU Technology Conference, 2019.
[cited by applicant]
B. Asgari et al., “Fafnir: Accelerating sparse gathering by using efficient near-memory intelligent reduction,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2021, pp. 908-920.
[cited by applicant]
J. L. Ba et al, “Layer normalization,” CoRR, vol. abs/1607.06450, Jul. 2016.
[cited by applicant]
T. Brown et al., “Language models are Few-Shot Learners,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 20…
[cited by applicant]
M. Caron et al., “Unsupervised learning of visual features by contrasting cluster assignments,” in Advances in Neural Information Processing Systems,H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin,Eds., …
[cited by applicant]
T. Chen et al., “TVM: An automated end-to-end optimizing compiler for deep learning,” in13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). Carlsbad, CA: USENIX Association, Oct. 2018, pp. 57…
[cited by applicant]
T. Chen et al., “Diannao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,” in Proceedings of the 19thInternational Conference on Architectural Support for Programming Languages and Operati…
[cited by applicant]
T. Chen et al., “A simple framework for contrastive learning of visual representations,” in Proceedings of the 37th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, H. D. III …
[cited by applicant]
T. Chen et al, “Big self-supervised models are strong semi-supervised learners,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33.Curra…
[cited by applicant]
Y.-H. Chen et al., “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE Journal of Solid-State Circuits, vol. 52, No. 1,pp. 127-138, Jan. 2017.
[cited by applicant]
S. Cho et al., “McDRAM v2:In-dynamic random access memory systolic array accelerator to address the large model problem in deep neural networks on the edge,” IEEE Access, vol. 8, pp. 135 223-135 243, Jul. 2020.
[cited by applicant]
E. Choukse et al., “Buddy compression: Enabling larger memory for deep learning and HPC workloads on GPUs,” in Proc. ACM/IEEE 47th Annu. Int. Symp. Comput. Archit., pp. 926-939, 2020.
[cited by applicant]
B. Dally et al., “Accelerating Intelligence”, GTC China 2020 keynote, https://investor.nvidia.com/events-and-presentations/events-andpresentations/event-details/2020/GTC-China-2020-Keynote-BillDally/default.aspx, Dec. 2…
[cited by applicant]
Q. Deng, et al., “DrAcc: a DRAM based Accelerator for Accurate CNN Inference”, in 2018 55thACM/ESDA/IEEE Design Automation Conference (DAC), 2018, pp. 1-6.
[cited by applicant]
F. Devaux et al., “True Processing In Memory with DRAM accelerator,” Hot Chips 31, UPMEM, 2019.
[cited by applicant]
J. Devlin, et al., “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics…
[cited by applicant]
V. Elango et al., “Diesel: DSL for linear algebra and neural net computations on GPUs,” in Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, ser. MAPL 2018, New Yor…
[cited by applicant]
M. Emani et al., “Accelerating scientific applications with SambaNova reconfigurable dataflow architecture”, Computing in Science Engineering, vol. 23,No. 2, pp. 114-119, 2021.
[cited by applicant]
M. Gao et al., “TETRIS: Scalable and efficient neural network acceleration with 3D memory,” in Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating S…
[cited by applicant]
N. Gebara et al., “In-network aggregation for shared machine learning clusters,” in Proceedings of Machine Learning and Systems, A. Smola, A. Dimakis, and I. Stoica, Eds., vol. 3, 2021, pp. 829-844.
[cited by applicant]
A. Gholami et al., “AI and Memory Wall,” RiseLab, Medium Post, https://medium.com/riselab/ai-and-memory-wall-2cb4265cb0b8, Mar. 2021.
[cited by applicant]
R. L. Graham et al., “Scalable hierarchical aggregation protocol (SHArP): A hardware architecture for efficient data reduction,” in 2016 First International Workshop on Communication Optimizations in HPC(COMHPC), 2016.
[cited by applicant]
J.-B. Grill et al., “Bootstrap your own latent—a new approach to self-supervised learning,” in Advances in Neural Information Processing Systems, vol. 33. Curran Associates, Inc., 2020.
[cited by applicant]
K. He et al., “Deep residual learning for image recognition,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2016.
[cited by applicant]
M. He et al., “Newton: A DRAM-maker's accelerator-in-memory (AiM) architecture for machine learning,” in Proc. 53rd Annual IEEE/ACM Int. Symp. Microarchitecture, 2020.
[cited by applicant]
M. Hibben, “TSMC, not intel, has the lead in semiconductor processes,” https://seekingalpha.com/article/4151376-tsmc-notintel-lead-in-semiconductor-processes, 2018.
[cited by applicant]
B. Hong et al., “Multi-dimensional parallel training of Winograd layer on memory-centric architecture,” in Proceedings of the 51st Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO-51. IEEE Press,…
[cited by applicant]
K. Hsieh et al., “Transparent offloading and mapping (TOM): Enabling programmer-transparent near-data processing in GPU systems,” SIGARCH Comput. Archit. News, vol. 44, No. 3, p. 204-216, Jun. 2016.
[cited by applicant]
J. Hu et al., “Squeeze-and-excitation networks,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.
[cited by applicant]
S. Ioffe et al., “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proc. 32nd Int. Conf. Int. Conf. Mach.Learn., 2015.
[cited by applicant]
A. Ishii et al., “NVSWITCH and DGX-2, NVLINK-Switching Chip and Scale-Up Compute Server,” HotChips, 2018.
[cited by applicant]
A. Ivanov et al., “Data movement is all you need: A case study on optimizing transformers,” in Proceedings of Machine Learning and Systems, vol. 3, 2021.
[cited by applicant]
S. Jeaugey, “Distributed Training and Fast Inter-GPU communication with NCCL,” NVIDIA GPU Technology Conference, 2020.
[cited by applicant]
Z. Jia et al., “Dissecting the graphcore ipu architecture via microbenchmarking,” Technical Report, CITADEL, High Performance Computing R&D Team, arXiv:1912.03413v1 [cs.DC] Dec. 7, 2019, Dec. 2019.
[cited by applicant]
L. Jiang et al., “XNOR-POP: A processing-in-memory architecture for binary convolutional neural networks in Wide-IO2 DRAMS,” in 2017 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), 2017.
[cited by applicant]
N. Jiang et al., “A detailed and flexible cycle-accurate network-on-chip simulator,” in 2013 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2013.
[cited by applicant]
N. P. Jouppi et al., “Ten lessons from three generations shaped Google's TPUv4i” ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), 2021.
[cited by applicant]
W. Jung et al., “Restructuring batch normalization to accelerate CNN training,” in Proceedings of Machine Learning and Systems, vol. 1, 2019.
[cited by applicant]
V. Kandiah et al., “AccelWattch: A power modeling framework for modern GPUs,” in MICRO-54: 54th AnnualIEEE/ACM International Symposium on Microarchitecture, ser. MICRO '21, New York, NY, USA, Association for Computing M…
[cited by applicant]
M. Khairy et al., “Accel-sim: An extensible simulation framework for validated GPU modeling,” in Proc. ACM/IEEE 47th Annu. Int. Symp. Comput. Archit., 2020.
[cited by applicant]
Y. Wu et al., “Tuning applications for efficient gpu offloading to in-memory processing,” in Proceedings of the 34th ACM International Conference on Supercomputing, ser. ICS'20. New York, NY, USA: Association for Comput…
[cited by applicant]
Y. Wu et al., “Group normalization,” in Proceedings of the European Conference on Computer Vision (ECCV), Sep. 2018.
[cited by applicant]
C. Xie et al., “Processing-in-memory enabled graphics processors for 3d rendering,” in 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2017.
[cited by applicant]
D. Zhang et al., “Top-pim: Throughput-oriented programmable processing in memory,” in Proceedings of the 23rd International Symposium on High-Performance Parallel and Distributed Computing, ser. HPDC '14. New York, NY, …
[cited by applicant]
H. Zhang et al., “Poseidon: An efficient communication architecture for distributed deep learning on gpu clusters,” in Proceedings of the 2017 USENIX Conference on Usenix Annual Technical Conference, ser. USENIX ATC '17…
[cited by applicant]
Keras API reference, “Batchnormalization layer,” https://keras.io/api/layers/normalization_layers/batch_normalization, 2021.
[cited by applicant]
Nvidia, “Nvidia a100 tensor core gpu architecture,” https://images.nvidia.com/aem-dam/en-zz/Solutions/datacenter/nvidia-ampere-architecture-whitepaper.pdf, 2020.
[cited by applicant]
Nvidia, “Nvidia a100 tensor core gpu,” https://www.nvid/cia.com/content/dam/en-zz/Solutions/DataCenter/a100/pdf/a100-80gb-datasheet-update-nvidia-us-1521051-r2-web.pdf, Jan. 2021.
[cited by applicant]
P. Brown, “Graphcore sets new ai performance standards with mk2ipu systems”, https://www.graphcore.ai/posts/graphcore-sets-new-aiperformance-standards-with-mk2-ipu-systems, Dec. 2020.
[cited by applicant]
D. Foley et al., “Ultra-performance pascal gpu and nvlink interconnect,” IEEE Micro, vol. 37, No. 2, pp. 7-17, 2017.
[cited by applicant]
N. P. Jouppi et al., “A domain-specific supercomputer for training deep neural networks,” Commun. ACM, vol. 63, No. 7, p. 67-78, Jun. 2020.
[cited by applicant]
W. Jung et al., “Deepcuts: A deep learning optimization framework for versatile gpu workloads,” in Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, ser. PLD…
[cited by applicant]
S. Knowles, “Graphcore Colossus Mk2 IPU,” in 2021 IEEE Hot Chips 33 Symposium(HCS), 2021, pp. 1-25.
[cited by applicant]
G. Koo, et al., “Access pattern-aware cache management for improving data utilization in gpu,” in Proceedings of the 44th Annual International Symposium on Computer Architecture, ser. ISCA '17. New York, NY, USA: Associ…
[cited by applicant]
K. Lakhotia et al., “In-network reductions on multi-dimensional hyperx,” in 2021 IEEE Symposium on High-Performance Interconnects (HOTI), 2021, pp. 1-8.
[cited by applicant]
S. Lee et al., “A 1ynm 1.25v 8gb, 16gb/s/pin gddr6-basedaccelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications,” in 2022 IEEE International Solid-State Cir…
[cited by applicant]
S. Lie, “Multi-Million Core, Multi-Wafer AI Cluster”, Cerebras Systems, 2021.
[cited by applicant]
J. Liu et al., “Processing-in-memory for energy-efficient neural network training: A heterogeneous approach,” in 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2018, pp. 655-668.
[cited by applicant]
P. Micikevicius et al., “Mixed precision training,” in 6th International Conference on Learning Representations, ICLR 2018.
[cited by applicant]
W. Niu et al., “Dnnfusion: Accelerating deep neural networks execution with advanced operator fusion,” in Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, s…
[cited by applicant]
P. M. Phothilimthana, et al., “A flexible approach to autotuning multi-pass machine learning compilers,” in 2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT), 2021.
[cited by applicant]
S. Rajbhandari et al., “Zero: Memory optimization towards training A trillion parameter models,” CoRR, vol. abs/1910.02054, http://arxiv.org/abs/1910.02054, 2019.
[cited by applicant]
J. Ren et al., “ZeRO-Offload: Democratizing Billion-Scale model training,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21). USENIX Association, Jul. 2021, pp. 551-564.
[cited by applicant]
F. Schuiki et al., “A scalable near-memory architecture for training deep neural networks on large in-memory datasets,” IEEE Transactions on Computers, vol. 68, No. 4, pp. 484-497, 2019.
[cited by applicant]
N. Vijaykumar et al., “The locality descriptor: A holistic cross-layer abstraction to express data locality in gpus,” in 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA), 2018, pp. 829-8…
[cited by applicant]
Z. Wang et al., “Enabling efficient large-scale deep learning training with cache coherent disaggregated memory systems,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2022.
[cited by applicant]
H. Zhang et al., “Context encoding for semantic segmentation,” in 2018IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Los Alamitos, CA, USA: IEEE Computer Society, Jun. 2018, pp. 7151-7160.
[cited by applicant]
Z. Zheng, et al., “Astitch: Enabling a new multi-dimensional optimization space for memory-intensive m Itraining and inference on modern simt architectures,” in Proceedings of the 27th ACM International Conference on Ar…
[cited by applicant]
D. Amodei et al., “Ai and compute,” https://openai.com/blog/ai-and-compute, 2018.
[cited by applicant]
A. Chaudhary, “The illustrated SimCLR framework”, https://amitness.com/2020/03/illustrated-simclr/, 2020.
[cited by applicant]
L. Ke et al., “RecNMP: Accelerating personalized recommendation with near-memory processing,” in Proc. ACM/IEEE 47th Annu. Int. Symp. Comput. Archit., 2020, pp. 790-803.
[cited by applicant]
N. Luehr, “NCCL: Accelerated collective communications for GPUS”, https://on-demand.gputechconf.com/gtc/2016/presentation/s6616-nathan-luehr-nccl.pdf, 2016.
[cited by applicant]
Nvidia, “Convolutional layers user guide,” Nvidia Docs, https://docs.nvidia.com/deeplearning/performance/dl-performanceconvolutional/index.html, 2021.
[cited by applicant]
R. Hwang et al., “Centaur: A chiplet-based, hybrid sparse-dense accelerator for personalized recommendations,” in Proc. ACM/IEEE 47th Annu. Int. Symp. Comput. Archit., 2020, pp. 968-981.
[cited by applicant]
D. Abts et al., “Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads,” IEEE Press, 2020, p. 145-158.
[cited by applicant]
Nvidia, “Nvidia tesla v100 GPU architecture,” Santa Clara, CA, USA, Nvidia, WhitePaper, https://images.nvidia.com/content/voltaarchitecture/pdf/volta-architecture-whitepaper.pdf, 2017.
[cited by applicant]
G. Kim et al., “Memory-centric system interconnect design with Hybrid Memory Cubes,” Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques, 2013, pp. 145-155, doi: 10.1109…
[cited by applicant]
G. Kim et al., “Multi-GPU System Design with Memory Networks,” 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, 2014, pp. 484-495, doi: 10.1109/MICRO.2014.55.
[cited by applicant]
G. Kim et al., “FlexiBuffer: Reducing leakage power in on-chip network routers,” 2011 48th ACM/EDAC/IEEE Design Automation Conference (DAC), 2011, pp. 936-941.
[cited by applicant]
G. Kim et al., “Contention-based congestion management in large-scale networks,” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2016, pp. 1-13, doi: 10.1109/MICRO.2016.7783733.
[cited by applicant]
G. Kim et al., “TCEP: Traffic Consolidation for Energy-Proportional High-Radix Networks,” 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA), 2018, pp. 712-725, doi: 10.1109/ISCA.2018.0006…
[cited by applicant]
G. Kim et al., “Automatically exploiting implicit Pipeline Parallelism from multiple dependent kernels for GPUs,” 2016 International Conference on Parallel Architecture and Compilation Techniques (PACT), 2016, pp. 339-3…
[cited by applicant]
S. Xie et al., “Aggregated Residual Transformations for Deep Neural Networks,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 5987-5995, doi: 10.1109/CVPR.2017.634.
[cited by applicant]
G. Huang et al., “Densely Connected Convolutional Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2261-2269, doi: 10.1109/CVPR.2017.243.
[cited by applicant]
G.-S. Xia et al., “DOTA: A Large-Scale Dataset for Object Detection in Aerial Images,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 3974-3983, doi: 10.1109/CVPR.2018.00418.
[cited by applicant]
J. Alammar, “How GPT3 Works—Visualizations and Animations”, http://jalammar.github.io/how-gpt3-works-visualizations-animations/, 2020.
[cited by applicant]
B. Kim et al., “Trim: Tensor reduction in memory,” IEEE Computer Architecture Letters, vol. 20, No. 1, pp. 5-8, 2021.
[cited by applicant]
D. Kim et al., “Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory,” ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016.
[cited by applicant]
G. Kim, et al., “Toward standardized near-data processing with unrestricted data placement for GPUs,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser.…
[cited by applicant]
H. Kim et al., “GradPIM: A practical processing-in-dram architecture for gradient descent,” IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2021.
[cited by applicant]
Y. Kim et al., “Ramulator: A fast and extensible DRAM simulator,” IEEE Comput. Archit. Lett., vol. 15, No. 1, pp. 45-49, Jan.-Jun. 2016.
[cited by applicant]
D. P. Kingma et al., “Adam: A method for stochastic optimization,” in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015.
[cited by applicant]
B. Klenk et al., “An in-network architecture for accelerating shared-memory multiprocessor collectives,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), 2020.
[cited by applicant]
M. Krause et al., “Gen-Z DRAM and Persistent Memory Theory of Operation,” Gen-Z Consortium, 2019.
[cited by applicant]
Y. Kwon et al., “TensorDIMM: A practical near-memory processing architecture for embeddings and tensor operations in deep learning,” in Proc. 52nd Annu. IEEE/ACM Int. Symp. Microarchit., 2019.
[cited by applicant]
Y. Kwon et al., “Tensor casting: Co-designing algorithm-architecture for personalized recommendation training,” in2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2021.
[cited by applicant]
Y. Kwon et al., “Beyond the memory wall: A case for memory-centric HPC system for deep learning,” in Proc. 51st Annu. IEEE/ACM Int.Symp. Microarchit., 2018.
[cited by applicant]
S. Lee et al., “Hardware architecture and software stack for pim based on commercial dram technology: Industrial product,” in 2021ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), 2021.
[cited by applicant]
S. Li et al., “Drisa: A dram-based reconfigurable in-situ accelerator,” in 201750th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2017.
[cited by applicant]
Y. Li et al., “Accelerating distributed reinforcement learning with in-switch computing,” in Proceedings of the 46th International Symposium on Computer Architecture, ser. ISCA '19. New York, NY, USA: Association for Co…
[cited by applicant]
K. Lim et al., “Disaggregated memory for expansion and sharing in blade servers,” in Proceedings of the 36th Annual International Symposium on Computer Architecture, ser. ISCA '09. New York, NY, USA: Association for Com…
[cited by applicant]
T. Lin et al., “Feature pyramid networks for object detection,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Los Alamitos, CA, USA: IEEE Computer Society, Jul. 2017.
[cited by applicant]
S. Liu et al., “Cambricon: An instruction set architecture for neural networks,” in2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016.
[cited by applicant]
S. Mach et al., “FPnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,” IEEE Trans. VLSI Syst., vol. 29, No. 4, Apr. 2021.
[cited by applicant]
S. A. Mojumder et al., “MGPU-TSM: A multi-gpu system with truly shared memory,” CoRR, vol. abs/2008.02300, 2020.
[cited by applicant]
N. Muralimanohar et al., “Cacti 6.0: A tool to model large caches,” HP Laboratories, Palo Alto, Ca, USA, HPL-2009-85, Tech. Rep. , Apr. 2009, vol. 27.
[cited by applicant]
R. Nair et al., “Active memory cube: A processing-in-memory architecture for exascale systems,” IBM Journal of Research and Development, vol. 59, No. 2/3, 2015.
[cited by applicant]
A. V. Nori et al., “Reduct: Keep it close, keep it cool! : Efficient scaling of dnn inference on multicore cpus with near-cache compute,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (IS…
[cited by applicant]
T. Park et al., “Semantic image synthesis with spatially-adaptive normalization,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),2019.
[cited by applicant]
A. Paszke et al., “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems 32. Curran Associates, Inc., 2019.
[cited by applicant]
P. Patarasuk et al., “Bandwidth optimal all-reduce algorithms for clusters of workstations,” J. Parallel Distrib. Comput., vol. 69,No. 2, Feb. 2009.
[cited by applicant]
A. Pattnaik et al., “Scheduling techniques for gpu architectures with processing-in-memory capabilities,” in Proceedings of the 2016 International Conference on Parallel Architectures and Compilation, ser. PACT '16. New…
[cited by applicant]
C. Peng et al., “Megdet: A large mini-batch object detector,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.
[cited by applicant]
S. Rashidi et al., “Enabling compute-communication overlap in distributed deep learning training platforms,” in 2021 ACM/IEEE48th Annual International Symposium on Computer Architecture(ISCA), 2021.
[cited by applicant]
O. Ronneberger et al., “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015, Cham, Springer International Publishing, 2015.
[cited by applicant]
N. Rotem et al., “Glow: Graph lowering compiler techniques for neural networks,” 2018.
[cited by applicant]
O. Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge, ”International Journal of Computer Vision (IJCV), vol. 115, No. 3, pp. 211-252, 2015.
[cited by applicant]
M. Sandler et al., “MobileNetV2: Inverted residuals and linear bottlenecks,” in Proc. IEEE/CVF Conf. Comput.Vis. Pattern Recognit., 2018.
[cited by applicant]
A. Sapio et al., “Scaling distributed machine learning with In-Network aggregation,” in 18thUSENIX Symposium on Networked Systems Design and Implementation (NSDI 21). USENIX Association, Apr. 2021.
[cited by applicant]
H. Shin et al., “McDRAM: Low latency and energy-efficient matrix computations in dram,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 37, No. 11, 2018.
[cited by applicant]
M. Shoeybi et al., “Megatron-Im: Training multi-billion parameter language models using model parallelism,” 2020.
[cited by applicant]
K. Simonyan et al., “Very deep convolutional networks for largescale image recognition,” in Proc. 3rd Int. Conf. Learn. Representations, https://dblp.org/rec/journals/corr/SimonyanZ14a.bib, 2015.
[cited by applicant]
G. Singh et al., “FPGA-based near-memory acceleration of modern data-intensive applications,” IEEE Micro, vol. 41, No. 4, 2021.
[cited by applicant]
G. Singh et al., “Nero: A near high-bandwidth memory stencil accelerator for weather prediction modeling,” in2020 30th International Conference on Field-Programmable Logic and Applications (FPL), 2020.
[cited by applicant]
D. Stosic, “Introduction to Mixed Precision Training,” ICCV'19 Tutorial on Accelerating Computer Vision with Mixed Precision, 2019.
[cited by applicant]
C. Sun et al., “DSENT—a tool connecting emerging photonics with electronics for opto-electronic networks-on-chip modeling,” in Proc. IEEE/ACM16th Int. Symp. Netw.-on-Chip, 2012.
[cited by applicant]
I. Sutskever et al., “On the importance of initialization and momentum in deep learning,” in Proceedings of the 30th International Conference on International Conference on Machine Learning—vol. 28, ser. ICML'13. JMLR.o…
[cited by applicant]
C. Szegedy et al., “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR), Jun. 2016.
[cited by applicant]
M. Tan et al., “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proceedings of the 36thInternational Conference on Machine Learning, ser. Proceedings of Machine Learning Research, K. Chaudh…
[cited by applicant]
M. Tan et al., “Efficientnetv2: Smaller models and faster training,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 139. PMLR, Jul. 18-24, 20…
[cited by applicant]
M. Tan et al., “Efficientdet: Scalable and efficient object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020.
[cited by applicant]
G. Urban et al., “Do deep convolutional nets really need to be deep and convolutional?” in 5thInternational Conference on Learning Representations, ICLR 2017, Toulon, France, Apr. 24-26, 2017, Conference Track Proceedin…
[cited by applicant]
A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30. Curran Associates, Inc., 2017.
[cited by applicant]
O. Villa et al., “Nvbit: Adynamic binary instrumentation framework for nvidia gpus,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, ser. Micro '52, New York, NY,USA, Association…
[cited by applicant]
G. Wang “Blink: Fast and generic collectives for distributed ml,” in Proceedings of Machine Learning and Systems, vol. 2, 2020.
[cited by applicant]
M. Wilkening et al., “RecSSD: Near data processing for solid state drive based recommendation inference,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Op…
[cited by applicant]