IP Library Granted Patent US 12,657,439
Granted Patent B2
US 12,657,439 · App. 17/974,358 · Granted Jun 16, 2026

Fused convolutions for fast deep neural network

Inventors: Swagath Venkataramani (White Plains, NY); Sarada Krithivasan (White Plains, NY); Vijayalakshmi Srinivasan (New York, NY)
Assignee: International Business Machines Corporation
G06N3/048G06N3/044G06N3/0464G06N3/063G06N3/08G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,439
App. No.
17/974,358
Granted
Jun 16, 2026
Kind
B2
Abstract

Fused channel and/or fused filter convolutions for fast deep neural network execution are provided. In one aspect, a system includes: a processor, connected to a memory, configured to: implement an approximated datapath in a deep neural network having a sequence of adders and multipliers for adding up operands to provide accumulated sums for two or more groups of neurons in the deep neural network, and multiplying the accumulated sums to obtain a product; and make an inference using the deep neural network based on the product from the approximated datapath. A method for approximation in a deep neural network is also provided.

Claims (50)

1 . A system comprising:

a processor;

a computer readable storage media; and

program instructions stored on the computer readable storage media for execution by the processor to perform operations comprising:

inputting an input into a deep neural network wherein the deep neural network comprises an intermediate processing element that receives multiple intermediate inputs, respectively, from multiple upstream processing elements disposed in the deep neural network upstream from the intermediate processing element within the deep neural network, wherein the multiple intermediate inputs are based on the input into the deep neural network, and wherein the intermediate processing element processes the multiple intermediate inputs through a sequence comprising:

firstly, adding up a first group of input activations from two or more of the multiple intermediate inputs to provide a first accumulated sum,

adding up a second group of weight tensors for the two or more of the multiple intermediate inputs to provide a second accumulated sum, wherein the weight tensors correspond to the input activations,

subsequently, multiplying the first accumulated sum against the second accumulated sum to obtain a product; and

forwarding an intermediate processing element output to one or more downstream processing elements of the deep neural network, wherein the intermediate processing element output is based on the product; and

receiving, from an output layer of the deep neural network, an inference for the input, wherein the deep neural network determines the inference based on the intermediate processing element output.

2 . A method comprising:

inputting an input into a deep neural network, wherein the deep neural network comprises an intermediate processing element that receives multiple intermediate inputs, respectively, from multiple upstream processing elements disposed in the deep neural network upstream from the intermediate processing element within the deep neural network, wherein the multiple intermediate inputs are based on the input into the deep neural network, and wherein the intermediate processing element processes the multiple intermediate inputs through a sequence comprising:

firstly adding up a first group of input activations from two or more of the multiple intermediate inputs to provide a first accumulated sum, adding up a second group of weight tensors for the two or more of the multiple intermediate inputs to provide a second accumulated sum, wherein the weight tensors correspond to the input activations,

subsequently, multiplying the first accumulated sum against the second accumulated sum to obtain a product, and

forwarding an intermediate processing element output to one or more downstream processing elements of the deep neural network, wherein the intermediate processing element output is based on the product; and

receiving, from an output layer of the deep neural network, an inference for the input, wherein the deep neural network determines the inference based on the intermediate processing element output.

3 . The method of claim 2 , wherein:

the adding up of the first group of input activations comprises adding up more than two input activations to produce the first accumulated sum, and

the adding up of the second group of weight tensors comprises adding up more than two corresponding weight tensors that correspond, respectively, to the more than two input activations to produce the second accumulated sum.

4 . The method of claim 3 , wherein the weight tensors are applied statically during the inference and are determined during a training stage of the deep neural network.

5 . The method of claim 2 , further comprising:

mapping different input channels, different filters, or a combination thereof to the input activations so that the different input channels, the different filters, or the combination thereof are fused.

6 . The method of claim 2 , wherein the intermediate processing element, the multiple upstream processing elements, and the one or more downstream processing elements are respective neurons of the deep neural network.

7 . The method of claim 2 , wherein the intermediate processing element, the multiple upstream processing elements, and the one or more downstream processing elements are respective analog crossbars of the deep neural network, and the deep neural network is a hardware accelerator.

8 . The method of claim 2 , wherein the intermediate processing element, the multiple upstream processing elements, and the one or more downstream processing elements are respective resistive processing units whose resistance is adjusted according to voltage applied.

9 . The method of claim 2 , wherein the product is input into at least one of a normalization function and an activation function whose functional output constitutes the intermediate processing element output.

10 . The method of claim 2 , wherein the intermediate processing element is within a first intermediate layer of the deep neural network and the one or more downstream processing elements are in a downstream layer of the deep neural network, the downstream layer being disposed downstream from the intermediate layer within the deep neural network.

11 . The method of claim 2 , wherein:

the first group of input activations comprises filter dimension input activations for the two or more multiple upstream processing elements in the deep neural network,

the second group of weight tensors comprises corresponding filter kernels, respectively, to the input activations, and

the multiple upstream processing elements are adjacent filters that the sequence fuses together.

12 . The method of claim 2 , wherein the input is an image and the sequence fuses adjacent pixels of an intermediate image produced within the deep neural network.

13 . The method of claim 2 , wherein:

the first group of input activations comprises activation pixels,

the second group of weight tensors comprises filter kernels, and

the filter kernels vary along both channel and filter dimensions.

14 . A computer program product comprising:

a computer readable storage medium; and

program instructions stored on the computer readable storage medium for performing operations comprising:

inputting an input into a deep neural network, wherein the deep neural network comprises an intermediate processing element that receives multiple intermediate inputs, respectively, from multiple upstream processing elements disposed in the deep neural network upstream from the intermediate processing element within the deep neural network, wherein the multiple intermediate inputs are based on the input into the deep neural network, and wherein the intermediate processing element processes the multiple intermediate inputs through a sequence comprising:

firstly, adding up a first group of input activations from two or more of the multiple intermediate inputs to provide a first accumulated sum,

adding up a second group of weight tensors for the two or more of the multiple intermediate inputs to provide a respective second accumulated sum, wherein the weight tensors correspond to the input activations,

multiplying the first accumulated sum against the second accumulated sum to obtain a product, and

forwarding an intermediate processing element output to one or more downstream processing elements of the deep neural network, wherein the intermediate processing element output is based on the product; and

receiving, from an output layer of the deep neural network, an inference for the input, wherein the deep neural network determines the inference based on the intermediate processing element output.

15 . The computer program product of claim 14 , wherein:

the adding up of the first group of input activations comprises adding up more than two input activations to produce the first accumulated sum, and

the adding up of the second group of weight tensors comprises adding up more than two corresponding weight tensors that correspond, respectively, to the more than two input activations to produce the second accumulated sum.

16 . The computer program product of claim 14 , wherein the operations further comprise:

mapping different input channels, different filters, or a combination thereof to the input activations so that the different input channels, the different filters, or the combination thereof are fused.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2022
From: VENKATARAMANI, SWAGATH; KRITHIVASAN, SARADA; SRINIVASAN, VIJAYALAKSHMI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061550/0629 →
Continuity (1)
Related Publication 20240143982A1 · May 2, 2024
References Cited (39)
US 9779786B1 · Wu · 2017 [cited by examiner]
US 10699447B2 · Dwivedi · 2020 [cited by applicant]
US 10706498B2 · Nurvitadhi et al. · 2020 [cited by applicant]
US 11443553B1 · Liu · 2022 [cited by examiner]
US 11468147B1 · Hofer · 2022 [cited by examiner]
US 11586417B2 · Hill · 2023 [cited by examiner]
US 12443835B2 · Hunter · 2025 [cited by examiner]
US 20180053086A1 · Liu · 2018 [cited by examiner]
US 20180137417A1 · Theodorakopoulos · 2018 [cited by examiner]
US 20180314928A1 · Li · 2018 [cited by examiner]
US 20190065192A1 · Tao · 2019 [cited by examiner]
US 20190187771A1 · Ambardekar · 2019 [cited by examiner]
US 20190340499A1 · Burger · 2019 [cited by examiner]
US 20210201003A1 · Banerjee · 2021 [cited by examiner]
US 20220101091A1 · Srinivasa · 2022 [cited by examiner]
US 20220108156A1 · Hunter · 2022 [cited by examiner]
US 20220129759A1 · Yao · 2022 [cited by examiner]
US 20230017662A1 · Kadri · 2023 [cited by examiner]
US 20240127069A1 · Najaf · 2024 [cited by examiner]
WO WO2018154494A1 · 2018 [cited by examiner]
WO WO2020069239A1 · 2020 [cited by examiner]
WO WO2021101790A1 · 2021 [cited by examiner]
WO WO2021102679A1 · 2021 [cited by examiner]
WO WO2022188135A1 · 2022 [cited by examiner]
David Beniaguev et al., “Single cortical neurons as deep artificial neural networks”, vol. 109, Issue 17p. 2727-2739.e3 Sep. 1, 2021, 2727-2739. [cited by examiner]
Junxue Zhang et al., “LiteFlow:towards high-performance adaptive neural networks for kernel datapath”, SIGCOMM '22: Proceedings of the ACM SIGCOMM 2022 Conference pp. 414-427. [cited by examiner]
Ramya Anasseriyil Viswambaran et al., “Evolving Deep Recurrent Neural Networks Using A New Variable-Length Genetic Algorithm”, 2020 IEEE Congress on Evolutionary Computation (CEC) (2020, pp. 1-8). [cited by examiner]
Sudip Paul et al., “Deep Learning and its Importance for Early Signature of Neuronal Disorders”, 2018 4th International Conference on Computing Communication and Automation (ICCCA) (2018, pp. 1-5). [cited by examiner]
Qingchen Zhang et al., “A Tensor-Train Deep Computation Model for Industry Informatics Big Data Feature Learning”, IEEE Transactions on Industrial Informatics (vol. 14, Issue: 7, 2018, pp. 3197-3204). [cited by examiner]
Gao, H. et al., “ChannelNets: Compact and Efficient Convolutional Neural Networks via channel-Wise Convolutions,” 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montreal, Canada (9 pages). [cited by applicant]
Burkov, E. et al., “Deep Neural Networks with Box Convolutions,” 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montreal, Canada (11 pages). [cited by applicant]
Desai, S., “Lecture: Deep Convolutional Neural Networks,” Dec. 2018 (32 pages). [cited by applicant]
Disclosed Anonymously, “Branching Neural Networks,” IPCOM000254791D (Aug. 2018) (11 pages). [cited by applicant]
Disclosed Anonymously, “Design of Neural Networks Based on Cost Estimation,” IPCOM000257359D (Feb. 2019) (11 pages). [cited by applicant]
Disclosed Anonymously, “Determining Priority Value of Processes Based on Usage History,” IPCOM000252344D (Jan. 2018) (39 pages). [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing,” NIST Special Publication 800-145, Sep. 2011 (7 pages). [cited by applicant]
Choi et al., “Accurate and Efficient 2-Bit Quantized Neural Networks,” Proceedings of the 2nd SysML Conference, Palo Alto, CA, USA, 2019 (12 pages). [cited by applicant]
Yang et al., “DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity Measures,” arXiv:1908.09979v2 (Jan. 2020) (18 pages). [cited by applicant]
Hua et al., “Channel Gating Neural Networks,” arXiv:1805.12549v2 (Oct. 2019) (11 pages). [cited by applicant]