IP Library › Granted Patent US 12,423,055
Granted Patent B2
US 12,423,055 · App. 16/683,626 · Granted Sep 23, 2025

Processing apparatus and method of processing add operation therein

Inventors: Deokjin Joo (Seongnam-si, KR); Hossein Moradian Sardroudi (Seoul, KR); Sujeong Jo (Hongseong-gun, KR); Kiyoung Choi (Seoul, KR)
Assignees: Samsung Electronics Co., Ltd.; Seoul National University R&DB Foundation
G06F7/505G06F7/504
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,055
App. No.
16/683,626
Granted
Sep 23, 2025
Kind
B2
Abstract

A method of processing an add operation in a processing apparatus includes acquiring sub-operands from input operands each having an n-bit precision, acquiring intermediate addition results by performing add operations of sub-operands in parallel by using adders, bit-shifting each of the intermediate addition results such that the intermediate addition results correspond to original bit positions in the input operands, and outputting a final addition result of the add operations of the input operands based on the bit-shifted intermediate addition results.

Claims (47)

1. A processor-implemented method of a processor including a plurality of adder circuitries, the method comprising:

generating a plurality of sub-operands, respectively corresponding to bit values of bit sections of a plurality of operands of an input feature map of a neural network, having a first pixel configuration, by dividing the plurality of operands into the bit sections each having a predetermined bit size, wherein each of the plurality of operands has an n-bit precision, and n is a natural number;

generating respective intermediate bit results by operating the plurality of adder circuitries in parallel through each of the plurality of adder circuitries being provided respective multiple sub-operands of the plurality of sub-operands;

generating an output feature map, having a second pixel configuration different from the first pixel configuration, by combining bit-shifted intermediate bit results that result from respective bit-shiftings of the respective intermediate bit results based on respective bit positions in the plurality of operands, such that the combined bit-shifted intermediate bit results correspond to a summation result for the plurality of operands; and

dependent on a performed determination that there is a zero-bit section, among the bit sections, in which a corresponding sub-operand has a zero-value:

in the operating of the plurality of adder circuitries in parallel, controlling a corresponding one adder circuitry, of the plurality of adder circuitries, allocated to the zero-bit section to not operate with respect to the corresponding sub-operand, thereby skipping performance of an add operation with respect to the corresponding sub-operand; and

controlling the corresponding one adder circuitry to operate with respect to an other operand, other than the plurality of operands, to decrease an idling of the corresponding one adder circuitry.

2. The method of claim 1 , wherein

the predetermined bit size is k bits,

the generating of the respective intermediate bit results comprises generating the respective intermediate bit results by using multi-input adder circuitries each having a k-bit precision performing add operations of the respective multiple sub-operands each having a k-bit precision, and

k is a natural number less than n.

3. The method of claim 2 , wherein

a total number of the plurality of operands is m,

each of the respective intermediate bit results comprises a result of executing an add operation of m sub-operands, each having the k-bit precision, acquired from each corresponding bit section in the m operands,

each of the respective intermediate bit results has a precision of (k+log 2m) bits, and

m is a natural number.

4. The method of claim 3 , wherein

each of the plurality of operands is divided into (n/k) bit sections, and

the respective intermediate bit results comprise intermediate bit results from an intermediate bit result of sub-operands of a first bit section of each of the plurality of operands to an intermediate bit result of sub-operands of (n/k)-th bit section of each of the plurality of operands.

5. The method of claim 4 , further comprising performing the determination that there is the zero-bit section by determining whether there is the zero-bit section in which the corresponding sub-operand has the zero-value from among the first bit section to the (n/k)-th bit section,

wherein the controlling of the corresponding one adder circuitry to not operate with respect to the corresponding sub-operand and to operate with respect to the other operand is performed for avoiding the corresponding one adder circuitry from being in an idling state during the generating of the respective intermediate bit results.

6. The method of claim 2 , further comprising performing the respective bit-shiftings of the respective intermediate bit results, including bit-shifting each of the respective intermediate bit results by integer-multiple-of-k bits to correspond to the respective bit positions in each of the plurality of operands.

7. A processing apparatus comprising:

one or more memories storing instructions; and

one or more processors configured to execute the instructions, wherein the execution of the instructions configures the one or more processors to:

generate a plurality of sub-operands, respectively corresponding to bit values of bit sections of a plurality of operands of an input feature map of a neural network, having a first pixel configuration, by dividing the plurality of operands into the bit sections each having a certain bit size, wherein each of the plurality of operands has an n-bit precision, and n is a natural number;

generate respective intermediate bit results by operating a plurality of adder circuitries in parallel through each of the plurality of adder circuitries being provided different respective multiple sub-operands of the plurality of sub-operands;

generate an output feature map, having a second pixel configuration different from the first pixel configuration, by combining bit-shifted intermediate bit results that result from respective bit-shiftings of the respective intermediate bit results based on respective bit positions in the plurality of operands, such that the combined bit-shifted intermediate bit results correspond to a summation result for the plurality of operands; and

dependent on a performed determination that there is a zero-bit section, among the bit sections, in which a corresponding sub-operand has a zero-value:

in the operating of the plurality of adder circuitries in parallel, control a corresponding one adder circuitry, of the plurality of adder circuitries, allocated to the zero-bit section to not operate with respect to the corresponding sub-operand, thereby skipping performance of an add operation with respect to the corresponding sub-operand; and

control the corresponding one adder circuitry to operate with respect to an other operand, other than the plurality of operands, to decrease an idling of the corresponding one adder circuitry.

8. The processing apparatus of claim 7 , wherein

the certain bit size is k bits,

the generation of the respective intermediate bit results comprises generation of the respective intermediate bit results by using multi-input adder circuitries each having a k-bit precision performing add operations of the different respective multiple sub-operands each having a k-bit precision, and

k is a natural number less than n.

9. The processing apparatus of claim 8 , wherein

a total number of the plurality of operands is m,

each of the respective intermediate bit results comprises a result of performing an add operation of m sub-operands each having the k-bit precision and obtained from corresponding bit sections in each of the m operands,

each of the respective intermediate bit results has a precision of (k+log 2m) bits, and

m is a natural number.

10. The processing apparatus of claim 9 , wherein

each of the plurality of operands is divided into (n/k) sub-operands, and

the respective intermediate bit results comprise

intermediate bit results from an intermediate bit result of sub-operands of a first bit section of each of the plurality of operands to an intermediate bit result of sub-operands of (n/k)-th bit section of each of the plurality of operands.

11. The processing apparatus of claim 10 , wherein the one or more processors is further configured to

perform the determination that there is the zero-bit section by determining whether there is the zero-bit section in which the corresponding sub-operand has the zero-value from among the first bit section to the (n/k)-th bit section, and

wherein the controlling of the corresponding one adder circuitry to not operate with respect to the corresponding sub-operand and to operate with respect to the other operand is performed for avoiding the corresponding one adder circuitry from being in an idling state during the generating of the respective intermediate bit results.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2019
From: JOO, DEOKJIN; MORADIAN SARDROUDI, HOSSEIN; JO, SUJEONG; CHOI, KIYOUNG
To: SAMSUNG ELECTRONICS CO., LTD.; SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Reel/Frame 051328/0431 →
Priority Claims (1)
KR 10-2019-0020053 · Feb 20, 2019 · national
Continuity (2)
Provisional Application 62767692 · Nov 15, 2018
Related Publication 20200159495A1 · May 21, 2020
References Cited (20)
US 5134579A · Oki · 1992 [cited by examiner]
US 8468192B1 · Langhammer · 2013 [cited by examiner]
US 10380753B1 · Csordás · 2019 [cited by examiner]
US 20090228538A1 · Nagano et al. · 2009 [cited by applicant]
US 20140122555A1 · Hickmann · 2014 [cited by examiner]
US 20140337262A1 · Kato et al. · 2014 [cited by applicant]
US 20150193202A1 · Chen et al. · 2015 [cited by applicant]
US 20170357891A1 · Judd et al. · 2017 [cited by applicant]
US 20190114140A1 · Langhammer · 2019 [cited by examiner]
KR 1020170023708A · 2017 [cited by applicant]
WO WO2007052499A1 · 2007 [cited by applicant]
Chen on “Parallel Adders” Lecture Notes. Retrieved on [Mar. 22, 2023]. Retrieved on the Internet <https://users.encs.concordia.ca/˜asim/COEN_6501/Lecture_Notes/L2_Notes.pdf> (Year: 2005). [cited by examiner]
Anonymous, “Parsing” Definition on Google.com (Year: 2024). [cited by examiner]
Anonymous, “Parsing” Definition on https://www.merriam-webster.com/dictionary/parsing (Year: 2024). [cited by examiner]
Anonymous, “Multiples” Definition on Google.com (Year: 2024). [cited by examiner]
Hennessy et al., “Computer Organization and Design: The Hardware/Software Interface”, Fifth Edition, Chapter 1 pp. 2-59, 2014. Retrieved from <https://ict.iitk.ac.in/wp-content/uploads/CS422-Computer-Architecture-Comput… [cited by examiner]
Peddawad et al., “Matrix-Matrix Multiplication Using Systolic Array Architecture in Bluespec,” Systolic Array & Bluespec, CS6230: CAD for VLSI, 8 pages in English. [cited by applicant]
Rastegari et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Allen Institute for AI, University of Washington, 17 pages in English. [cited by applicant]
Courbariaux et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” Advances in neural information processing systems 2016, 11 pages in English. [cited by applicant]
Moradian et al., “Reconfigurable Multi-Input Adder Design for Deep Neural Network Accelerators,” 2018 IEEE, ISOCC 2018, pp. 212-213. [cited by applicant]
Cited By (1)
US 12,693,990