IP Library Granted Patent US 12,417,078
Granted Patent B2
US 12,417,078 · App. 17/943,078 · Granted Sep 16, 2025

Floating point accumulater with a single layer of shifters in the significand feedback

Inventors: Vojin G. Oklobdzija (Berkeley, CA); Matthew M. Kim (Frankston South, AU)
Assignee: SambaNova Systems, Inc.
G06F7/5443G06F7/483
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,078
App. No.
17/943,078
Granted
Sep 16, 2025
Kind
B2
Abstract

A floating-point accumulator circuit includes a floating-point input having an input significand field and a first shifter coupled to the input significand field and providing an output of the input significand field shifted by a first amount. A carry-save adder has a first, second, and third input and an output. The first input is coupled to the output of the first shifter and the output provides carry bits and sum bits representing a summation of the first input, the second input, and the third input as a significand of the accumulated value. Shifters are also coupled to the carry bits and the sum bits of the output of the carry-save adder to respectively provide the carry bits and the sum bits, both shifted by a second amount, to the second input and the third input of the carry-save adder.

Claims (74)

1. A floating-point accumulator circuit comprising:

a floating-point input having an input significand field;

a first shifter coupled to the input significand field and providing an output of the input significand field shifted by a first amount;

a three-input carry-save adder having a first input coupled to the output of the first shifter, a second input, a third input, and an output providing carry bits and sum bits representing a summation of the first input, the second input, and the third input, as a significand of an accumulated value;

a second shifter coupled to the carry bits of the output of the three-input carry-save adder and providing an output of the carry bits shifted by a second amount to the second input of the three-input carry-save adder; and

a third shifter coupled to the sum bits of the output of the three-input carry-save adder and providing an output of the sum bits shifted by the second amount to the third input of the three-input carry-save adder.

2. The floating-point accumulator circuit of claim 1 , further comprising:

an input pipeline register clocked by a first clock and having an input coupled to the input significand field of the floating-point input and a registered significand output coupled to an input of the first shifter;

a fractional carry pipeline register clocked by the first clock and having an input coupled to the carry bits of the output of the three-input carry-save adder and a registered output coupled to an input of the second shifter; and

a fractional sum pipeline register clocked by the first clock and having an input coupled to the sum bits of the output of the three-input carry-save adder and a registered output coupled to an input of the third shifter.

3. The floating-point accumulator circuit of claim 2 , further comprising:

a first circuit path from the registered significand output of the input pipeline register to the input of the fractional sum pipeline register that passes through only one shifter and the three-input carry-save adder, the only one shifter being the first shifter.

4. The floating-point accumulator circuit of claim 2 , further comprising:

a second circuit path from the registered output of the fractional carry pipeline register to the input of the fractional carry pipeline register that passes through only one shifter and the three-input carry-save adder, the only one shifter being the second shifter; and

a third circuit path from the registered output of the fractional sum pipeline register to the input of the fractional sum pipeline register that passes through only one shifter and the three- input carry-save adder, the only one shifter being the third shifter.

5. The floating-point accumulator circuit of claim 4 , the third circuit path further comprising only one multiplexor.

6. The floating-point accumulator circuit of claim 2 , further comprising a carry-save conversion pipeline stage that includes:

an adder coupled to the registered output of the fractional carry pipeline register and the registered output of the fractional sum pipeline register and having an output to provide their sum; and

a pipeline register clocked by the first clock with an input coupled to the output of the adder and a registered output to provide a significand value of the accumulated value.

7. The floating-point accumulator circuit of claim 6 , wherein the significand value of the accumulated value includes a sign value and an unsigned magnitude value.

8. The floating-point accumulator circuit of claim 2 , further comprising:

an addend input having an addend signficand field; and

a multiplexor having a first input coupled to the addend signficand field of the addend input, a second input coupled to the registered output of the fractional sum pipeline register, and an output coupled to the input of the second shifter.

9. The floating-point accumulator circuit of claim 1 , wherein the floating-point input has one and only one input significand field to represent a magnitude of a significand of a floating-point value received.

10. The floating-point accumulator circuit of claim 1 , wherein the input significand field of the floating-point input uses a 2′s compliment representation of the significand of a floating-point value received.

11. The floating-point accumulator circuit of claim 1 , further comprising:

one and only one gate per bit of the output of the second shifter interposed between a bit of the output of the second shifter and a respective bit of the second input of the three-input carry-save adder; and

one and only one gate per bit of the output of the third shifter interposed between a bit of the output of the third shifter and a respective bit of the third input of the three-input carry-save adder.

12. The floating-point accumulator circuit of claim 11 , wherein the one and only one gate per bit of the output of the second shifter and the one and only one gate per bit of the output of the third shifter each consist of a two-input AND gate.

13. An integrated circuit comprising:

a floating-point input having an input significand field;

a first shifter coupled to the input significand field and providing an output of the input significand field shifted by a first amount;

a three-input carry-save adder having a first input coupled to the output of the first shifter, a second input, a third input, and an output providing carry bits and sum bits representing a summation of the first input, the second input, and the third input, as a significand of an accumulated value;

a second shifter coupled to the carry bits of the output of the three-input carry-save adder and providing an output of the carry bits shifted by a second amount to the second input of the three-input carry-save adder; and

a third shifter coupled to the sum bits of the output of the three-input carry-save adder and providing an output of the sum bits shifted by the second amount to the third input of the three-input carry-save adder.

14. The integrated circuit of claim 13 , further comprising:

an input pipeline register clocked by a first clock and having an input coupled to the floating-point input and a registered output coupled to an input of the first shifter;

a fractional carry pipeline register clocked by the first clock and having an input coupled to the carry bits of the output of the three-input carry-save adder and a registered output coupled to an input of the second shifter; and

a fractional sum pipeline register clocked by the first clock and having an input coupled to the sum bits of the output of the three-input carry-save adder and a registered output coupled to an input of the third shifter.

15. The integrated circuit of claim 14 , further comprising:

a first circuit path from the registered output of the input pipeline register to the input of the fractional sum pipeline register that passes through only one shifter and the three-input carry-save adder, the only one shifter being the first shifter.

16. The integrated circuit of claim 14 , further comprising:

a second circuit path from the registered output of the fractional carry pipeline register to the input of the fractional carry pipeline register that passes through only one shifter and the three-input carry-save adder, the only one shifter being the second shifter; and

a third circuit path from the registered output of the fractional sum pipeline register to the input of the fractional sum pipeline register that passes through only one shifter and the three-input carry-save adder, the only one shifter being the third shifter.

17. The integrated circuit of claim 16 , the third circuit path further comprising only one multiplexor.

18. The integrated circuit of claim 14 , further comprising a carry-save conversion pipeline stage that includes:

an adder coupled to the registered output of the fractional carry pipeline register and the registered output of the fractional sum pipeline register and having an output to provide their sum; and

a pipeline register clocked by the first clock with an input coupled to the output of the adder and a registered output to provide a significand value of the accumulated value.

19. The integrated circuit of claim 18 , wherein the significand value of the accumulated value includes a sign value and an unsigned magnitude value.

20. The integrated circuit of claim 14 , further comprising:

an addend input having an addend signficand field; and

a multiplexor having a first input coupled to the addend signficand field of the addend input, a second input coupled to the registered output of the fractional sum pipeline register, and an output coupled to the input of the second shifter.

21. The integrated circuit of claim 13 , wherein the input significand field of the floating-point input uses a 2′s compliment representation of a significand of a floating- point value received.

22. The integrated circuit of claim 13 , further comprising:

one and only one gate per bit of the output of the second shifter interposed between a bit of the output of the second shifter and a respective bit of the second input of the three-input carry-save adder; and

one and only one gate per bit of the output of the third shifter interposed between a bit of the output of the third shifter and a respective bit of the third input of the three-input carry-save adder.

23. The integrated circuit of claim 22 , wherein the one and only one gate per bit of the output of the second shifter and the one and only one gate per bit of the output of the third shifter each consist of a two-input AND gate.

24. A method of accumulating floating-point values in a computing device, the method comprising:

latching, in a first clock cycle, into a register of the computing device:

a floating-point input having a significand field to generate a latched input significand field,

accumulated carry bits from a three-input carry-save adder circuit of the computing device to generate latched accumulated carry bits, and

accumulated sum bits from the three-input carry-save adder circuit to generate latched accumulated sum bits;

shifting the latched input significand field by a first amount to generate a shifted input significand field;

shifting the latched accumulated carry bits by a second amount to generate shifted carry bits;

shifting the latched accumulated sum bits by the second amount to generate shifted sum bits; and

providing the shifted input significand field, the shifted carry bits, and the shifted sum bits to a first input, a second input, and a third input, respectively of the three-input carry-save adder circuit.

25. The method of claim 24 , wherein the significand field of the floating-point input uses a 2′s compliment representation of a significand of a floating-point value received.

26. The method of claim 24 , further comprising:

adding the latched accumulated carry bits and the latched accumulated sum bits to generate a significand value; and

latching, in a second clock cycle immediately following the first clock cycle, the significand value into a pipleline register of the computing device.

27. The method of claim 26 , wherein the significand value includes a sign value and an unsigned magnitude value.

28. The method of claim 24 , further comprising:

latching, in the first clock cycle, a second floating-point input having an addend signficand field; and

selecting between the addend signficand field of the second floating-point input, the latched accumulated sum bits to shift by the second amount and to provide to the third input of the three-input carry-save adder circuit.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2024
From: OKLOBDZIJA, VOJIN G.; KIM, MATTHEW M.
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 066827/0334 →
Continuity (8)
Continuation 17534376 · Nov 23, 2021
Continuation In Part 17397241 · Aug 9, 2021
Provisional Application 63239384 · Aug 31, 2021
Provisional Application 63190749 · May 19, 2021
Provisional Application 63174460 · Apr 13, 2021
Provisional Application 63166221 · Mar 25, 2021
Provisional Application 63165073 · Mar 23, 2021
Related Publication 20230004353A1 · Jan 5, 2023
References Cited (63)
US 4841467A · Ho et al. · 1989 [cited by applicant]
US 6636223B1 · Morein · 2003 [cited by applicant]
US 6889241B2 · Pangal et al. · 2005 [cited by applicant]
US 6904446B2 · Dibrino · 2005 [cited by applicant]
US 6947962B2 · Hoskote · 2005 [cited by applicant]
US 6988119B2 · Hoskote et al. · 2006 [cited by applicant]
US 7024439B2 · Hoskote · 2006 [cited by applicant]
US 7369637B1 · Mauer · 2008 [cited by applicant]
US 8037119B1 · Oberman et al. · 2011 [cited by applicant]
US 8301681B1 · Lee et al. · 2012 [cited by applicant]
US 20020002573A1 · Landers et al. · 2002 [cited by applicant]
US 20020194240A1 · Pangal et al. · 2002 [cited by applicant]
US 20030154227A1 · Howard · 2003 [cited by examiner]
US 20040068612A1 · Stolowitz · 2004 [cited by applicant]
US 20040128338A1 · Even et al. · 2004 [cited by applicant]
US 20040208173A1 · Gregorio · 2004 [cited by applicant]
US 20070011396A1 · Singh et al. · 2007 [cited by applicant]
US 20100161948A1 · Abdallah · 2010 [cited by applicant]
US 20100191878A1 · Nandagopalan et al. · 2010 [cited by applicant]
US 20140208078A1 · Bradbury et al. · 2014 [cited by applicant]
US 20160126975A1 · Lutz et al. · 2016 [cited by applicant]
US 20180121168A1 · Langhammer · 2018 [cited by applicant]
US 20180129624A1 · Krutsch et al. · 2018 [cited by applicant]
US 20180157464A1 · Lutz et al. · 2018 [cited by applicant]
US 20180204117A1 · Brevdo · 2018 [cited by applicant]
US 20190057053A1 · Tsuchida et al. · 2019 [cited by applicant]
US 20190392296A1 · Brady et al. · 2019 [cited by applicant]
US 20200241797A1 · Kanno et al. · 2020 [cited by applicant]
US 20200387397A1 · Ohta et al. · 2020 [cited by applicant]
US 20210042259A1 · Koeplinger et al. · 2021 [cited by applicant]
US 20210081769A1 · Chen et al. · 2021 [cited by applicant]
US 20210182021A1 · Wang et al. · 2021 [cited by applicant]
US 20220057958A1 · Patriarca et al. · 2022 [cited by applicant]
US 20220066739A1 · Croxford et al. · 2022 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
WO 2019202216A2 · 2019 [cited by applicant]
WO 2021073918A1 · 2021 [cited by applicant]
U.S. Appl. No. 17/942, 941—Notice of Allowance dated Jan. 8, 2024, 9 pages. [cited by applicant]
Bewick, Fast Multiplication: Algorithms and Implementation, Doctoral Dissertation Stanford University, Feb. 1994, 170 pages. [cited by applicant]
Chandrakala, et al., “Design and Implementation of 4-2 Compressor Design with New Xor-Xnor,” Int'l Advanced Research J. in Sci., Engineering and Technology IARJSET, vol. 3, issue 7, Jul. 2016, 4 pages. [cited by applicant]
Cloud TPU, The bfloat16 numerical format, Google Cloud, retrieved on Feb. 14, 2022, 3 pages. Retrieved from the Internet[URL: https://cloud.google.com/tpu/docs/bfloat16 ]. [cited by applicant]
Dadda, Some schemes for parallel multipliers, Alta Fre-quenza, vol. 34, dated Mar. 1965, pp. 349-356. [cited by applicant]
IEEE Computer Society, IEEE Standard for Floating-Point Arithmetic, IEEE Std 754—2008, dated Aug. 29, 2008, 70 pages. [cited by applicant]
Intel BLOAT16-Hardware Numerics Definition White Paper, Rev. 1.0, Nov. 2018, 7 pages. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Oklobdzija et al., A Method for Speed Optimized Partial Product Reduction and Generation of Fast Parallel Multipliers Using an Algorithmic Approach, IEEE Transactions on Computers, vol. 45, No. 3, Mar. 1996, pp. 294-306. [cited by applicant]
PCT/US2022/021223—International Search Report and Written Opinion, dated Jun. 28, 2022, 15 pages. [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Sarma R et al: “A MUX based signed-floating-point MAC architecture using UCM algorithm”, Bulletin of the Polish Academy ofSciences. Technical sciences, Jan. 1, 2020 (Jan. 1, 2020), p. 835, XP055931150, Warsaw.DOI: 10.24… [cited by applicant]
U.S. Appl. No. 17/397,241—Notice of Allowance dated Apr. 14, 2022, 14 pages. [cited by applicant]
U.S. Appl. No. 17/397,241—Office Action dated Feb. 10, 2022, 21 pages. [cited by applicant]
U.S. Appl. No. 17/397,241—Response to Office Action dated Feb. 10, 2022, filed Mar. 18, 2022, 11 pages. [cited by applicant]
U.S. Appl. No. 17/465,558—Notice of Allowance, dated Feb. 25, 2022, 8 pages. [cited by applicant]
U.S. Appl. No. 17/465,558—Office Action dated Jan. 13, 2022, 13 pages. [cited by applicant]
U.S. Appl. No. 17/465,558—Response to Office Action dated Jan. 13, 2022, filed Jan. 31, 2022, 9 pages. [cited by applicant]
U.S. Appl. No. 17/534,376—Notice of Allowance dated Apr. 13, 2022, 17 pages. [cited by applicant]
U.S. Appl. No. 17/534,376—Office Action dated Feb. 18, 2022, 10 pages. [cited by applicant]
U.S. Appl. No. 17/534,376—Response to Office Action dated Feb. 18, 2022, filed Mar. 15, 2022, 10 pages. [cited by applicant]
Vangal et al. “A 6.2-GFlops Floating-Point Multiply-Accumulator With Conditional Normalization,” in IEEE Journal of Solid-State Circuits, vol. 41, No. 10, Oct. 2006, pp. 2314-2323. [cited by applicant]
Vangal et al., “An 80-Tile Sub-100-W TeraFLOPS Processor in 65-nm CMOS,” in IEEE Journal of Solid-State Circuits, vol. 43, No. 1, Jan. 2008, pp. 29-41. [cited by applicant]
Wikipedia, bfloat16 floating-point format, 5 pages, retreived on Dec. 3, 2021. Retrieved from the internet [URL: https://en.wikipedia.org/wiki/Bfloat16_floating-point_format ]. [cited by applicant]
Wikipedia, IEEE 754, retrieved on Feb. 11, 2022, 15 pages Retrieved from the internet [URL: https://en.wikipedia.org/wiki/IEEE_754]. [cited by applicant]
Zhao et al. “Serving Recurrent Neural Networks Efficiently with a Spatial Accelerator”, 2019, Proceedings of the 2nd sysML Conference. (Year: 2019). [cited by applicant]