IP Library Granted Patent US 12,340,185
Granted Patent B2
US 12,340,185 · App. 18/746,561 · Granted Jun 24, 2025

Processing core with data associative adaptive rounding

Inventors: Ljubisa Bajic (Toronto, CA); Alex Cejkov (Toronto, CA); Lejla Bajic (Toronto, CA)
Assignee: Tenstorrent AI ULC
G06F7/49947G06N3/082G06F2207/4824
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,185
App. No.
18/746,561
Granted
Jun 24, 2025
Kind
B2
Abstract

Processing cores with data associative adaptive rounding and associated methods are disclosed herein. One disclosed processing core comprises an arithmetic logic unit cluster configured to generate a value for a unit of directed graph data using input directed graph data, a comparator coupled to a threshold register and a data register, a core controller configured to load a threshold value into the threshold register when the value for the unit of directed graph data is loaded into the data register, and a rounding circuit. The rounding circuit is configured to receive the value for the unit of directed graph data from the arithmetic logic unit cluster and conditionally round the value for the unit of directed graph data based on a comparator output from the comparator.

Claims (104)

1. A method, in which each step is conducted by a processing core, comprising:

executing, using a core controller on the processing core, a directed graph;

associating a value with a unit of directed graph data on the processing core;

generating, while executing the directed graph and using a computational unit in the processing core, a value for the unit of directed graph data;

calculating, while executing the directed graph and using the core controller, a threshold value, wherein the calculating of the threshold value uses the value associated with the unit of directed graph data;

conducting, using a comparator on the processing core, a comparison of the value for the unit of directed graph data and the threshold value; and

conditionally rounding, using a rounding circuit on the processing core and based on the comparison, the value for the unit of directed graph data.

2. The method of claim 1 ,

wherein: the calculating of the threshold value uses the value for the unit of directed graph data.

3. The method of claim 2 , wherein:

the directed graph is of an artificial neural network; and

the value for the unit of directed graph data is an activation value.

4. The method of claim 1 ,

wherein: (i) the generating of the value for the unit of directed graph data uses a value for a unit of input directed graph data; and (ii) the calculating of the threshold value uses the value associated with the unit of directed graph data and the value for the unit of input directed graph data.

5. The method of claim 4 , wherein:

the directed graph is of an artificial neural network; and

the value for the unit of directed graph data is an activation value.

6. The method of claim 1 , further comprising:

generating, while executing the directed graph, a value for a second unit of directed graph data;

calculating, while executing the directed graph, a second threshold value;

conducting a second comparison of the value for the second unit of directed graph data and the second threshold value; and

conditionally rounding, based on the second comparison, the value for the second unit of directed graph data.

7. The method of claim 1 , wherein:

the directed graph is of an artificial neural network; and

the value for the unit of directed graph data is an activation value.

8. The method of claim 7 , further comprising:

determining the value associated with the unit of directed graph data using the artificial neural network and test inputs.

9. A processing core, comprising:

a core controller configured to control an execution of a directed graph and a calculation of a threshold value while the directed graph is being executed;

a computational unit configured to generate a value for a unit of directed graph data using input directed graph data during the execution of the directed graph;

a comparator configured to generate a comparator output based on: (i) the threshold value; and (ii) the value for the unit of directed graph data; and

a rounding circuit configured to conditionally round the value for the unit of directed graph data based on the comparator output from the comparator;

wherein: (i) the directed graph is of an artificial neural network; and (ii) the value for the unit of directed graph data is an activation value.

10. The processing core of claim 9 , further comprising:

an association between the threshold value and the value for the unit of directed graph data;

wherein the core controller uses the association to provide the threshold value, along with the value for the unit of directed graph data, to the comparator.

11. The processing core of claim 9 , further comprising:

an association between a value and the value for the unit of directed graph data;

wherein the calculation of the threshold value uses the value.

12. The processing core of claim 9 , further comprising:

an association between a value and the value for the unit of directed graph data;

wherein the calculation of the threshold value uses the value and the value for the unit of directed graph data.

13. The processing core of claim 9 , further comprising:

an association between a value and the value for the unit of directed graph data;

wherein: (i) generating the value for the unit of directed graph data uses a value for a unit of input directed graph data; and (ii) calculating the threshold value uses the value and the value for the unit of input directed graph data.

14. The processing core of claim 9 , wherein:

the computational unit is configured to generate, while executing the directed graph, a value for a second unit of directed graph data;

the core controller is configured to control the calculation of a second threshold value, while executing the directed graph;

the comparator is configured to conduct a second comparison of the value for the second unit of directed graph data and the second threshold value; and

the rounding circuit is configured to conditionally round, based on the second comparison, the value for the second unit of directed graph data.

15. The processing core of claim 9 , wherein:

the core controller is configured to determine the activation value using the artificial neural network and test inputs.

16. A processing core, comprising:

a core controller configured to control an execution of a directed graph and a calculation of a threshold value while the directed graph is being executed;

a computational unit configured to generate a value for a unit of directed graph data using a value for a unit of input directed graph data;

an association between a value and the value for the unit of directed graph data, wherein the processing core is configured to calculate the threshold value using the value;

a comparator configured to generate a comparator output based on the threshold value and the value for the unit of directed graph data; and

a rounding circuit configured to conditionally round the value for the unit of directed graph data based on the comparator output from the comparator.

17. A method, in which each step is conducted by a processing core, comprising:

executing, using a core controller on the processing core, a directed graph;

associating a value with a unit of directed graph data on the processing core;

generating, while executing the directed graph and using a computational unit in the processing core, a value for the unit of directed graph data, wherein the generating of the value for the unit of directed graph data uses a value for a unit of input directed graph data;

calculating, while executing the directed graph and using the core controller, a threshold value, wherein the calculating of the threshold value uses the value associated with the unit of directed graph data and the value for the unit of input directed graph data;

conducting, using a comparator on the processing core, a comparison of the value for the unit of directed graph data and the threshold value; and

conditionally rounding, using a rounding circuit on the processing core and based on the comparison, the value for the unit of directed graph data.

18. A method, in which each step is conducted by a processing core, comprising:

executing, using a core controller on the processing core, a directed graph;

associating a value with a unit of directed graph data on the processing core;

generating, while executing the directed graph and using a computational unit in the processing core, a value for the unit of directed graph data;

calculating, while executing the directed graph and using the core controller, a threshold value, wherein the calculating of the threshold value uses the value associated with the unit of directed graph data;

conducting, using a comparator on the processing core, a comparison of the value for the unit of directed graph data and the threshold value;

conditionally rounding, using a rounding circuit on the processing core and based on the comparison, the value for the unit of directed graph data;

generating, while executing the directed graph a value for a second unit of directed graph data;

calculating, while executing the directed graph, a second threshold value;

conducting a second comparison of the value for the second unit of directed graph data and the second threshold value; and

conditionally rounding, based on the second comparison, the value for the second unit of directed graph data.

19. A method, in which each step is conducted by a processing core, comprising:

executing, using a core controller on the processing core, a directed graph;

associating a value with a unit of directed graph data on the processing core;

generating, while executing the directed graph and using a computational unit in the processing core, a value for the unit of directed graph data;

calculating, while executing the directed graph and using the core controller, a threshold value, wherein the calculating of the threshold value uses the value associated with the unit of directed graph data;

conducting, using a comparator on the processing core, a comparison of the value for the unit of directed graph data and the threshold value; and

conditionally rounding, using a rounding circuit on the processing core and based on the comparison, the value for the unit of directed graph data;

wherein: (i) the directed graph is of an artificial neural network; and (ii) the value for the unit of directed graph data is an activation value.

20. The method of claim 19 , wherein the calculating of the threshold value uses the value for the unit of directed graph data.

21. A processing core, comprising:

a core controller configured to control an execution of a directed graph and a calculation of a threshold value while the directed graph is being executed;

an association between the threshold value and a value for a unit of directed graph data;

a computational unit configured to generate the value for the unit of directed graph data using input directed graph data during the execution of the directed graph;

a comparator configured to generate a comparator output based on: (i) the threshold value; and (ii) the value for the unit of directed graph data; and

a rounding circuit configured to conditionally round the value for the unit of directed graph data based on the comparator output from the comparator;

wherein the core controller uses the association to provide the threshold value, along with the value for the unit of directed graph data, to the comparator.

22. A processing core, comprising:

a core controller configured to control an execution of a directed graph and a calculation of a threshold value while the directed graph is being executed;

a computational unit configured to generate a value for a unit of directed graph data using input directed graph data during the execution of the directed graph;

an association between a value and the value for the unit of directed graph data, wherein the calculation of the threshold value uses the value and the value for the unit of directed graph data;

a comparator configured to generate a comparator output based on: (i) the threshold value; and (ii) the value for the unit of directed graph data; and

a rounding circuit configured to conditionally round the value for the unit of directed graph data based on the comparator output from the comparator.

23. The processing core of claim 22 , wherein:

the directed graph is of an artificial neural network; and

the value for the unit of directed graph data is an activation value.

24. The processing core of claim 22 , wherein:

generating the value for the unit of directed graph data uses a value for a unit of input directed graph data; and

calculating the threshold value uses the value and the value for the unit of input directed graph data.

Assignments (3)
CHANGE OF NAME Recorded Feb 23, 2025
From: TENSTORRENT INC.
To: TENSTORRENT AI INC.
Reel/Frame 070298/0922 →
CHANGE OF NAME Recorded Feb 23, 2025
From: TENSTORRENT AI INC.
To: TENSTORRENT AI ULC
Reel/Frame 070298/0944 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2024
From: BAJIC, LJUBISA; CEJKOV, ALEX; BAJIC, LEJLA
To: TENSTORRENT INC.
Reel/Frame 067770/0112 →
Continuity (5)
Continuation 18132962 · Apr 10, 2023
Continuation 17321925 · May 17, 2021
Continuation 16573728 · Sep 17, 2019
Provisional Application 62738286 · Sep 28, 2018
Related Publication 20240338176A1 · Oct 10, 2024
References Cited (40)
US 5787408A · DeAngelis · 1998 [cited by applicant]
US 6044392A · Anderson et al. · 2000 [cited by applicant]
US 20130007076A1 · Wegener · 2013 [cited by applicant]
US 20140365548A1 · Mortensen · 2014 [cited by applicant]
US 20180018558A1 · Lee et al. · 2018 [cited by applicant]
US 20180046903A1 · Yao et al. · 2018 [cited by applicant]
US 20180232640A1 · Ji et al. · 2018 [cited by applicant]
US 20190080226A1 · Lee et al. · 2019 [cited by applicant]
US 20190228274A1 · Georgiadis et al. · 2019 [cited by applicant]
US 20190230380A1 · Ogasawara · 2019 [cited by applicant]
US 20190362235A1 · Xu et al. · 2019 [cited by applicant]
US 20190377549A1 · Alben et al. · 2019 [cited by applicant]
CN 101882064A · 2010 [cited by applicant]
CN 107251059A · 2017 [cited by applicant]
WO 2018000309A1 · 2018 [cited by applicant]
Communication under Rule 71(3) EPC dated Dec. 13, 2024 from European Application No. 19198828.6, 50 pages. [cited by applicant]
Alemdar et al., “Ternary Neural Networks for Resource-Efficient AI Applications”, ARXIV. Org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 1, 2016, XP080723676, 13 pages. [cited by applicant]
Choi et al., “Fixed-Point Roundoff Error Analysis of Large Feedforward Neural Networks”, Proceedings of 1993 International Joint Conference on Neural Networks, pp. 1947-1950. [cited by applicant]
Esser et al., “Convolutional Networks for Fast, Energy-Efficient Neuromorphic Computing”, Mar. 28, 2016, pp. 1-7, XP055799677, retrieved from the Internet: URL:https://arxiv.org/pdf/1603.08270v1.pdf on Apr. 28, 2021. [cited by applicant]
Examination Report dated Oct. 25, 2021 from European Application No. 19198828.6, 7 pages. [cited by applicant]
Examination Report from EP Application No. 19198828.6 dated May 11, 2023, 4 pages. [cited by applicant]
Extended European Search Report dated Mar. 5, 2020 from European Application No. 19198828.6, 11 pages. [cited by applicant]
First Office Action dated Aug. 27, 2021 from Chinese Patent Application No. 201910936061.1, 30 pages. [cited by applicant]
Gupta et al., “Deep Learning with Limited Numerical Precision”, arXiv:1502.02551v1 [cs.LG], Feb. 9, 2015, 10 pages. [cited by applicant]
Han et al., “Learning Both Weights and Connections for Efficient Neural Networks”, Computer Science, Neural and Evolutionary Computing, arXiv:1506.02626v3, Oct. 30, 2015, 9 pages. [cited by applicant]
Huan et al., “A Multiplication Reduction Technique with Near-Zero Approximation for Embedded Learning in IoT Devices”, IEEE, 978-1-5090-1367, Aug. 2016, pp. 102-107. [cited by applicant]
Köster et al., “Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks”, arXiv:1711.02213v2 [cs.LG] Dec. 2, 2017, 14 pages. [cited by applicant]
Na et al., “On-Chip Training of Recurrent Neural Networks with Limited Numerical Precision”, IEEE, 978-1-5090-6182, Feb. 2017, pp. 3716-3723. [cited by applicant]
Nomura et al., “A Convolutional Neural Network VLSI Architecture Using Sorting Model for Reducing Multiply-and-Accumulation Operations”, ICNC 2005, LNCS 3612, pp. 1006-1014, 2005. [cited by applicant]
Nonfinal Office Action dated Feb. 4, 2021 from U.S. Appl. No. 16/573,728, 38 pages. [cited by applicant]
Nonfinal Office Action dated Nov. 23, 2022 from U.S. Appl. No. 17/321,925, 15 pages. [cited by applicant]
Non-Final Office Action dated Nov. 24, 2023 from U.S. Appl. No. 18/132,962, 17 pages. [cited by applicant]
Notice of allowance and fee due from U.S. Appl. No. 17/321,925 dated Mar. 13, 2023, 6 pages. [cited by applicant]
Notice of Allowance dated Mar. 1, 2021 from U.S. Appl. No. 16/573,728, 18 pages. [cited by applicant]
Notice of Allowance issued Dec. 8, 2021 from Chinese Application No. 201910936061.1, 7 pages. [cited by applicant]
Ott et al., “Recurrent Neural Networks with Limited Numerical Precision”, arXiv:1608.06902v2 [cs.NE] Feb. 26, 2017, 11 pages. [cited by applicant]
Reagen et al., “Minerva: Enabling Low-Power, Highly Accurate Deep Neural Network Accelerators”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Jun. 22, 2016, 12 pages. [cited by applicant]
Wielgosz et al., “Compressing 3D-CNNs for Hardware-Efficient Object Classification in Video Streams”, 2018 International Conference on Signals and Electronic Systems, Sep. 10-12, 2018, 6 pages. [cited by applicant]
English Translation of the First Office Action from the Chinese Application No. 202210166427.3 dated Feb. 6, 2025, 13 pages. [cited by applicant]
First Office Action from the Chinese Application No. 202210166427.3 dated Feb. 6, 2025, 8 pages. [cited by applicant]