IP Library › Granted Patent US 12,254,400
Granted Patent B2
US 12,254,400 · App. 16/244,267 · Granted Mar 18, 2025

Optimizing artificial neural network computations based on automatic determination of a batch size

Inventors: Benoit Chappet de Vangel (Paris, FR); Thomas Cagnac (Longpont sur Orge, FR); Benjamin Poumarede (Paris, FR); Ludovic Larzul (El Dorado Hills, CA)
Assignee: Mipsology SAS
G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,400
App. No.
16/244,267
Granted
Mar 18, 2025
Kind
B2
Abstract

Systems and methods for optimizing artificial neural network (ANN) computations based on automatic determination of a batch size are disclosed. An example method may comprise receiving, by an optimization module, an ANN structure associated with the ANN, and generating, based on the ANN structure, a configuration for a computation engine capable of performing computation of the layers of the ANN. The configuration may include information concerning a batch size of one or more layers of the ANN. The batch size of a layer can be determined based on a bandwidth required to read data related to layer, a number of parameters associated with the layer, and a time the layer processes one input dataset from the batch. The batch size of the layer can differ from the batch size of the ANN. The batch size of the layer may differ from a batch size of another layer of ANN.

Claims (56)

1. A system comprising:

a computation circuit configured to perform computations of one or more layers of an artificial neural network (ANN) for a series of input datasets; and

one or more processors in communication with the computation circuit, wherein the one or more processors are configured to initiate operations including:

receiving an ANN structure and a user input, the ANN structure being associated with the ANN and the user input including a specified performance measure for the ANN; and

generating, based on the ANN structure and the user input, a configuration for the computation circuit and a memory configuration associated with the ANN, the memory configuration including a plurality of memories allocated to storing data associated with the one or more layers of the ANN, wherein the configuration corresponds to the specified performance measure and includes information concerning batch sizes of the one or more layers of the ANN, and wherein:

the one or more layers of the ANN are assigned the batch sizes based on the configuration;

the ANN includes a first layer, a second layer, and a third layer, wherein an output of the first layer is an input to the second layer and the third layer; and

the computation circuit is configured to:

determine a first latency of storing the data associated with the one or more layers in the plurality of memories when a computation of the first layer is executed prior to a computation of the second layer;

determine a second latency of storing the data associated with the one or more layers in the plurality of memories when the computation of the first layer is executed after the computation of the second layer;

determine that the second latency is less than the first latency; and

in response to the determination that the second latency is less than the first latency, perform at least one computation of the first layer before a computation of the second layer and at least one further computation of the first layer after the computation of the second layer and prior to a computation of the third layer.

2. The system of claim 1 , wherein the one or more processors are configured to determine the batch sizes of the one or more layers based on a bandwidth required to read the data associated with the one or more layers.

3. The system of claim 1 , wherein the one or more processors are configured to determine the batch sizes of the one or more layers based on a number of parameters associated with the one or more layers.

4. The system of claim 1 , wherein the one or more processors are configured to determine the batch sizes of the one or more layers based on a time needed for the one or more layers to process one input dataset from a batch.

5. The system of claim 1 , wherein the one or more processors are configured to determine a batch size for the ANN based on the batch sizes of the one or more layers of the ANN.

6. The system of claim 5 , wherein at least one of the batch sizes of the one or more layers differs from the batch size of the ANN.

7. The system of claim 1 , wherein a batch size of the first layer differs from a batch size of the second layer.

8. The system of claim 1 , wherein the computation circuit is configured to repeat computations of a subpart of the ANN for different input datasets from the series of input datasets.

9. The system of claim 1 , wherein the computation circuit is configured to perform, concurrently, a computation of the first layer of the ANN for a first input dataset of the series of input datasets and a computation of the first layer for a second input dataset of the series of input datasets prior to the computation of the second layer of the ANN, wherein an input dataset of the second layer includes an output dataset of the first layer.

10. The system of claim 1 , wherein the computation circuit is implemented on a field-programmable gate array.

11. The system of claim 1 , wherein the one or more processors are configured to perform one or more iterations of selecting the batch sizes of the one or more layers of the ANN to optimize a performance measure, the performance measure being a function of one or more of: a batch size of the ANN, a latency of the ANN, and a throughput of the ANN.

12. The system of claim 11 , wherein the one or more processors are configured to perform the one or more iterations until a number of the one or more iterations exceeds a predetermined threshold or the performance measure matches the desired specified performance measure.

13. The system of claim 11 , wherein a batch size of the second layer is less than a batch size of the third layer.

14. The system of claim 11 , wherein the one or more processors are configured to select the batch sizes of the one or more layers of the ANN based on a heuristic algorithm.

15. A method comprising:

receiving, by one or more processors in communication with a computation circuit configured to perform computations of one or more layers of an artificial neural network (ANN) for a series of input datasets, an ANN structure and a user input, the ANN structure being associated with the ANN and the user input including a specified performance measure for the ANN;

generating, by the one or more processors and based on the ANN structure and the user input, a configuration for the computation circuit and a memory configuration associated with the ANN, the memory configuration including a plurality of memories allocated to storing data associated with the one or more layers of the ANN, wherein the configuration corresponds to the specified performance measure and includes information concerning batch sizes of the one or more layers of the ANN, and wherein:

the one or more layers of the ANN are assigned the batch sizes based on the configuration;

the ANN includes a first layer, a second layer, and a third layer, wherein an output of the first layer is an input to the second layer and the third layer; and

the computation circuit is configured to:

determine a first latency of storing the data associated with the one or more layers in the plurality of memories when a computation of the first layer is executed prior to a computation of the second layer;

determine a second latency of storing the data associated with the one or more layers in the plurality of memories when the computation of the first layer is executed after the computation of the second layer;

determine that the second latency is less than the first latency; and

in response to the determination that the second latency is less than the first latency, perform at least one computation of the first layer before a computation of the second layer and at least one further computation of the first layer after the computation of the second layer and prior to a computation of the third layer; and

determining, by the one or more processors and based on the batch sizes of the one or more layers of the ANN, a batch size for the ANN.

16. The method of claim 15 , wherein the batch sizes of the one or more layers are determined based on one or more of:

a bandwidth required to read the data associated with the one or more layers;

a number of parameters associated with the one or more layers; and

a time needed for the one or more layers to process one input dataset from a batch.

17. The method of claim 15 , wherein at least one of the batch sizes of the one or more layers differs from the batch size of the ANN.

18. The method of claim 15 , wherein a batch size of the first layer is different from a batch size of the second layer.

19. The method of claim 15 , further comprising performing, by the one or more processors, one or more iterations of selecting the batch sizes of the one or more layers of the ANN to optimize a performance measure for the ANN, the performance measure being a function of one or more of: a latency of the ANN, a throughput of the ANN, and a user-specified batch size of the ANN.

20. A system comprising:

a computation circuit configured to perform computations of one or more layers of an artificial neural network (ANN) for a series of input datasets; and

one or more processors in communication with the computation circuit, wherein the one or more processors are configured to initiate operations including:

receiving an ANN structure and a user input, the ANN structure being associated with the ANN and the user input including a specified performance measure for the ANN;

generating, based on the ANN structure and the user input, a configuration for the computation circuit and a memory configuration associated with the ANN, the memory configuration including a plurality of memories allocated to storing data associated with the one or more layers of the ANN, wherein the configuration corresponds to the specified performance measure and includes information concerning batch sizes of the one or more layers of the ANN, and wherein:

the one or more layers of the ANN are assigned the batch sizes based on the configuration;

the ANN includes a first layer, a second layer, and a third layer, wherein an output of the first layer is an input to the second layer and the third layer; and

the computation circuit is configured to:

determine a first latency of storing the data associated with the one or more layers in the plurality of memories when a computation of the first layer is executed prior to a computation of the second layer;

determine a second latency of storing the data associated with the one or more layers in the plurality of memories when the computation of the first layer is executed after the computation of the second layer;

determine that the second latency is less than the first latency; and

in response to the determination that the second latency is less than the first latency, perform at least one computation of the first layer before a computation of the second layer and at least one further computation of the first layer after the computation of the second layer and prior to a computation of the third layer; and

determining a batch size for the ANN based on the batch sizes of the one or more layers of the ANN.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2026
From: MIPSOLOGY SAS
To: XILINX, INC.
Reel/Frame 073663/0308 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2019
From: DE VANGEL, BENOIT CHAPPET; CAGNAC, THOMAS; POUMAREDE, BENJAMIN; LARZUL, LUDOVIC
To: MIPSOLOGY SAS
Reel/Frame 047949/0484 →
Continuity (1)
Related Publication 20200226458A1 · Jul 16, 2020
References Cited (18)
US 7711797B1 · Huang · 2010 [cited by examiner]
US 10846096B1 · Chung · 2020 [cited by examiner]
US 20090138770A1 · Nakaya · 2009 [cited by examiner]
US 20170228645A1 · Wang · 2017 [cited by examiner]
US 20170286861A1 · Kelly · 2017 [cited by examiner]
US 20170344882A1 · Ambrose · 2017 [cited by examiner]
US 20180307894A1 · Lim · 2018 [cited by examiner]
US 20180373976A1 · Woo · 2018 [cited by examiner]
US 20190087721A1 · Prakash · 2019 [cited by examiner]
US 20190228298A1 · Suzuki · 2019 [cited by examiner]
US 20200042362A1 · Cui · 2020 [cited by examiner]
US 20200042856A1 · Datta · 2020 [cited by examiner]
US 20200125926A1 · Choudhury · 2020 [cited by examiner]
US 20200175374A1 · Hestness · 2020 [cited by examiner]
EP 3346425A1 · 2018 [cited by applicant]
A. Devarakonda et al. AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks. Feb. 14, 2018. [retrieved from arXiv on Jan. 29, 2022] <URL: https://arxiv.org/pdf/1712.02029.pdf> (Year: 2018). [cited by examiner]
NIST. Engineering Statistics Handbook: 5.5.3. How do you optimize a process? [archived on Feb. 19, 2001] [retrieved on Jun. 27, 2022] <URL: https://web.archive.org/web/20010219000909/https://www.itl.nist.gov/div898/hand… [cited by examiner]
DT. Vooturi et al. Efficient Inferencing of Compressed Deep Neural Networks. arXiv. Nov. 1, 2017. [retrieved on Jun. 27, 2022] <URL: https://arxiv.org/pdf/1711.00244.pdf> (Year: 2017). [cited by examiner]