IP Library › Granted Patent US 12,737,615
Granted Patent B2
US 12,737,615 · App. 17/544,688 · Granted Sep 15, 2026

Method and apparatus for optimizing batch size for artificial neural network accelerator

Inventor: Hyun Mi Kim (Daejeon, KR)
Assignee: Electronics and Telecommunications Research Institute
G06N3/08G06F9/5027G06F18/217
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,615
App. No.
17/544,688
Filed
Dec 7, 2021
Granted
Sep 15, 2026
Kind
B2
Examiner
LO, ANN J
Art Unit
2159
USPC
706/15
Abstract

A method for optimizing a batch size for an artificial neural network accelerator that processes at least one batch in an apparatus for optimizing a batch size is provided. The method for optimizing a batch size includes: receiving information from an artificial neural network to determine a batch size; and determining the batch size for optimizing basic performance of the artificial neural network according to the artificial neural network.

Claims (26)

1 . A method performed by an artificial neural network accelerator, the method comprising:

receiving information via an input interface from an artificial neural network to determine a batch size for batches within internal memory banks of the artificial neural network,

wherein the information from the artificial neural network includes information associated with a hardware configuration of the internal memory banks of the artificial neural network;

optimizing a throughput based performance metric of the artificial neural network to determine the batch size for the artificial neural network, by:

determining a plurality of candidate numbers of batches constrained by a total number of the internal memory banks;

evaluating the throughput based performance metric of the artificial neural network for each of the plurality of the candidate numbers of batches;

selecting, as the batch size, a number of batches corresponding to a maximum value of the throughput based performance metric among the plurality of candidate numbers of batches; and

performing operations within the artificial neural network accelerator based on the determined batch size in parallel using the internal memory banks.

2 . The method of claim 1 , wherein the determining of the batch size includes determining one batch size for the entire artificial neural network.

3 . The method of claim 1 , wherein the determining of the batch size includes determining a different batch size for each layer of the artificial neural network.

4 . The method of claim 3 , wherein the determining of a different batch size for each layer includes:

optimizing the throughput based performance metric of the artificial neural network based on each batch size for each layer;

determining a batch size having a best throughput maximum value of the throughput based performance metric for each layer as the batch size of the layer.

5 . An apparatus, comprising:

an input interface that receives information from an artificial neural network to determine a batch size,

wherein the information from the artificial neural network includes information associated with internal memory banks of the artificial neural network, including an internal memory bank structure and an internal memory bank size of the artificial neural network;

a processor optimizes a throughput based performance metric of the artificial neural network to determine the batch size for the artificial neural network, by:

determining a plurality of candidate numbers of batches constrained by a total number of internal memory banks;

evaluating the throughput based performance metric of the artificial neural network for each of the plurality of the candidate numbers of batches;

selecting, as the batch size, a number of batches corresponding to a maximum value of the throughput based performance metric among the plurality of candidate numbers of batches; and

performing operations within a artificial neural network accelerator based on the determined batch size in parallel using the internal memory banks; and

an output interface that transmits the determined batch size to the artificial neural network accelerator.

6 . The apparatus of claim 5 , wherein the processor determines a different batch size for each layer of the artificial neural network.

7 . The apparatus of claim 6 , wherein the processor optimizes a throughput based performance metric for the artificial neural network based on each batch size for each layer and determines a batch size having a maximum value of the throughput based performance metric for each layer as the batch size of the layer.

8 . The apparatus of claim 5 , wherein the processor buffers input data to match the determined batch size.

9 . The apparatus of claim 5 , wherein the processor optimizes the throughput based performance metric of the artificial neural network by changing a number of batches associated with the batch size to a total number of internal memory banks of the artificial neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2021
From: KIM, HYUN MI
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 058335/0680 →
Continuity (1)
Related Publication 20220180192A1 · Jun 9, 2022
References Cited (21)
US 9842293B2 · Young · 2017 [cited by applicant]
US 10083395B2 · Young · 2018 [cited by applicant]
US 11049006B2 · Langford et al. · 2021 [cited by applicant]
US 20160342890A1 · Young · 2016 [cited by applicant]
US 20180150745A1 · Shirahata · 2018 [cited by examiner]
US 20190079801A1 · Lyuh et al. · 2019 [cited by applicant]
US 20190122107A1 · Young · 2019 [cited by applicant]
US 20200151568A1 · Lee et al. · 2020 [cited by applicant]
US 20200226458A1 · de Vangel · 2020 [cited by examiner]
US 20200372013A1 · Lee et al. · 2020 [cited by applicant]
US 20200372343A1 · Mikami · 2020 [cited by applicant]
US 20210224654A1 · Young · 2021 [cited by applicant]
CN 104035751A · 2014 [cited by examiner]
JP 202042591A · 2020 [cited by applicant]
JP 202042753A · 2020 [cited by applicant]
KR 1020190054449A · 2019 [cited by applicant]
KR 1020200045017A · 2020 [cited by applicant]
KR 102103644B1 · 2020 [cited by applicant]
KR 102106144B1 · 2020 [cited by applicant]
KR 1020200134944A · 2020 [cited by applicant]
Machine Translation of CN104035751 (Year: 2014). [cited by examiner]