IP Library › Granted Patent US 12,632,303
Granted Patent B2
US 12,632,303 · App. 18/337,723 · Granted May 19, 2026

Accelerator, method of operating the same, and electronic device including the same

Inventors: Wookeun Jung (Seoul, KR); Jaejin Lee (Seoul, KR); Seung Wook Lee (Suwon-si, KR)
Assignees: Samsung Electronics Co., Ltd.; Seoul National University R & DB Foundation
G06F9/5027G06F9/30014G06F9/3802G06F9/3836G06F9/48G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,303
App. No.
18/337,723
Granted
May 19, 2026
Kind
B2
Abstract

A processor-implemented accelerator method includes: reading, from a memory, an instruction to be executed in an accelerator; reading, from the memory, input data based on the instruction; and performing, on the input data and a parameter value included in the instruction, an inference task corresponding to the instruction.

Claims (51)

1 . A processor-implemented method of operating an electronic device, the method comprising:

reading an instruction to be executed in an accelerator included in the electronic device, the instruction embedding a parameter value of at least some portion of layers in a neural network for an inference task, wherein the reading of the instruction comprises reading a plurality of instructions each including a partial value of the parameter value to be used in the inference task, wherein a maximum length of the partial value of the parameter value included in each of the plurality of instructions is less than a preset threshold value;

reading input data based on the instruction;

reading, from the instruction, the parameter value for the inference task; and

performing, on the input data and the parameter value embedded in the instruction, the inference task instructed by the instruction, and

wherein the parameter value is a value fixed for the inference task.

2 . The method of claim 1 , wherein the instruction is determined by substituting, with the parameter value, an indicator of the parameter value included in an initial code.

3 . The method of claim 2 , wherein

the indicator included in the initial code uses a loop variable index, and

the loop variable index is converted to an invariable index through loop unrolling.

4 . The method of claim 1 , wherein the instruction is determined such that a same parameter value is used in a thread block including a plurality of threads to be processed in the accelerator.

5 . The method of claim 4 , wherein an operation unit included in the accelerator and having an instruction cache is configured to process the thread block.

6 . The method of claim 1 ,

wherein the parameter value is determined based on each of the partial values respectively included in a corresponding one of the plurality of instructions.

7 . The method of claim 6 , wherein each of the plurality of instructions includes information indicating which part of the parameter value corresponds to the partial value of the parameter value included in the respective instruction.

8 . The method of claim 6 , wherein

the plurality of instructions includes a first instruction including a mantissa part of the parameter value and a second instruction including an exponent part of the parameter value, and

the performing of the inference task comprises performing a multiplication operation on the mantissa part in response to the first instruction being read, and performing an addition operation on the exponent part in response to the second instruction being read.

9 . The method of claim 1 , wherein

a maximum length of a parameter value portion of an instruction is less than a length of the parameter value,

the parameter value is stored in a memory as a plurality of parameter value portions, each of the parameter value portions being included in a respective instruction,

the reading of the instruction comprises reading the respective instructions, and

the parameter value is determined based on the parameter value portions.

10 . The method of claim 1 , wherein the parameter value is a parameter included in a neural network, and further comprising performing any one of speech recognition, machine translation, machine interpretation, object recognition, and pattern recognition based on a result of the performing of the inference task.

11 . A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the method of claim 1 .

12 . An accelerator comprising:

at least one processing element configured to:

read an instruction to be executed in an accelerator, the instruction embedding a parameter value of at least some portion of layers in a neural network for an inference task, wherein the reading of the instruction comprises reading a plurality of instructions each including a partial value of the parameter value to be used in the inference task, wherein a maximum length of the partial value of the parameter value included in each of the plurality of instructions is less than a preset threshold value;

read input data based on the instruction;

reading, from the instruction, the parameter value for the inference task; and

perform, on the input data and the parameter value embedded in the instruction, the inference task instructed by the instruction, and

wherein the parameter value is a value fixed for the inference task.

13 . The accelerator of claim 12 , wherein the instruction is determined by substituting, with the parameter value, an indicator of the parameter value included in an initial code.

14 . The accelerator of claim 13 , wherein

the indicator included in the initial code uses a loop variable index, and

the loop variable index is converted to an invariable index through loop unrolling.

15 . The accelerator of claim 12 , wherein the instruction is determined such that a same parameter value is used in a thread block including a plurality of threads to be processed in the accelerator.

16 . The accelerator of claim 15 , wherein an operation unit included in the accelerator and having an instruction cache is configured to process the thread block.

17 . The accelerator of claim 12 ,

wherein the parameter value is determined based on each of the partial values respectively included in a corresponding one of the plurality of instructions.

18 . The accelerator of claim 17 , wherein each of the plurality of instruction includes information indicating which part of the parameter value corresponds to the partial value included in the respective instruction.

19 . An electronic device comprising a memory storing the instruction and the accelerator of claim 13 .

20 . An electronic device comprising:

a memory configured to store an instruction to be executed in an accelerator; and

the accelerator configured to:

read the instruction from the memory, the instruction embedding a parameter value of at least some portion of layers in a neural network for an inference task,

wherein the reading of the instruction comprises reading a plurality of instructions each including a partial value of the parameter value to be used in the inference task, wherein a maximum length of the partial value of the parameter value included in each of the plurality of instructions is less than a preset threshold value;

read input data based on the instruction;

read, from the instruction, the parameter value for the inference task; and

perform, on the input data and the parameter value embedded in the instruction, the inference task instructed by the instruction, and

wherein the parameter value is a value fixed for the inference task.

Priority Claims (1)
KR 10-2020-0075682 · Jun 22, 2020 · national
Continuity (2)
Continuation 17145958 · Jan 11, 2021
Related Publication 20230333899A1 · Oct 19, 2023
References Cited (19)
US 9558094B2 · Zhou · 2017 [cited by applicant]
US 9972063B2 · Ashari et al. · 2018 [cited by applicant]
US 10157045B2 · Venkataramani et al. · 2018 [cited by applicant]
US 10223762B2 · Ashari et al. · 2019 [cited by applicant]
US 20170032487A1 · Ashari et al. · 2017 [cited by applicant]
US 20180136912A1 · Venkataramani et al. · 2018 [cited by applicant]
US 20180189633A1 · Henry et al. · 2018 [cited by applicant]
US 20180322381A1 · Liu · 2018 [cited by examiner]
US 20190114529A1 · Ng · 2019 [cited by examiner]
US 20190220278A1 · Adelman et al. · 2019 [cited by applicant]
US 20200257986A1 · Diamantopoulos et al. · 2020 [cited by applicant]
US 20200295778A1 · Ramesh · 2020 [cited by applicant]
US 20200304393A1 · Gama et al. · 2020 [cited by applicant]
JP 6635265B2 · 2020 [cited by applicant]
JP 2020517030A · 2020 [cited by applicant]
KR 1020190085444A · 2019 [cited by applicant]
KR 1020190088643A · 2019 [cited by applicant]
KR 1020200044103A · 2020 [cited by applicant]
Korean Office Action issued on Sep. 28, 2024, in counterpart Korean Patent Application No. 10-2020-0075682 (2 pages in English, 5 pages in Korean). [cited by applicant]