IP Library Granted Patent US 12682236
Granted Patent B2
US 12682236 · App. 18/270,416 · Granted Jul 14, 2026

Method and system for lightweighting artificial neural network model, and non-transitory computer-readable recording medium

Inventor: Jin Woo Park (Seoul, KR)
Assignee: MAY-I INC.
G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682236
App. No.
18/270,416
Granted
Jul 14, 2026
Kind
B2
Abstract

A method for light-weighting an artificial neural network model, the method comprising is provided. The method includes the steps of: learning, on the basis of a channel length of a 1×1 pruning unit applied in a channel direction to each of a plurality of kernels included in an artificial neural network model, a pruning factor for each of the pruning units and a weight for each of the pruning units; and determining, on the basis of at least one of the learned pruning factors and weights, which of the pruning units is to be removed from the artificial neural network model, wherein the channel length of the pruning unit is shorter than a channel length of at least a part of the plurality of kernels.

Claims (19)

1 . A method for a device connecting to and communicating with an

artificial neural network model light-weighting system, wherein the device includes an application for a user to receive distribution of a light-weighted artificial neural network model services, the method comprising the steps of:

downloading, from the artificial neural network model light-weighting system, the application;

learning, by the application, on the basis of a channel length of a 1×1 pruning unit applied in a channel direction to each of a plurality of kernels included in an artificial neural network model, a pruning factor for each of the pruning units and a weight for each of the pruning units; and

determining, on the basis of at least one of the learned pruning factors and weights, which of the pruning units is to be removed from the artificial neural network model, wherein the channel length of the pruning unit is shorter than a channel length of at least a part of the plurality of kernels, wherein

artificial neural network models of various structures are distributed to the device with only one-time learning and ensures

that artificial neural network models suitable for various edge computing environments.

2 . The method of claim 1 , wherein the channel length of the pruning unit is commonly applied to a plurality of convolution layers included in the artificial neural network model.

3 . The method of claim 2 , wherein the plurality of convolution layers include a first convolution layer and a second convolution layer having a different channel length from the first convolution layer, and

wherein in the determining step, the pruning unit to be removed is shared in the first convolution layer and the second convolution layer.

4 . The method of claim 1 , wherein the channel length of the pruning unit is a power of 2.

5 . A non-transitory computer-readable recording medium having stored thereon a computer program for executing the method of claim 1 .

6 . An artificial neural network model light-weighting system where a device connects to and communicating with

artificial neural network model light-weighting system, wherein the device includes an application for a user to receive distribution of a light-weighted artificial neural network model services, the system comprising:

storing the application for downloading to the device;

a pruning factor learning unit configured to learn using the application, on the basis of a channel length of a 1×1 pruning unit applied in a channel direction to each of a plurality of kernels, a pruning factor for each of the pruning units and a weight for each of the pruning units; and

an artificial neural network model management unit configured to determine, on the basis of at least one of the learned pruning factors and weights, which of the pruning units is to be removed from an artificial neural network model, wherein the channel length of the pruning unit is shorter than a channel length of at least a part of the plurality of kernels, wherein

artificial neural network models of various structures are distributed to the device with only one-time learning and ensures that artificial

neural network models suitable for various edge computing environments.