IP Library Granted Patent US 12675697
Granted Patent B2
US 12675697 · App. 17/886,499 · Granted Jul 7, 2026

Computer-readable recording medium having stored therein machine learning program, method for machine learning, and information processing apparatus

Inventor: Yasufumi Sakai (Fuchu, JP)
Assignee: Fujitsu Limited
G06N3/082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12675697
App. No.
17/886,499
Granted
Jul 7, 2026
Kind
B2
Abstract

A computer-readable recording medium has stored therein a machine learning program for causing a computer to execute a process including: selecting a reduction ratio of each element of a plurality of layers in a trained model of a neural network including the plurality of layers; and adjusting, when the neural network includes a calculating process that outputs a tensor serving as a result of a given calculation on a tensor from a first layer and one or more tensors of one or more second layers preceding the first layer, a first reduction ratio and one or more second reduction ratios based on one or more elements to be reduced in the first layer at the first reduction ratio and one or more elements to be reduced in each of the one or more second layers at the one or more second reduction ratios.

Claims (38)

1 . A non-transitory computer-readable recording medium having stored therein a machine learning program comprising instructions which, when the machine learning program is executed by a computer, cause the computer to execute a process comprising:

calculating a plurality of thresholds of errors in tensors between before and after reduction one for each element of a plurality of layers in a trained model of a neural network including the plurality of layers, the neural network including a calculating process that outputs a tensor serving as a result of a given calculation on a tensor from a first layer and one or more tensors of one or more second layers preceding the first layer;

selecting, as a pruning rate, one from a plurality of pruning rate candidates each presenting a rate of elements to be pruned from one or more elements in each of the plurality of layers based on the plurality of thresholds and errors in tensors between before and after reduction in cases where the elements are pruned by each of the plurality of pruning rate candidates in each of the plurality of layers;

adjusting a first pruning rate and one or more second pruning rates by determining elements to be pruned in the first layer and the one or more second layers based on one or more elements to be reduced when the first layer is pruned at a first pruning rate selected as the pruning rate for the first layer and one or more elements to be reduced when the one or more second layers are pruned at the one or more second pruning rates selected as the pruning rates for the one or more second layers;

retraining a pruned model to obtain a retrained pruned model, the pruned model being obtained by pruning each element of the plurality of layers in the trained model according to the pruning rates selected or adjusted; and

determining the pruning rate to be applied to each of the plurality of layers based on inference accuracy of the trained model not being subjected to pruning and inference accuracy of the retrained pruned model.

2 . The non-transitory computer-readable recording medium according to claim 1 , wherein the adjusting includes adjusting the first pruning rate and the one or more second pruning rates such that a number of elements of the tensor from the first layer matches a number of elements of each of the one or more tensors from the one or more second layers.

3 . The non-transitory computer-readable recording medium according to claim 1 , wherein the adjusting includes adjusting the first pruning rate and the one or more second pruning rates such that one or more elements to be reduced in all of the first layer and the one or more second layers is regarded as one or more elements to be reduced when each of the first layer and the one or more second layers are pruned and that one or more elements not to be reduced in at least one of the first layer and the one or more second layers is excluded from one or more elements to be reduced when each of the first layer and the one or more second layers are pruned.

4 . The non-transitory computer-readable recording medium according to claim 1 , wherein the element is one selected from a group consisting of a channel, a weight, and a node.

5 . A computer-implemented method for machine learning comprising:

calculating a plurality of thresholds of errors in tensors between before and after reduction one for each element of a plurality of layers in a trained model of a neural network including the plurality of layers, the neural network including a calculating process that outputs a tensor serving as a result of a given calculation on a tensor from a first layer and one or more tensors of one or more second layers preceding the first layer;

selecting, as a pruning rate, one from a plurality of pruning rate candidates each presenting a rate of elements to be pruned from one or more elements in each of the plurality of layers based on the plurality of thresholds and errors in tensors between before and after reduction in cases where the elements are pruned by each of the plurality of pruning rate candidates in each of the plurality of layers;

adjusting a first pruning rate and one or more second pruning rates by determining elements to be pruned in the first layer and the one or more second layers based on one or more elements to be reduced when the first layer is pruned at a first pruning rate selected as the pruning rate for the first layer and one or more elements to be reduced when the one or more second layers are pruned at the one or more second pruning rates selected as the pruning rates for the one or more second layers;

retraining a pruned model to obtain a retrained pruned model, the pruned model being obtained by pruning each element of the plurality of layers in the trained model according to the pruning rates selected or adjusted; and

determining the pruning rate to be applied to each of the plurality of layers based on inference accuracy of the trained model not being subjected to pruning and inference accuracy of the retrained pruned model.

6 . The computer-implemented method according to claim 5 , wherein the adjusting includes adjusting the first pruning rate and the one or more second pruning rates such that a number of elements of the tensor from the first layer matches a number of elements of each of the tensors from the one or more second layers.

7 . The computer-implemented method according to claim 5 , wherein the adjusting includes adjusting the first pruning rate and the one or more second pruning rates such that one or more elements to be reduced in all of the first layer and the one or more second layers is regarded as one or more elements to be reduced when each of the first layer and the one or more second layers are pruned and that one or more elements not to be reduced in at least one of the first layer and the one or more second layers is excluded from one or more elements to be reduced when each of the first layer and the one or more second layers are pruned.

8 . The computer-implemented method according to claim 5 , wherein the element is one selected from a group consisting of a channel, a weight, and a node.

9 . An information processing apparatus comprising:

a memory; and

a processor coupled to the memory, the processor being configured to execute a process comprising:

calculating a plurality of thresholds of errors in tensors between before and after reduction one for each element of a plurality of layers in a trained model of a neural network including the plurality of layers, the neural network including a calculating process that outputs a tensor serving as a result of a given calculation on a tensor from a first layer and one or more tensors of one or more second layers preceding the first layer;

selecting, as a pruning rate, one from a plurality of pruning rate candidates each presenting a rate of elements to be pruned from one or more elements in each of the plurality of layers based on the plurality of thresholds and errors in tensors between before and after reduction in cases where the elements are pruned by each of the plurality of pruning rate candidates in each of the plurality of layers;

adjusting a first pruning rate and one or more second pruning rates by determining elements to be pruned in the first layer and the one or more second layers based on one or more elements to be reduced when the first layer is pruned at a first pruning rate selected as the pruning rate for the first layer and one or more elements to be reduced when the one or more second layers are pruned at the one or more second pruning rates selected as the pruning rates for the one or more second layers;

retraining a pruned model to obtain a retrained pruned model, the pruned model being obtained by pruning each element of the plurality of layers in the trained model according to the pruning rates selected or adjusted; and

determining the pruning rate to be applied to each of the plurality of layers based on inference accuracy of the trained model not being subjected to pruning and inference accuracy of the retrained pruned model.

10 . The information processing apparatus according to claim 9 , wherein the adjusting includes adjusting the first pruning rate and the one or more second pruning rates such that a number of elements of the tensor from the first layer matches a number of elements of each of the tensors from the one or more second layers.

11 . The information processing apparatus according to claim 9 , wherein the adjusting includes adjusting the first pruning rate and the one or more second pruning rates such that one or more elements to be reduced in all of the first layer and the one or more second layers is regarded as one or more elements to be reduced when each of the first layer and the one or more second layers are pruned and that one or more elements not to be reduced in at least one of the first layer and the one or more second layers is excluded from one or more elements to be reduced when each of the first layer and the one or more second layers are pruned.

12 . The information processing apparatus according to claim 9 , wherein the element is one selected from a group consisting of a channel, a weight, and a node.

13 . The non-transitory computer-readable recording medium according to claim 1 , wherein

the selecting and the adjusting are repeated while changing the plurality of thresholds according to inference accuracy of the trained model not being subjected to pruning and inference accuracy of a pruned model after machine learning, and

the process further comprises determining the pruning rate to be applied to each of the plurality of layers based on a result of the repeating.

14 . The computer-implemented method according to claim 6 , wherein

the selecting and the adjusting are repeated while changing the plurality of thresholds according to inference accuracy of the trained model not being subjected to pruning and inference accuracy of a pruned model after machine learning, and

the process further comprises determining the pruning rate to be applied to each of the plurality of layers based on a result of the repeating.

15 . The information processing apparatus according to claim 9 , wherein

the selecting and the adjusting are repeated while changing the plurality of thresholds according to inference accuracy of the trained model not being subjected to pruning and inference accuracy of a pruned model after machine learning, and

the process further comprises determining the pruning rate to be applied to each of the plurality of layers based on a result of the repeating.