IP Library › Granted Patent US 12,488,249
Granted Patent B2
US 12,488,249 · App. 17/465,439 · Granted Dec 2, 2025

Secure, accurate and fast neural network inference by replacing at least one non-linear activation channel

Inventors: Qian Lou (Mountain View, CA); Yilin Shen (Mountain View, CA); Hongxia Jin (Mountain View, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06N3/082H04L67/34
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,249
App. No.
17/465,439
Granted
Dec 2, 2025
Kind
B2
Abstract

A method of a server device is provided. The method of a server device includes retrieving a prediction input and a prediction setting, replacing at least one non-linear activation channel in a neural network with at least one replacement channel based on the received prediction setting, generating a prediction based on the received prediction input based on the neural network with the at least one replacement channel, and outputting the generated prediction.

Claims (44)

1 . A method of a server device, the method comprising:

receiving a prediction input and a prediction setting comprising an accuracy threshold;

providing a neural network operating on the server device and comprising a plurality of channels, the neural network being configured to generate a prediction based on the received prediction input;

based on the received prediction setting, selecting at least one non-linear activation channel among the plurality of channels that uses a non-linear activation and has a contribution to prediction accuracy of the neural network below the accuracy threshold;

replacing the at least one non-linear activation channel with at least one replacement channel having at least one approximated linear polynomial that utilizes fewer resources of the server device than the non-linear activation;

generating a prediction based on the received prediction input using the neural network with the at least one replacement channel; and

outputting the generated prediction.

2 . The method of claim 1 , wherein the prediction setting comprises a prediction generation speed.

3 . The method of claim 1 , wherein the prediction setting comprises a prediction generation accuracy.

4 . The method of claim 1 , wherein the selecting the at least one non-linear activation channel comprises determining a replacement ratio based on the prediction setting.

5 . The method of claim 4 , wherein the selecting the at least one non-linear activation channel further comprises selecting a percentage of non-linear activation channels in a layer of the neural network based on the determined replacement ratio.

6 . The method of claim 5 , wherein the replacing the at least one non-linear activation channel further comprises determining a replacement option based on the prediction setting.

7 . The method of claim 6 , wherein the replacing the at least one non-linear activation channel further comprises replacing the percentage of non-linear activation channels in the layer of the neural network with a polynomial approximation function with a degree that is determined based on the replacement option.

8 . The method of claim 1 , wherein the prediction setting is generated based on a user input to an interface of a client device.

9 . The method of claim 1 , wherein the at least one non-linear activation channel comprises any one or any combination of a rectified linear unit (ReLU) activation layer, a sigmoid function activation layer, a tangent function activation layer, a softmax function activation layer, and an exponential linear unit (ELU) layer.

10 . A method of a client device, the method comprising:

obtaining a prediction input;

generating a prediction setting comprising an accuracy threshold, based on a user input to an interface of the client device;

inputting the prediction input to a neural network comprising a plurality of channels, the neural network being configured to generate a prediction based on the prediction input;

inputting the prediction setting to modify the neural network; and

receiving a prediction generated from the modified neural network,

wherein the neural network is modified by:

based on the prediction setting, selecting at least one non-linear activation channel among the plurality of channels that uses a non-linear activation and has contribution to prediction accuracy of the neural network below the accuracy threshold; and

replacing the at least one non-linear activation channel with at least one replacement channel, and

wherein the at least one replacement channel has at least one approximated linear polynomial that utilizes fewer resources than the non-linear activation.

11 . The method of claim 10 , wherein the prediction setting comprises a prediction generation speed.

12 . The method of claim 10 , wherein the prediction setting comprises a prediction generation accuracy.

13 . The method of claim 10 , wherein the at least one non-linear activation channel is selected by determining a replacement ratio based on the prediction setting and selecting a percentage of non-linear activation channels in a layer of the neural network based on the determined replacement ratio.

14 . The method of claim 13 , wherein the at least one non-linear activation channel is replaced by determining a replacement option based on the prediction setting and replacing the percentage of non-linear activation channels in the layer of the neural network with a polynomial approximation function with a degree that is determined based on the replacement option.

15 . The method of claim 10 , wherein the at least one non-linear activation channel comprises any one or any combination of a rectified linear unit (ReLU) activation layer, a sigmoid function activation layer, a tangent function activation layer, a softmax function activation layer, and an exponential linear unit (ELU) layer.

16 . A server device, comprising:

a planner;

a neural network comprising a plurality of channels, the neural network being configured to generate a prediction;

at least one processor; and

memory that stores instructions that, when executed, cause the at least one processor to:

receive a prediction input and a prediction setting comprising an accuracy threshold;

select, with the planner, at least one non-linear activation channel among the plurality of channels that uses a non-linear activation and has a contribution to prediction accuracy of the neural network below the accuracy threshold, based on the received prediction setting;

replace, with the planner, the at least one non-linear activation channel in with at least one replacement channel having at least one approximated linear polynomial that utilizes fewer resources of the server device than the non-linear activation;

generate a prediction based on the received prediction input using the neural network with the at least one replacement channel; and

output the generated prediction.

17 . The server device of claim 16 , wherein the prediction setting comprises a prediction generation speed.

18 . The server device of claim 16 , wherein the prediction setting comprises a prediction generation accuracy.

19 . The server device of claim 16 , wherein the at least one non-linear activation channel is selected by determining a replacement ratio based on the prediction setting and selecting a percentage of non-linear activation channels in a layer of the neural network based on the determined replacement ratio.

20 . The server device of claim 19 , wherein the at least one non-linear activation channel is replaced by determining a replacement option based on the prediction setting and replacing the percentage of non-linear activation channels in the layer of the neural network with a polynomial approximation function with a degree that is determined based on the replacement option.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 2, 2021
From: LOU, QIAN; SHEN, YILIN; JIN, HONGXIA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 057375/0302 →
Continuity (2)
Provisional Application 63093949 · Oct 20, 2020
Related Publication 20220121947A1 · Apr 21, 2022
References Cited (30)
US 10891537B2 · Wang · 2021 [cited by examiner]
US 11853401B1 · Nookula · 2023 [cited by examiner]
US 20150254555A1 · Williams, Jr. · 2015 [cited by examiner]
US 20160350648A1 · Gilad-Bachrach et al. · 2016 [cited by applicant]
US 20170228639A1 · Hara · 2017 [cited by examiner]
US 20170372201A1 · Gupta · 2017 [cited by examiner]
US 20180060278A1 · Lin · 2018 [cited by examiner]
US 20190272309A1 · Chung · 2019 [cited by examiner]
US 20200036510A1 · Gomez · 2020 [cited by examiner]
US 20210027166A1 · Gorokhov et al. · 2021 [cited by applicant]
US 20210042559A1 · Xu · 2021 [cited by examiner]
US 20210042624A1 · Matveev et al. · 2021 [cited by applicant]
US 20210056352A1 · Boustati et al. · 2021 [cited by applicant]
US 20210081789A1 · Chai et al. · 2021 [cited by applicant]
US 20210081806A1 · Chai et al. · 2021 [cited by applicant]
US 20210089922A1 · Lu et al. · 2021 [cited by applicant]
US 20210209247A1 · Mohassel · 2021 [cited by examiner]
US 20210397988A1 · Sarpatwar · 2021 [cited by examiner]
US 20230118109A1 · Mohassel · 2023 [cited by examiner]
KR 1020190110068A · 2019 [cited by applicant]
“Chen et al., Deep Neural Network Acceleration Based on Low-Rank Approximated Channel Pruning, 2020, IEEE” (Year: 2020). [cited by examiner]
“Nikolaev et al., Learning polynomial feedforward neural networks by genetic programming and backpropagation, IEEE, 3003” (Year: 2003). [cited by examiner]
https://arxiv.org/pdf/1711.08797 Practical Hash Functions for Similarity Estimation and Dimensionality Reduction* Søren Dahlgaard1,2, Mathias Bæk Tejs Knudsen1,2, and Mikkel Thorup (Year: 2017). [cited by examiner]
https://arxiv.org/pdf/1704.02685 Learning Important Features Through Propagating Activation Differences Avanti Shrikumar 1 Peyton Greenside 1 Anshul Kundaje 1 (Year: 2019). [cited by examiner]
Sze, Vivienne, et al. “Efficient processing of deep neural networks: A tutorial and survey.” Proceedings of the IEEE 105.12 (2017): 2295-2329. (Year: 2017). [cited by examiner]
Lou, Qian, and Lei Jiang. “She: A fast and accurate privacy-preserving deep neural network via leveled tfhe and logarithmic data representation.” arXiv preprint arXiv:1906.00148 (2019). (Year: 2019). [cited by examiner]
AboulAtta, Moustafa, Matthias Ossadnik, and Seyed-Ahmad Ahmadi. “Stabilizing inputs to approximated nonlinear functions for inference with homomorphic encryption in deep neural networks.” arXiv preprint arXiv:1902.01870… [cited by examiner]
Riazi, M. Sadegh, Bita Darvish Rouani, and Farinaz Koushanfar. “Deep learning on private data.” IEEE Security & Privacy 17.6 (2019): 54-63. (Year: 2019). [cited by examiner]
Wright, Sarah, and Tshilidzi Marwala. “Artificial intelligence techniques for steam generator modelling.” arXiv preprint arXiv:0811.1711 (2008). (Year: 2008). [cited by examiner]
International Search Report dated Jan. 11, 2022 issued by the International Searching Authority in counterpart International Application No. PCT/KR2021/014460 (PCT/ISA/210). [cited by applicant]