IP Library › Granted Patent US 12,271,818
Granted Patent B1
US 12,271,818 · App. 17/330,048 · Granted Apr 8, 2025

Implementation-tuned architecture for neural network processing in a learned transform domain

Inventors: Kristof Denolf (Longmont, CO); Alireza Khodamoradi (San Diego, CA); Kornelis A. Vissers (Sunnyvale, CA)
Assignee: XILINX, INC.
G06N3/082G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,271,818
App. No.
17/330,048
Granted
Apr 8, 2025
Kind
B1
Abstract

Embodiments herein describe a learnable transform block disposed before, or in between, the neural network layers to transform received data into a more computational-friendly domain while preserving discriminative features required for the neural network to generate accurate results. In one embodiment, during a training phase, an AI system learns parameters for the transform block that are then used during the inference phase to transform received data into the computational-friendly domain that has a reduced size input. The transformed data may require less compute resources or less memory usage to process by the underlying hardware device that hosts the neural network.

Claims (49)

1. A method, comprising:

receiving training data at a transform block;

transforming the training data using the transform block to generate transformed data, wherein the transformed data requires at least one of less compute resources or less memory to process by a hardware device hosting a neural network;

inputting the transformed data to a layer in the neural network; and

learning parameters for the transform block during a training phase of the neural network, wherein adjusting the parameters for the transform block adjusts an amount of compute resources or memory used by the hardware device when processing the transformed data.

2. The method of claim 1 , wherein learning the parameters for the transform block is based on a multi-objective cost function that balances a cost of implementing the neural network on the hardware device with an accuracy of the neural network.

3. The method of claim 2 , wherein the multi-objective cost function maximizes the accuracy of the neural network while minimizing the cost of implementing the neural network on the hardware device.

4. The method of claim 1 , wherein learning the parameters for the transform block comprises:

identifying discriminative features in the training data that have, according to a threshold, a substantial impact on an accuracy of the neural network.

5. The method of claim 4 , wherein adjusting the parameters is performed to keep the discriminative features in the training data.

6. The method of claim 1 , further comprising:

providing a second transform block, wherein the second transform block is disposed between two layers in the neural network;

transforming the training data using the second transform block to generate second transformed data; and

learning second parameters for the second transform block during the training phase of the neural network.

7. The method of claim 6 , further comprising:

upon determining that an inherent computational cost of at least one of the transform block or the second transform block outweigh its computational savings from transforming the training data, removing the at least one of the transform block of the second transform block before entering an inference phase.

8. A computing system, comprising:

a processor;

memory storing an application which, when executed by the processor, performs an operation, the operation comprising:

receiving training data at a transform block;

transforming the training data using the transform block to generate transformed data, wherein the transformed data requires at least one of less compute resources or less memory to process by a hardware device hosting a neural network;

inputting the transformed data to a layer in the neural network; and

learning parameters for the transform block during a training phase of the neural network, wherein adjusting the parameters for the transform block adjusts an amount of compute resources or memory used by the hardware device when processing the transformed data.

9. The computing system of claim 8 , wherein learning the parameters for the transform block is based on a multi-objective cost function that balances a cost of implementing the neural network on the hardware device with an accuracy of the neural network.

10. The computing system of claim 9 , wherein the multi-objective cost function maximizes the accuracy of the neural network while minimizing the cost of implementing the neural network on the hardware device.

11. The computing system of claim 8 , wherein learning the parameters for the transform block comprises:

identifying discriminative features in the training data that have, according to a threshold, a substantial impact on an accuracy of the neural network.

12. The computing system of claim 11 , wherein adjusting the parameters is performed to keep the discriminative features in the training data.

13. The computing system of claim 8 , wherein the operation further comprises:

providing a second transform block, wherein the second transform block is disposed between two layers in the neural network;

transforming the training data using the second transform block to generate second transformed data; and

learning second parameters for the second transform block during the training phase of the neural network.

14. The computing system of claim 13 , wherein the operation further comprises:

upon determining that an inherent computational cost of at least one of the transform block or the second transform block outweigh its computational savings from transforming the training data, removing the at least one of the transform block of the second transform block before entering an inference phase.

15. A non-transitory computer readable medium having program instructions embodied therewith, the program instructions executable by a processor to perform an operation, the operation comprising:

receiving training data at a transform block;

transforming the training data using the transform block to generate transformed data, wherein the transformed data requires at least one of less compute resources or less memory to process by a hardware device hosting a neural network;

inputting the transformed data to a layer in the neural network; and

learning parameters for the transform block during a training phase of the neural network, wherein adjusting the parameters for the transform block adjusts an amount of compute resources or memory used by the hardware device when processing the transformed data.

16. The non-transitory computer readable medium of claim 15 , wherein learning the parameters for the transform block is based on a multi-objective cost function that balances a cost of implementing the neural network on the hardware device with an accuracy of the neural network.

17. The non-transitory computer readable medium of claim 15 , wherein learning the parameters for the transform block comprises:

identifying discriminative features in the training data that have, according to a threshold, a substantial impact on an accuracy of the neural network.

18. The non-transitory computer readable medium of claim 17 , wherein adjusting the parameters is performed to keep the discriminative features in the training data.

19. The non-transitory computer readable medium of claim 15 , wherein the operation further comprises:

providing a second transform block, wherein the second transform block is disposed between two layers in the neural network;

transforming the training data using the second transform block to generate second transformed data; and

learning second parameters for the second transform block during the training phase of the neural network.

20. The non-transitory computer readable medium of claim 19 , wherein the operation further comprises:

upon determining that an inherent computational cost of at least one of the transform block or the second transform block outweigh its computational savings from transforming the training data, removing the at least one of the transform block of the second transform block before entering an inference phase.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: DENOLF, KRISTOF; KHODAMORADI, ALIREZA; VISSERS, KORNELIS A.
To: XILINX, INC.
Reel/Frame 059101/0194 →
References Cited (14)
US 10853726B2 · Zoph · 2020 [cited by examiner]
US 20200104715A1 · Denolf et al. · 2020 [cited by applicant]
CN 106796668B · 2019 [cited by examiner]
CN 112313666A · 2021 [cited by examiner]
CN 112347550A · 2021 [cited by examiner]
CN 115238883A · 2022 [cited by examiner]
CN 111488976B · 2023 [cited by examiner]
CN 111488963B · 2023 [cited by examiner]
WO WO2021022903A1 · 2021 [cited by examiner]
WO WO2021036892A1 · 2021 [cited by examiner]
Fujieda, S. et al. “Wavelet Convolutional Neural Networks.” ArXiv abs/1805.08620 (2018), 10 pages. [cited by applicant]
M. X. Bastidas Rodriguez et al., “Deep Adaptive Wavelet Network,” 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), 2020, pp. 3111-3119. [cited by applicant]
Xu, Kai et al. “Learning in the Frequency Domain.” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020): 1740-1749. [cited by applicant]
Hou, Yunzhong et al. “Learning to Structure an Image With Few Colors.” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020): 10116-10125. [cited by applicant]