IP Library Patent Application 17124238
Patent Application
App. No. 17/124,238

RESOURCE AWARE NEURAL NETWORK MODEL DYNAMIC UPDATING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/124,238
Abstract

Resources of an embedded system, such as RAM utilization and available processor cycles or bandwidth are monitored. Neural network models of varying size and computational load for given neural networks are utilized in conjunction with this resource monitoring. The neural network model used for a particular neural network is dynamically varied based on the resource monitoring. In one example, neural network models of varying precision are stored and the best model for the available RAM and processor cycles is loaded. In one example, neural network model weight values are quantized before being loaded for use, the level of quantization being based on the available RAM and processor cycles. This dynamic adaption of the neural network models allows other processes in the embedded system to operate normally and yet allows the neural network to operate at the maximum capability allowed for a given period.

Claims (53)

1 . A method of operating a device which includes operating neural networks, the device having a processor and RAM and executing a plurality of modules of varying functionality, including at least one neural network, the method comprising:

periodically determining RAM utilization and available processor cycles of the device;

selecting a neural network model for the at least one neural network based on the periodic determination of RAM utilization and available processor cycles; and

executing the selected neural network model as the at least one neural network.

2 . The method of claim 1 , wherein there are a plurality of neural networks executing on the device, and

wherein the selecting a neural network model and executing the selected neural network model are performed for each of the plurality of neural networks.

3 . The method of claim 1 , wherein there are a plurality of neural networks executing on the device, and

wherein the selecting a neural network model and executing the selected neural network model are performed for at least one neural network but less than all of the plurality of neural networks.

4 . The method of claim 1 , wherein there are a plurality of neural network models for the at least one neural network, the plurality of neural network models differing in precision of the weights, and

wherein the selecting a neural network model includes selecting one of the plurality of neural network models based on the precision of the neural network model.

5 . The method of claim 4 , wherein the precisions differ by bit sizes and floating point or integer.

6 . The method of claim 4 , wherein the neural network model weight values are quantized,

wherein the selecting a neural network model includes determining a level of quantization of the neural network model weight values, and

wherein both the selection of the precision and the level of quantization are based on the RAM utilization and available processor cycles.

7 . The method of claim 1 , wherein the neural network model weight values are quantized,

wherein the selecting a neural network model includes determining a level of quantization of the neural network model weight values, and

wherein the level of quantization is based on the RAM utilization and available processor cycles.

8 . A device comprising:

RAM;

a processor coupled to the RAM for executing programs; and

memory coupled to the processor for storing programs executed by the processor, the memory storing programs executed by the processor to perform the operations of:

executing a plurality of programs of varying functionality, including at least one neural network;

periodically determining RAM utilization and available processor cycles of the device;

selecting a neural network model for the at least one neural network based on the periodic determination of RAM utilization and available processor cycles; and

executing the selected neural network model as the at least one neural network.

9 . The device of claim 8 , wherein there are a plurality of neural networks executing on the device, and

wherein the selecting a neural network model and executing the selected neural network model are performed for each of the plurality of neural networks.

10 . The device of claim 8 , wherein there are a plurality of neural networks executing on the device, and

wherein the selecting a neural network model and executing the selected neural network model are performed for at least one neural network but less than all of the plurality of neural networks.

11 . The device of claim 8 , wherein there are a plurality of neural network models for the at least one neural network, the plurality of neural network models differing in precision of the weights,

wherein the selecting a neural network model includes selecting one of the plurality of neural network models based on the precision of the neural network model, and

wherein each of the plurality of neural network models is stored in the memory.

12 . The device of claim 11 , wherein the precisions differ by bit sizes and floating point or integer.

13 . The device of claim 11 , wherein the neural network model weight values are quantized,

wherein the selecting a neural network model includes determining a level of quantization of the neural network model weight values, and

wherein both the selection of the precision and the level of quantization are based on the RAM utilization and available processor cycles.

14 . The device of claim 8 , wherein the neural network model weight values are quantized,

wherein the selecting a neural network model includes determining a level of quantization of the neural network model weight values, and

wherein the level of quantization is based on the RAM utilization and available processor cycles.

15 . A non-transitory processor readable memory containing programs that when executed cause a processor to perform the following method of operating a device which includes operating neural networks, the device having a processor and RAM and executing a plurality of modules of varying functionality, including at least one neural network, the method comprising:

periodically determining RAM utilization and available processor cycles of the device;

selecting a neural network model for the at least one neural network based on the periodic determination of RAM utilization and available processor cycles; and

executing the selected neural network model as the at least one neural network.

16 . The non-transitory processor readable memory of claim 15 , wherein there are a plurality of neural networks executing on the device, and

wherein the selecting a neural network model and executing the selected neural network model are performed for each of the plurality of neural networks.

17 . The non-transitory processor readable memory of claim 15 , wherein there are a plurality of neural networks executing on the device, and

wherein the selecting a neural network model and executing the selected neural network model are performed for at least one neural network but less than all of the plurality of neural networks.

18 . The non-transitory processor readable memory of claim 15 , wherein there are a plurality of neural network models for the at least one neural network, the plurality of neural network models differing in precision of the weights, and

wherein the selecting a neural network model includes selecting one of the plurality of neural network models based on the precision of the neural network model.

19 . The non-transitory processor readable memory of claim 18 , wherein the precisions differ by bit sizes and floating point or integer.

20 . The non-transitory processor readable memory of claim 15 , wherein the neural network model weight values are quantized,

wherein the selecting a neural network model includes determining a level of quantization of the neural network model weight values, and

wherein the level of quantization is based on the RAM utilization and available processor cycles.

Assignments (4)
NUNC PRO TUNC ASSIGNMENT Recorded Nov 13, 2023
From: PLANTRONICS, INC.
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 065549/0065 →
RELEASE OF PATENT SECURITY INTERESTS Recorded Aug 30, 2022
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: PLANTRONICS, INC.; POLYCOM, INC.
Reel/Frame 061356/0366 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Oct 6, 2021
From: PLANTRONICS, INC.; POLYCOM, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 057723/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2020
From: YAN, YONG; BRYAN, DAVID A.
To: PLANTRONICS, INC.
Reel/Frame 054672/0788 →