IP Library Granted Patent US 12711416
Granted Patent B2
US 12711416 · App. 17/314,244 · Granted Aug 18, 2026

Model modification and deployment

Inventors: Alessandro Montanari (Cambridge, GB); Fahim Kawsar (Cambridge, GB); Akhil Mathur (London, GB); Chulhong Min (Trumpington, GB)
Assignee: NOKIA TECHNOLOGIES OY
G06N20/00G06F18/2178
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711416
App. No.
17/314,244
Granted
Aug 18, 2026
Kind
B2
Abstract

An apparatus, method and computer program is described comprising: determining an initial performance of a first model, wherein determining the initial performance comprises deploying the first model at a first device; determining one or more operations for modifying the first model based on at least the initial performance of the first model and one or more user requirements; modifying the first model by performing the one or more operations; determining whether a performance of the modified first model satisfies the one or more user requirements, wherein the determining comprises deploying the modified first model at the first device; and in the event that the modified first model does not satisfy the one or more user requirements, further modifying the first model by performing one or more further operations until the performance of the modified first model satisfies the one or more user requirements, wherein the determining further one or more operations based on at least the performance of the modified first model and the one or more user requirements.

Claims (46)

1 . An apparatus comprising:

at least one processor; and

at least one memory comprising computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform:

determining an initial performance of a first model, wherein determining the initial performance comprises deploying the first model at a first device;

determining one or more operations for modifying the first model based on at least the initial performance of the first model and one or more user requirements, wherein the one or more operations comprise quantisation of the first model and/or causing concurrent execution of a plurality of models, including the first model, to improve use of memory at the first device;

modifying the first model by performing the one or more operations;

determining whether a performance of the modified first model satisfies the one or more user requirements, wherein the determining whether the performance of the modified first model satisfies one or more user requirements comprises deploying the modified first model at the first device; and

in the event that the modified first model does not satisfy the one or more user requirements, further modifying the first model by performing one or more further operations until the performance of the modified first model satisfies the one or more user requirements, wherein the determining one or more operations further comprises determining one or more operations based on at least the performance of the modified first model and the one or more user requirements.

2 . An apparatus as claimed in claim 1 , wherein determining whether the performance of the modified first model satisfies the one or more user requirements further comprises:

running a first number of inferences of the deployed first model at the first device;

collecting performance values of the modified first model; and

comparing the performance values with the one or more user requirements.

3 . An apparatus as claimed in claim 1 , wherein the one or more user requirements comprise one or more of accuracy requirements, latency requirements, memory consumption requirements, and/or energy consumption requirements.

4 . An apparatus as claimed in claim 1 , wherein the at least one memory and the computer program code are configured to, with the at least one processor, further cause the apparatus to perform:

retraining the modified first model.

5 . An apparatus as claimed claim 1 , wherein the one or more operations for modifying the first model comprises operations for optimising one or more of accuracy, latency, memory consumption, and/or energy consumption of the first model based on the one or more user requirements.

6 . An apparatus as claimed claim 1 , wherein the one or more operations for modifying the first model further comprises one or more of:

modification of a size of the first model;

replacing one or more first actions comprised in the execution of the first model at the first device with one or more equivalent second actions, wherein the one or more first actions are unsupported by the first device, and the one or more second actions are supported by the first device.

7 . An apparatus as claimed in claim 1 , wherein deploying the modified first model at the first device further comprises:

receiving, from the first device, requirements of the first device, wherein the requirements are based at least in part on hardware of the first device;

determining a compilation flow for deployment of the modified first model in the first device based, at least in part, on the received requirements;

generating a compiled first model binary based, at least in part, on the compilation flow; and

deploying the compiled first model binary at the first device.

8 . An apparatus as claimed in claim 7 , wherein generating the compiled first model binary further comprising performing, depending on the determined compilation flow, one of a pre-training quantization and post training quantization.

9 . An apparatus as claimed in claim 7 , wherein generating the compiled first model binary further comprises, depending on the determined compilation flow, performing one or more format conversion actions.

10 . An apparatus as claimed in claim 7 , wherein the compilation flow is determined based at least in part on an accelerator of the first device.

11 . An apparatus as claimed in claim 1 , wherein at least some of said means are remote from the first device.

12 . A method comprising:

determining an initial performance of a first model, wherein determining the initial performance comprises deploying the first model at a first device;

determining one or more operations for modifying the first model based on at least the initial performance of the first model and one or more user requirements, wherein the one or more operations comprise quantisation of the first model and/or causing concurrent execution of a plurality of models, including the first model, to improve use of memory at the first device;

modifying the first model by performing the one or more operations;

determining whether a performance of the modified first model satisfies the one or more user requirements, wherein the determining whether the performance of the modified first model satisfies one or more user requirements comprises deploying the modified first model at the first device; and

in the event that the modified first model does not satisfy the one or more user requirements, further modifying the first model by performing one or more further operations until the performance of the modified first model satisfies the one or more user requirements, wherein the determining one or more operations further comprises determining one or more operations based on at least the performance of the modified first model and the one or more user requirements.

13 . A method as claimed in claim 12 , wherein deploying the modified first model at the first device further comprises performing:

receiving, from the first device, requirements of the first device, wherein the requirements are based at least in part on hardware of the first device;

determining a compilation flow for deployment of the modified first model in the first device based, at least in part, on the received requirements;

generating a compiled first model binary based, at least in part, on the compilation flow;

and

deploying the compiled first model binary at the first device.

14 . A non-transitory computer-readable storage medium including a computer program comprising instructions for causing an apparatus to perform at least the following:

determining an initial performance of a first model, wherein determining the initial performance comprises deploying the first model at a first device;

determining one or more operations for modifying the first model based on at least the initial performance of the first model and one or more user requirements, wherein the one or more operations comprise quantisation of the first model and/or causing concurrent execution of a plurality of models, including the first model, to improve use of memory at the first device;

modifying the first model by performing the one or more operations;

determining whether a performance of the modified first model satisfies the one or more user requirements, wherein the determining whether the performance of the modified first model satisfies one or more user requirements comprises deploying the modified first model at the first device; and

in the event that the modified first model does not satisfy the one or more user requirements, further modifying the first model by performing one or more further operations until the performance of the modified first model satisfies the one or more user requirements, wherein the determining one or more operations further comprises determining one or more operations based on at least the performance of the modified first model and the one or more user requirements.