IP Library › Granted Patent US 12,474,914
Granted Patent B1
US 12,474,914 · App. 18/067,203 · Granted Nov 18, 2025

Concurrent machine learning model prediction on edge devices

Inventors: Chao Zhou (Fremont, CA); Ravish Hastantram (Fremont, CA); Peter Arnold Zientara (New Lenox, IL); Patrick Sisterhen (Georgetown, TX); Surya Kari (Newark, CA)
Assignee: Amazon Technologies, Inc.
G06F8/65G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,474,914
App. No.
18/067,203
Granted
Nov 18, 2025
Kind
B1
Abstract

In some examples, a method includes receiving a request to load a first machine learning model on an edge device, wherein the request includes an external identifier of the first machine learning model. The method further includes loading at least one instance of the first machine learning model onto the edge device, wherein the first machine learning model is loaded into a model pool having at least a second machine learning model using the same external identifier as the first machine learning model.

Claims (36)

1 . A computer-implemented method comprising:

receiving a request to load a first machine learning (ML) model into a ML model pool of an edge device, wherein the request includes an external identifier of the first ML model that identifies the ML model pool;

in response to the request, the edge device executing edge compute code using one or more processors to load at least one instance of the first ML model into the ML model pool of the edge device, wherein the ML model pool has at least a second ML model using the external identifier and each of the first and second ML models have different internal aliases; and

a runtime orchestrator executed by the one or more processors of the edge device handling translation of prediction requests to a model layer to underlying runtime application programming interfaces (APIs).

2 . The computer-implemented method of claim 1 , wherein the first ML model is an update of the second ML model wherein only weights of the ML models are different.

3 . The computer-implemented method of claim 1 , wherein the request includes one or more of: an identification of a model handle, an identification of a model alias, an indication of a location of a model and/or its runtime, an indication of input and/or output buffer sizes, an indication of a desired execution environment, and an indication of a number of instances to load.

4 . A computer-implemented method comprising:

receiving a request to load a first machine learning (ML) model on an edge device, wherein the request includes an external identifier of the first ML model;

in response to the request, the edge device executing edge compute code using one or more processors to load at least one instance of the first ML model onto the edge device, wherein the first ML model is loaded into a ML model pool having at least a second ML model using the external identifier and each of the first and second ML models have different internal aliases; and

a runtime orchestrator executed by the one or more processors of the edge device handling translation of prediction requests to a model layer to underlying runtime application programming interfaces (APIs).

5 . The computer-implemented method of claim 4 , wherein the first ML model is an update of the second ML model wherein only weights of the ML models are different.

6 . The computer-implemented method of claim 4 , wherein the first ML model has a different architecture than the second ML model.

7 . The computer-implemented method of claim 4 , further comprising:

receiving data from at least one application to use in a prediction using the first ML model according to a prediction request, wherein the prediction request identifies the external identifier and not the internal alias of the first ML model; and

performing inference using the first ML model.

8 . The computer-implemented method of claim 4 , further comprising:

receiving data from at least one application to use in a prediction using the first ML model according to a prediction request, wherein the prediction request identifies the external identifier and not the internal alias of the first ML model;

pre-processing the data to fit the first ML model; and

performing inference using the first ML model.

9 . The computer-implemented method of claim 4 , further comprising unloading the second ML model upon loading the first ML model.

10 . The computer-implemented method of claim 4 , wherein the request includes one or more of: an identification of a model handle, an identification of a model alias, an indication of a location of a model and/or its runtime, an indication of input and/or output buffer sizes, an indication of a desired execution environment, and an indication of a number of instances to load.

11 . The computer-implemented method of claim 4 , wherein return status codes for the request include at least one of load is successful, an unknown error has occurred, an internal error has occurred, the first ML model does not exist, the first ML model already exists, memory is not available to load the first ML model, or the first ML model is not compiled for the machine.

12 . The computer-implemented method of claim 4 , further comprising batching tensor data to be used by the first ML model during inferencing.

13 . The computer-implemented method of claim 4 , further comprising performing inference using the first ML model and the second ML model.

14 . The computer-implemented method of claim 4 , wherein the request to load the first ML model on the edge device is received by a provider network to invoke loading on the edge device.

15 . A system comprising:

a first one or more electronic devices to implement an edge device service in a multi-tenant provider network; and

a second one or more electronic devices to implement edge compute software on an edge device, the edge compute software including instructions that upon execution by one or more processors cause the edge device to:

receive a request from the edge device service to load a first machine learning (ML) model on the edge device, wherein the request includes an external identifier of the first ML model;

in response to the request, execute the edge compute software to load at least one instance of the first ML model onto the edge device, wherein the first ML model is loaded into a ML model pool having at least a second ML model using the external identifier and each of the first and second ML models have different internal aliases; and

a runtime orchestrator executed by the one or more processors of the edge device handling translation of prediction requests to a model layer to underlying runtime application programming interfaces (APIs).

16 . The system of claim 15 , wherein the request includes one or more of: an identification of a model handle, an identification of a model alias, an indication of a location of a model and/or its runtime, an indication of input and/or output buffer sizes, an indication of a desired execution environment, and an indication of a number of instances to load.

17 . The system of claim 15 , wherein the first ML model is an update of the second ML model wherein only weights of the ML models are different.

18 . The system of claim 15 , wherein the first ML model has a different architecture than the second ML model.

19 . The system of claim 15 , wherein return status codes for the request include at least one of load is successful, an unknown error has occurred, an internal error has occurred, the first ML model does not exist, the first ML model already exists, memory is not available to load the first ML model, or the first ML model is not compiled for the machine.

20 . The system of claim 15 , wherein the edge compute software includes further instructions that upon execution by the one or more processors further cause the edge device to perform inference using the first ML model and the second ML model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: ZHOU, CHAO; HASTANTRAM, RAVISH; ZIENTARA, PETER ARNOLD; SISTERHEN, PATRICK; KARI, SURYA
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069097/0292 →
References Cited (55)
US 7480907B1 · Marolia et al. · 2009 [cited by applicant]
US 10963810B2 · Dirac · 2021 [cited by examiner]
US 11182691B1 · Zhang · 2021 [cited by examiner]
US 11301762B1 · Chen · 2022 [cited by examiner]
US 11516311B2 · Klein et al. · 2022 [cited by applicant]
US 20160078361A1 · Brueckner · 2016 [cited by examiner]
US 20160217388A1 · Okanohara et al. · 2016 [cited by applicant]
US 20180032915A1 · Nagaraju et al. · 2018 [cited by applicant]
US 20190087239A1 · Adibowo · 2019 [cited by examiner]
US 20190139441A1 · Akella et al. · 2019 [cited by applicant]
US 20190220783A1 · Nookula · 2019 [cited by examiner]
US 20190251469A1 · Wagstaff et al. · 2019 [cited by applicant]
US 20190370687A1 · Pezzillo · 2019 [cited by examiner]
US 20200349465A1 · Hechtman · 2020 [cited by applicant]
US 20210064941A1 · Khan et al. · 2021 [cited by applicant]
US 20210157312A1 · Cella et al. · 2021 [cited by applicant]
US 20210201526A1 · Moloney et al. · 2021 [cited by applicant]
US 20210232981A1 · Sundaresan · 2021 [cited by examiner]
US 20210312277A1 · Prabhudesai et al. · 2021 [cited by applicant]
US 20210334629A1 · Yuan et al. · 2021 [cited by applicant]
US 20220012012A1 · Langhammer et al. · 2022 [cited by applicant]
US 20220091837A1 · Chai · 2022 [cited by examiner]
US 20220092480A1 · Mahadik · 2022 [cited by examiner]
US 20220108035A1 · Mehta · 2022 [cited by examiner]
US 20220129785A1 · Vogeti · 2022 [cited by examiner]
US 20220129787A1 · Vogeti · 2022 [cited by examiner]
US 20220180178A1 · Tasinga · 2022 [cited by examiner]
US 20220343137A1 · Surendran et al. · 2022 [cited by applicant]
US 20220366302A1 · Gilad · 2022 [cited by examiner]
US 20220382539A1 · Gumashta · 2022 [cited by examiner]
US 20220382601A1 · Feldman · 2022 [cited by examiner]
US 20220405619A1 · Ramamurthy · 2022 [cited by examiner]
US 20230068386A1 · Akdeniz et al. · 2023 [cited by applicant]
US 20230370476A1 · Bakshi · 2023 [cited by examiner]
US 20230409876A1 · Agrawal et al. · 2023 [cited by applicant]
US 20240119003A1 · Tobkin · 2024 [cited by examiner]
US 20240135241A1 · Yang · 2024 [cited by examiner]
CN 114661455A · 2022 [cited by examiner]
Nikita Kotsehub, FLoX: Federated Learning with FaaS at the Edge, 2022, pp. 1-10. https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9973578 (Year: 2022). [cited by examiner]
Pierrick Pochelu, An efficient and flexible inference system for serving heterogeneous ensembles of deep neural networks, 2021, pp. 1-8. https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9671725 (Year: 2021). [cited by examiner]
English translation, Guim (CN 114661455 A), 2022, pp. 1-15. (Year: 2022). [cited by examiner]
Chen, Zhuo, et al. “An empirical study of latency in an emerging class of edge computing applications for wearable cognitive assistance.” Proceedings of the Second ACM/IEEE Symposium on Edge Computing. 2017. pp. 1-14 (Y… [cited by applicant]
Notice of Allowance, U.S. App. No. 18/067,171, Nov. 25, 2024, 11 pages. [cited by applicant]
Qolomany, Basheer, et al. “Leveraging machine learning and big data for smart buildings: A comprehensive survey.” IEEE access 7 (2019): pp. 90316-90356. (Year: 2019). [cited by applicant]
Wang, Jin, et al. “Fast adaptive task offloading in edge computing based on meta reinforcement learning.” IEEE Transactions on Parallel and Distributed Systems 32.1 (2020): pp. 242-253. (Year: 2020). [cited by applicant]
Chen, Jiasi, and Xukan Ran. “Deep learning with edge computing: A review.” Proceedings of the IEEE 107.8 (2019): pp. 1655-1674. (Year: 2019). [cited by applicant]
Duc, Thang Le, et al. “Machine learning methods for reliable resource provisioning in edge-cloud computing: A survey.” ACM Computing Surveys (CSUR) 52.5 (2019): pp. 1-39. (Year: 2019). [cited by applicant]
Li, En, Zhi Zhou, and Xu Chen. “Edge intelligence: On-demand deep learning model co-inference with device-edge synergy.” Proceedings of the 2018 workshop on mobile edge communications. 2018. pp. 31-36 (Year: 2018). [cited by applicant]
Li, He, Kaoru Ota, and Mianxiong Dong. “Learning IoT in edge: Deep learning for the Internet of Things with edge computing.” IEEE network 32.1 (2018): 96-101. (Year: 2018). [cited by applicant]
Li, Tian, et al. “Ease, ml: Towards multi-tenant resource sharing for machine learning workloads.” Proceedings of the VLDB Endowment 11.5 (2018): pp. 607-620. (Year: 2018). [cited by applicant]
Non-Final Office Action, U.S. App. No. 18/067,171, Jun. 10, 2024, 26 pages. [cited by applicant]
Notice of Allowance, U.S. App. No. 18/067,171, Aug. 21, 2024, 13 pages. [cited by applicant]
Xu, Dianlei, et al. “Edge intelligence: Empowering intelligence to the edge of network.” Proceedings of the IEEE 109.11 (2021): pp. 1778-1837. (Year: 2021). [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/067,231, Jun. 30, 2025, 27 pages. [cited by applicant]
Visengeriyeva, Larysa, Kammer Kammer, Isabel Bar, Alexander Kniesz, and Michael Plod. “ML-Ops.Org.” ML Ops: Machine Learning Operations, Dec. 2, 2020. https://ml-ops.org/content/three-levels-of-ml-software. (Year: 2020). [cited by applicant]