IP Library › Granted Patent US 12,260,236
Granted Patent B1
US 12,260,236 · App. 18/067,171 · Granted Mar 25, 2025

Machine learning model replacement on edge devices

Inventors: Chao Zhou (Fremont, CA); Maxwell Edward Chapman Nuyens (Redwood City, CA); Ravish Hastantram (Fremont, CA)
Assignee: Amazon Technologies, Inc.
G06F9/45508
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,236
App. No.
18/067,171
Granted
Mar 25, 2025
Kind
B1
Abstract

Techniques for machine learning prediction using edge devices are described. In some examples a method of use includes receiving a request to load a second model onto an edge device while a first model has already been loaded on the edge device, wherein the second model and the first model share an external handle; loading at least one instance the second model into memory of the edge device; and after the second model has been loaded into memory of the edge device, directing a prediction request to the shared external handle to the second model instead of the first model.

Claims (38)

1. A computer-implemented method comprising:

receiving a request to load a second machine learning (ML) model onto an edge device while a first ML model has already been loaded into memory of the edge device, wherein the second ML model and the first ML model share an external handle but have different internal aliases;

in response to the request, the edge device executing edge compute code using one or more processors to load at least one instance of the second ML model into the memory of the edge device while the first ML model is still loaded into the memory of the edge device;

after the second ML model has been loaded into the memory of the edge device, receiving a prediction request at the shared external handle of the first ML model and the second ML model;

in response to the prediction request, the edge device executing the edge compute code using the one or more processors to perform a handle-to-alias translation to determine to direct the prediction request to the second ML model instead of the first ML model;

directing the prediction request to the second ML model instead of the first ML model; and

after directing the prediction request to the second ML model instead of the first ML model, unloading the first ML model from the memory of the edge device.

2. The computer-implemented method of claim 1 , wherein the request to load the second ML model onto the edge device while the first ML model has already been loaded into the memory of the edge device includes at least one of an identification of a model handle, an identification of a model alias, an indication of a location of a model and/or its runtime, an indication of input and/or output buffer sizes, or an indication of a desired execution environment.

3. The computer-implemented method of claim 1 , wherein the second ML model is a variant of the first ML model.

4. A computer-implemented method comprising:

receiving a request to load a second machine learning (ML) model onto an edge device while a first ML model has already been loaded into memory of the edge device, wherein the second ML model and the first ML model share an external handle but have different internal aliases;

in response to the request, the edge device executing edge compute code using one or more processors to load at least one instance of the second ML model into the memory of the edge device while the first ML model is still loaded into the memory of the edge device;

after the second ML model has been loaded into the memory of the edge device, receiving a prediction request at the shared external handle of the first ML model and the second ML model;

in response to the prediction request, the edge device executing the edge compute code using the one or more processors to perform a handle-to-alias translation to determine to direct the prediction request to the second ML model instead of the first ML model; and

directing the prediction request to the second ML model instead of the first ML model.

5. The computer-implemented method of claim 4 , wherein the second ML model is a variant of the first ML model.

6. The computer-implemented method of claim 4 , wherein routing is based on the aliases.

7. The computer-implemented method of claim 4 , wherein loading the at least one instance of the second ML model into the memory of the edge device at least includes loading a runtime.

8. The computer-implemented method of claim 4 , further comprising batching requests to the shared external handle, wherein directing the prediction request to the second ML model instead of the first ML model is a part of a batch.

9. The computer-implemented method of claim 4 , further comprising unloading the first ML model from the memory of the edge device.

10. The computer-implemented method of claim 9 , wherein the first ML model is unloaded in response to an unload request.

11. The computer-implemented method of claim 10 , wherein the unload request includes an identification of the shared external handle and a model alias.

12. The computer-implemented method of claim 9 , wherein the first ML model is unloaded to free the memory into which the first ML model had been loaded.

13. The computer-implemented method of claim 4 , wherein the request to load the second ML model onto the edge device while the first ML model has already been loaded into the memory of the edge device includes at least one of an identification of a model handle, an identification of a model alias, an indication of a location of a model and/or its runtime, an indication of input and/or output buffer sizes, or an indication of a desired execution environment.

14. The computer-implemented method of claim 4 , wherein loading the second ML model comprises downloading and verifying the second ML model.

15. A system comprising:

a first one or more electronic devices to implement an edge device service in a multi-tenant provider network; and

a second one or more electronic devices to implement edge compute software in the multi-tenant provider network, the edge compute software including instructions that upon execution by one or more processors cause the edge compute software to:

receive a request to load a second machine learning (ML) model onto the edge device while a first ML model has already been loaded into memory of the edge device, wherein the second ML model and the first ML model share an external handle but have different internal aliases;

in response to the request, the second one or more electronic devices executing the edge compute software using the one or more processors to load at least one instance of the second ML model into the memory of the edge device while the first ML model is still loaded into the memory of the edge device;

after the second ML model has been loaded into the memory of the edge device, receiving a prediction request at the shared external handle of the first ML model and the second ML model;

in response to the prediction request, the edge device executing the edge compute software using the one or more processors to perform a handle-to-alias translation to determine to direct the prediction request to the second ML model instead of the first ML model; and

direct the prediction request to the second ML model instead of the first ML model.

16. The system of claim 15 , wherein the second ML model is a variant of the first ML model.

17. The system of claim 15 , wherein routing is based on the aliases.

18. The system of claim 15 , wherein the first ML model is to be unloaded in response to an unload request.

19. The system of claim 18 , wherein the unload request includes an identification of the shared external handle and a model alias.

20. The system of claim 15 , wherein the request to load the second ML model onto the edge device while the first ML model has already been loaded into the memory of the edge device includes at least one of an identification of a model handle, an identification of a model alias, an indication of a location of a model and/or its runtime, an indication of input and/or output buffer sizes, or an indication of a desired execution environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2022
From: ZHOU, CHAO; NUYENS, MAXWELL EDWARD CHAPMAN; HASTANTRAM, RAVISH
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 062131/0762 →
References Cited (29)
US 7480907B1 · Marolia · 2009 [cited by examiner]
US 11301762B1 · Chen et al. · 2022 [cited by applicant]
US 11516311B2 · Klein · 2022 [cited by examiner]
US 20160078361A1 · Brueckner et al. · 2016 [cited by applicant]
US 20160217388A1 · Okanohara · 2016 [cited by examiner]
US 20180032915A1 · Nagaraju · 2018 [cited by examiner]
US 20190087239A1 · Adibowo · 2019 [cited by applicant]
US 20190251469A1 · Wagstaff · 2019 [cited by examiner]
US 20210064941A1 · Khan · 2021 [cited by examiner]
US 20210312277A1 · Prabhudesai · 2021 [cited by examiner]
US 20220092480A1 · Mahadik · 2022 [cited by applicant]
US 20220180178A1 · Tasinga et al. · 2022 [cited by applicant]
US 20220366302A1 · Gilad · 2022 [cited by examiner]
US 20220382601A1 · Feldman et al. · 2022 [cited by applicant]
US 20220405619A1 · Ramamurthy et al. · 2022 [cited by applicant]
US 20230409876A1 · Agrawal · 2023 [cited by examiner]
US 20240119003A1 · Tobkin et al. · 2024 [cited by applicant]
Duc, Thang Le, et al. “Machine learning methods for reliable resource provisioning in edge-cloud computing: A survey.” ACM Computing Surveys (CSUR) 52.5 (2019): pp. 1-39. (Year: 2019). [cited by examiner]
Xu, Dianlei, et al. “Edge intelligence: Empowering intelligence to the edge of network.” Proceedings of the IEEE 109.11 (2021): pp. 1778-1837. (Year: 2021). [cited by examiner]
Li, Tian, et al. “Ease. ml: Towards multi-tenant resource sharing for machine learning workloads.” Proceedings of the VLDB Endowment 11.5 (2018): pp. 607-620. (Year: 2018). [cited by examiner]
Chen, Jiasi, and Xukan Ran. “Deep learning with edge computing: A review.” Proceedings of the IEEE 107.8 (2019): pp. 1655-1674. (Year: 2019). [cited by examiner]
Li, He, Kaoru Ota, and Mianxiong Dong. “Learning IoT in edge: Deep learning for the Internet of Things with edge computing.” IEEE network 32.1 (2018): 96-101. (Year: 2018). [cited by examiner]
Li, En, Zhi Zhou, and Xu Chen. “Edge intelligence: On-demand deep learning model co-inference with device-edge synergy.” Proceedings of the 2018 workshop on mobile edge communications. 2018.pp. 31-36 (Year: 2018). [cited by examiner]
Chen, Zhuo, et al. “An empirical study of latency in an emerging class of edge computing applications for wearable cognitive assistance.” Proceedings of the Second ACM/IEEE Symposium on Edge Computing. 2017. pp. 1-14 (Y… [cited by examiner]
Qolomany, Basheer, et al. “Leveraging machine learning and big data for smart buildings: A comprehensive survey.” IEEE access 7 (2019): pp. 90316-90356. (Year: 2019). [cited by examiner]
Wang, Jin, et al. “Fast adaptive task offloading in edge computing based on meta reinforcement learning.” IEEE Transactions on Parallel and Distributed Systems 32.1 (2020): pp. 242-253. (Year: 2020). [cited by examiner]
Nikita Kotsehub, FLoX: Federated Learning with FaaS at the Edge, 2022, pp. 1-10. https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9973578 (Year: 2022). [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/067,203, Oct. 10, 2024, 24 pages. [cited by applicant]
Pierrick Pochelu, An efficient and flexible inference system for serving heterogeneous ensembles of deep neural networks, 2021, pp. 1-8. https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9671725 (Year: 2021). [cited by applicant]
Cited By (3)
US 12,645,455 US 12,657,024 US 12,669,996