IP Library › Granted Patent US 12,536,467
Granted Patent B2
US 12,536,467 · App. 17/471,816 · Granted Jan 27, 2026

Merging models on an edge server

Inventors: Ganesh Ananthanarayanan (Seattle, WA); Anand Padmanabha Iyer (Redmond, WA); Yuanchao Shu (Kirkland, WA); Nikolaos Karianakis (Sammamish, WA); Arthi Hema Padmanabhan (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06N20/00G06F18/214G06F18/217G06V20/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,467
App. No.
17/471,816
Granted
Jan 27, 2026
Kind
B2
Abstract

Systems and methods are provided for merging models for use in an edge server under the multi-access edge computing environment. In particular, a model merger selects a layer of a model based on a level of memory consumption in the edge server and determines sharable layers based on common properties of the selected layer. The model merger generates a merged model by generating a single instantiation of a layer that corresponds to the sharable layers. A model trainer trains the merged model based on training data for the respective models to attain a level of accuracy of data analytics above a predetermined threshold. The disclosed technology further refreshes the merged model upon observing a level of data drift that exceeds a predetermined threshold. The refreshing of the merged model includes detaching and/or splitting consolidated sharable layers of sub-models in the merged model. By merging models, the disclosed technology reduces memory footprints of models used in the edge server, rectifying memory scarcity issues in the edge server.

Claims (61)

1 . A computer-implemented method for merging at least a part of a plurality of models, the method comprising:

selecting, based on memory consumption of a first layer of a first model in a memory of an edge server, the first layer of the first model as a first sharable layer, wherein the plurality of models comprises the first model and a second model;

determining, based on layer properties of the second model, a second layer of the second model in the memory of the edge server as a second sharable layer, wherein the second layer of the second model indicates one or more same layer properties as the first layer of the first model;

generating a third model as a merged model of the first model and the second model, wherein the third model comprises a third layer as a common layer representing a single instantiation of both the first layer and the second layer; and

transmitting the third model to the edge server, thereby causing the edge server to replace a respective instantiation of the first model and the second model in the memory of the edge server with the third model, and further causing the edge server to execute the third model to perform data analytics on the edge server, wherein the third model consumes less memory on the edge server than a combination of both the first model and the second model.

2 . The computer-implemented method of claim 1 , wherein the first model and the second model correspond to video analytic models.

3 . The computer-implemented method of claim 1 , wherein each of the first layer and the second layer is based on properties including one or more of:

input size,

output size,

kernel size, or

stride length.

4 . The computer-implemented method of claim 1 , wherein the first model and the second model are distinct.

5 . The computer-implemented method of claim 1 , the method further comprising:

receiving data drift information associated with a first sub-model of the third model;

generating, based on the data drift information associated with the first sub-model of the third model, a fourth model, wherein the fourth model includes the first sub-model;

updating the third model by detaching the first sub-model from the third model; and

transmitting the third model and the fourth model.

6 . The computer-implemented method of claim 1 , the method further comprising:

training the third model based on training data associated with the first model and the second model.

7 . The computer-implemented method of claim 6 , wherein training the third model comprises determining that the third model meets an accuracy threshold for performing the data analytics.

8 . The computer-implemented method of claim 1 , wherein the first layer and the second layer are not located in corresponding locations within the first model and the second model, respectively.

9 . A system for merging a plurality of models for use in an edge server, the system comprising:

a processor; and

a memory storing computer-executable instructions that when executed by the processor cause the system to:

select, based on memory consumption of a first layer of a first model in a memory of the edge server, the first layer of the first model as a first sharable layer, wherein the plurality of models comprises the first model and a second model;

determine, based on layer properties of the second model, a second layer of the second model in the memory of the edge server as a second sharable layer, wherein the second layer of the second model indicates one or more same layer properties as the first layer of the first model;

generate a third model as a merged model for the first model and the second model, wherein the third model comprises a third layer as a common layer representing a single instantiation of both the first layer and the second layer; and

transmit the third model to the edge server, thereby causing the edge server to replace a respective instantiation of the first model and the second model in the memory of the edge server with the third model, and further causing the edge server to execute the third model to perform data analytics on the edge server, wherein the third model consumes less memory on the edge server than a combination of both the first model and the second model.

10 . The system of claim 9 , wherein the first model and the second model correspond to video analytic models.

11 . The system of claim 9 , wherein each of the first layer and the second layer is based on properties including one or more of:

input size,

output size,

kernel size, or

stride length.

12 . The system of claim 9 , the computer-executable instructions when executed further cause the system to:

receive data drift information associated with a first sub-model of the third model;

generate, based on the data drift information associated with the first sub-model of the third model, a fourth model, wherein the fourth model includes the first sub-model;

update the third model by detaching the first sub-model from the third model; and

transmit the third model and the fourth model.

13 . The system of claim 9 , the computer-executable instructions when executed further cause the system to:

train the third model based on training data associated with the first model and the second model.

14 . The system of claim 13 , wherein training the third model comprises determining that the third model meets an accuracy threshold for performing the data analytics.

15 . The system of claim 9 , wherein the first layer and the second layer are not located in corresponding locations within the first model and the second model, respectively.

16 . A computer-readable recording medium storing computer-executable instructions that when executed by a processor cause a system to:

select, based on memory consumption of a first layer of a first model of a plurality of models in a memory of an edge server, the first layer of the first model as a first sharable layer, wherein the plurality of models comprises the first model and a second model;

determine, based on layer properties of the second model, a second layer of the second model in the memory of the edge server as a second sharable layer, wherein the second layer of the second model indicates one or more same layer properties as the first layer of the first model;

generate a third model as a merged model of the first model and the second model, wherein the third model comprises a third layer as a common layer representing a single instantiation of both the first layer and the second layer; and

transmit the third model to the edge server, thereby causing the edge server to replace a respective instantiation of the first model and the second model in the memory of the edge server with the third model, and further causing the edge server to execute the third model to perform data analytics on the edge server, and wherein the third model consumes less memory on the edge server than a combination of both the first model and the second model.

17 . The computer-readable recording medium of claim 16 , wherein the first model and the second model correspond to video analytic models.

18 . The computer-readable recording medium of claim 16 , wherein each of the first layer and the second layer is based on properties including one or more of:

input size,

output size,

kernel size, or

stride length.

19 . The computer-readable recording medium of claim 16 , the computer-executable instructions when executed further cause the system to:

receive data drift information associated with a first sub-model of the third model;

generate, based on the data drift information associated with the first sub-model of the third model, a fourth model, wherein the fourth model includes the first sub-model;

update the third model by detaching the first sub-model from the third model; and

transmit the third model and the fourth model.

20 . The computer-readable recording medium of claim 16 , the computer-executable instructions when executed further cause the system to:

train the third model based on training data associated with the first model and the second model, wherein training the third model comprises determining that the third model meets an accuracy threshold for performing the data analytics.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2021
From: ANANTHANARAYANAN, GANESH; PADMANABHA IYER, ANAND; SHU, YUANCHAO; KARIANAKIS, NIKOLAOS; PADMANABHAN, ARTHI HEMA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 057446/0694 →
Continuity (2)
Provisional Application 63195118 · May 31, 2021
Related Publication 20220383188A1 · Dec 1, 2022
References Cited (63)
US 20170344881A1 · Okuno · 2017 [cited by examiner]
US 20180232663A1 · Ross et al. · 2018 [cited by applicant]
US 20190130261A1 · Rallapalli et al. · 2019 [cited by applicant]
US 20190260827A1 · Tajima · 2019 [cited by examiner]
US 20210110045A1 · Buesser et al. · 2021 [cited by applicant]
US 20220024032A1 · Singh · 2022 [cited by examiner]
US 20220374635A1 · Xiong · 2022 [cited by examiner]
Yu, et al., “Semantic drift compensation for class-incremental learning”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 14, 2020, pp. 6982-6991. [cited by applicant]
Zhang, et al., “Live Video Analytics at Scale with Approximation and Delay-Tolerance”, In Proceedings of the14th USENIX Symposium on Networked Systems Design and Implementation, Mar. 27, 2017, pp. 377-392. [cited by applicant]
Zhang, et al., “The Design and Implementation of a Wireless Video Surveillance System”, In Proceedings of the 21st Annual International Conference on Mobile Computing and Networking, Sep. 7, 2015, pp. 426-438. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/027742”, Mailed Date: Aug. 17, 2022, 11 Pages. [cited by applicant]
“AWS Outposts”, Retrieved from: https://web.archive.org/web/20201208190445/https://aws.amazon.com/outposts/, Dec. 8, 2020, 9 Pages. [cited by applicant]
“Azure Stack Edge”, Retrieved from: https://web.archive.org/web/20200919135652/https:/azure.microsoft.com/en-in/products/azure-stack/edge/, Sep. 19, 2020, 17 Pages. [cited by applicant]
“Jetson Nano”, Retrieved from: https://web.archive.org/web/20210215213314/https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-nano/product-development/, Feb. 15, 2021, 4 Pages. [cited by applicant]
“Nvidia Multi-Process Service”, Retrieved from: https://docs.nvidia.com/deploy/pdfCUDA_Multi_Process_Service_Overview.pdf, Jun. 2020, 28 Pages. [cited by applicant]
“Nvidia TensorRT”, Retrieved from: https://developer.nvidia.com/tensorrt, Retrieved Date: May 12, 2021, 7 Pages. [cited by applicant]
“PyTorch”, Retrieved from: https://web.archive.org/web/20190101155119/https://pytorch.org/, Jan. 1, 2019, 3 Pages. [cited by applicant]
Ananthanarayan, “Real-Time Video Analytics: The Killer App for Edge Computing”, In Journal of Computer, vol. 50, Issue: 10, Oct. 2017, pp. 58-67. [cited by applicant]
Ananthanarayanan, et al., “Demo: Video Analytics—Killer App for Edge Computing”, In Proceedings of the 17th Annual International Conference on Mobile Systems, Applications, and Services, Jun. 12, 2019, pp. 695-696. [cited by applicant]
Ananthanarayanan, et al., “Traffic Video Analytics—Case Study Report”, In Technical report of MSR-TR-1970-3, Dec. 2019, 2 Pages. [cited by applicant]
Bai, et al., “PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications”, In Proceedings of 14th USENIX Symposium on Operating Systems Design and Implementation, Nov. 4, 2020, pp. 499-514. [cited by applicant]
Bhardwaj, et al., “Ekya: Continuous Learning of Video Analytics Models on Edge Compute Servers”, In Repository of arXiv:2012.10557v1, Dec. 19, 2020, 15 Pages. [cited by applicant]
Cai, et al., “Learning Complexity-Aware Cascades for Deep Pedestrian Detection”, In Proceedings of IEEE International Conference on Computer Vision, Dec. 7, 2015, pp. 3361-3369. [cited by applicant]
Canel, et al., “Scaling Video Analytics on Constrained Edge Nodes”, In Proceedings of the 2nd SysML Conference, Mar. 31, 2019, 12 Pages. [cited by applicant]
Caruana, Rich, “Multitask learning”, In Journal of Machine learning vol. 28, No. 1, Jul. 1, 1997, pp. 41-75. [cited by applicant]
Du, et al., “Server-Driven Video Streaming for Deep Learning Inference”, In Proceedings of the Annual conference of the ACM Special Interest Group on Data Communication on the applications, technologies, architectures, … [cited by applicant]
Emmons, et al., “Cracking Open the DNN Black-Box: Video Analytics with DNNs across the Camera-Cloud Boundary”, In Proceedings of the Workshop on Hot Topics in Video Analytics and Intelligent Edges, Oct. 21, 2019, pp. 27… [cited by applicant]
Gao, et al., “Estimating GPU Memory Consumption of Deep Learning Models”, In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering… [cited by applicant]
Harwell, et al., “Microsoft Rocket Video Analytics Platform”, Retrieved from: https://github.com/microsoft/Microsoft-Rocket-Video-Analytics-Platform, Jun. 2, 2021, 8 Pages. [cited by applicant]
He, et al., “Mask R-CNN”, In Journal of The Computing Research Repository, Mar. 20, 2017, 9 Pages. [cited by applicant]
Hsieh, et al., “Focus: Querying Large Video Datasets with Low Latency and Low Cost”, In Proceedings of 13th Symposium on Operating Systems Design and Implementation, Oct. 18, 2018, pp. 269-286. [cited by applicant]
Huang, et al., “SwapAdvisor: Pushing Deep Learning Beyond the GPU Memory Limit via Smart Swapping”, In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Oper… [cited by applicant]
Hung, et al., “VideoEdge: Processing Camera Streams using Hierarchical Clusters”, In Proceedings of IEEE/ACM Symposium on Edge Computing, Oct. 25, 2018, pp. 115-131. [cited by applicant]
Jain, et al., “Spatula: Efficient Cross-camera Video Analytics on Large Camera Networks”, In Proceedings of ACM/IEEE Symposium on Edge Computing (SEC)., Nov. 12, 2020, 15 Pages. [cited by applicant]
Jeon, et al., “Multi-tenant GPU Clusters for Deep Learning Workloads: Analysis and Implications”, In Microsoft Research Technical Report, May 2018, 14 Pages. [cited by applicant]
Jeon, et al., “Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads”, In Proceedings of the Usenix Conference on Usenix Annual Technical Conference, Jul. 2019, pp. 947-960. [cited by applicant]
Jiang, et al., “Chameleon: Scalable Adaptation of Video Analytics”, In Proceedings of the Conference of the ACM Special Interest Group on Data Communication, Aug. 20, 2018, pp. 253-266. [cited by applicant]
Jiang, et al., “Mainstream: Dynamic Stem-sharing for Multi-tenant Video Processing”, In Proceedings of the USENIX Annual Technical Conference (USENIX ATC '18), Jul. 11, 2018, pp. 29-41. [cited by applicant]
Jiang, et al., “Networked Cameras Are the New Big Data Clusters”, In Proceedings of the Workshop on Hot Topics in Video Analytics and Intelligent Edges, Oct. 21, 2019, pp. 1-17. [cited by applicant]
Kang, et al., “NoScope: Optimizing Neural Network Queries over Video at Scale”, In Proceedings of the VLDB Endowment, Aug. 2017, pp. 1586-1597. [cited by applicant]
Kewalramani, Avi, “Live Video Analytics with Microsoft Rocket for reducing edge compute costs”, Retrieved from: https://techcommunity.microsoft.com/t5/internet-of-things/live-video-analytics-with-microsoft-rocket-for-re… [cited by applicant]
Li, et al., “A convolutional neural network cascade for face detection”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 7, 2015, pp. 5325-5334. [cited by applicant]
Li, et al., “Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video Analytics”, In Proceedings of the Annual Conference of the ACM Special Interest Group on Data Communication on the Applications, Technolog… [cited by applicant]
Lin, et al., “Feature Pyramid Networks for Object Detection”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 22, 2017, pp. 2117-2125. [cited by applicant]
Regazzoni, et al., “Video Analytics for Surveillance: Theory and Practice”, In Journal of IEEE Signal Processing Magazine vol. 27, Issue 5, Sep. 2010, pp. 16-17. [cited by applicant]
Russakovsky, et al., “ImageNet Large Scale Visual Recognition Challenge”, In Publication of International Journal of Computer Vision (IJCV), vol. 115, Issue 3, Dec. 2015, 43 Pages. [cited by applicant]
Sarwar, et al., “Incremental learning in deep convolutional neural networks using partial network sharing”, In Journal of IEEE Access, vol. 8, Dec. 30, 2019, pp. 4615-4628. [cited by applicant]
Senior, et al., “Video analytics for retail”, In Proceedings of IEEE Conference on Advanced Video and Signal Based Surveillance, Sep. 5, 2007, pp. 423-428. [cited by applicant]
Shriram, et al., “Dynamic memory management for gpu-based training of deep neural networks”, In Proceedings of IEEE International Parallel and Distributed Processing Symposium, May 24, 2019, pp. 200-209. [cited by applicant]
Sun, et al., “Adashare: Learning what to share for efficient deep multi-task learning”, In Repository of arXiv:1911.12423v1, Nov. 27, 2019, pp. 1-12. [cited by applicant]
Vahl, et al., “PyTorch-YOLOv3”, Retrieved from: https://github.com/eriklindernoren/PyTorch-YOLOv3, Jul. 11, 2021, 8 Pages. [cited by applicant]
Vandenhende, et al., “Branched multi-task networks: deciding what layers to share”, In Repository of arXiv:1904.02920v5, Aug. 13, 2020, pp. 1-19. [cited by applicant]
Wang, et al., “Characterizing Deep Learning Training Workloads on Alibaba-PAI”, In Repository of arXiv:1910.05930v1, Oct. 14, 2019, 14 Pages. [cited by applicant]
Wang, et al., “Towards Scalable Edge-Native Applications”, In Proceeding of ACM/IEEE Symposium on Edge Computing, Nov. 7, 2019, 14 Pages. [cited by applicant]
Wu, et al., “Irina: Accelerating DNN Inference with Efficient Online Scheduling”, In Proceedings of 4th Asia-Pacific Workshop on Networking, Aug. 3, 2020, pp. 36-43. [cited by applicant]
Xiao, et al., “Gandiva: Introspective Cluster Scheduling for Deep Learning”, In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation, Oct. 8, 2018, pp. 595-610. [cited by applicant]
Yi, et al., “Lavea: Latency-aware video analytics on edge computing platform”, In Proceedings of the Second ACM/IEEE Symposium on Edge Computing, Oct. 12, 2017, 13 Pages. [cited by applicant]
Yosinski, et al., “How transferable are features in deep neural networks?”, In Proceedings of 27th Annual Conference on Neural Information Processing Systems, Nov. 6, 2014, pp. 1-9. [cited by applicant]
Yousefpour, et al., “All one needs to know about fog computing and related edge computing paradigms: A complete survey”, In Journal of Journal of Systems Architecture, vol. 98, Sep. 2019, pp. 289-330. [cited by applicant]
Yu, et al., “A Survey on the Edge Computing for the Internet of Things”, In Journal of IEEE Access, vol. 6, Nov. 29, 2017, pp. 6900-6919. [cited by applicant]
Yu, et al., “Salus: Fine-Grained GPU Sharing Primitives for Deep Learning Applications”, In Proceedings of Machine Learning and Systems, vol. 2, Mar. 2020, 14 Pages. [cited by applicant]
Office Action Received for European Application No. 22726333.2, mailed on Jan. 9, 2024, 3 pages. [cited by applicant]
Communication under Rule 71(3) Received in European Patent Application No. 22726333.2, mailed on Jan. 3, 2025, 08 pages. [cited by applicant]