IP Library › Granted Patent US 12,566,626
Granted Patent B2
US 12,566,626 · App. 18/657,377 · Granted Mar 3, 2026

Edge cloud hierarchical language model design

Inventors: Shadi Abdollahian Noghabi (Bellevue, WA); Ranveer Chandra (Kirkland, WA); Leonardo de Oliveira Nunes (Rio de Janeiro, BR); Alexander Steven Crown (Bellevue, WA); Vinamra Benara (Berkeley, CA)
Assignee: Microsoft Technology Licensing, LLC
G06F9/48G06N3/082H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,626
App. No.
18/657,377
Granted
Mar 3, 2026
Kind
B2
Abstract

The present disclosure relates to systems and methods for using language models in locations with limited network connectivity. The systems and methods include a hierarchical edge architecture with a plurality of language models with diverse compute capabilities. The systems and methods dynamically select a language model from the plurality of language models to use to respond to a query received by a user in response to determining a level of network connectivity available at a user device.

Claims (36)

1 . A method comprising:

receiving, at a user device, a query;

selecting, from a plurality of language models with diverse compute capabilities, a small language model (SLM) on the user device based on data available on the user device, query parameters of the query, a required accuracy for the query, and available network connectivity;

retrieving user context and industry context relevant to the query from the data on the user device, wherein the user context and the industry context is automatically generated from collected data and provided in vectorized databases for storage on the user device;

using the SLM to provide a response to the query tailored to the user context and the industry context; and

providing the response to a user.

2 . The method of claim 1 , wherein the plurality of language models include the SLM at the user device, a medium language model at a device of a back office edge in communication with the user device, and a large language model (LLM) at a device of a cloud network in communication with the user device.

3 . The method of claim 2 , wherein the SLM has less compute and memory footprint as compared to the medium language model and the LLM, and the medium language model has less compute and memory footprint as compared to the LLM.

4 . The method of claim 1 , further comprising:

selecting, from the plurality of language models, a medium language model on a device at a back office edge in communication with the user device to provide the response to the query in response to determining that the SLM is unable to provide the response to the query within a latency threshold, wherein the medium language model includes more compute and memory footprint as compared to the SLM.

5 . The method of claim 1 , further comprising:

selecting, from the plurality of language models, a large language model (LLM) at a cloud network in communication with the user device to provide the response to the query in response to determining that the SLM is unable to provide an accurate response to the query, wherein the LLM includes more compute and memory footprint as compared to the SLM.

6 . The method of claim 1 , further comprising:

using, by the SLM, the user context and the industry context stored on the user device to provide the response to the query.

7 . The method of claim 6 , wherein the user context includes information obtained from private documents of the user with data specific to the user and the response is tailored to the user using the data specific to the user.

8 . The method of claim 6 , wherein the industry context includes data specific to an industry that is related to the user.

9 . The method of claim 6 , further comprising:

periodically receiving the context from a device at a back office edge in communication with the user device or devices on a cloud network in communication with the user device, wherein the context is obtained from external sources and information specific to the user.

10 . The method of claim 6 , wherein the user context and the industry context is automatically generated by a device at a back office edge in communication with the user device or devices on a cloud network in communication with the user device in response to sensory data obtained by sensors at a location.

11 . A user device comprising:

a memory to store data and instructions; and

a processor operable to communicate with the memory, wherein the processor is operable to:

receive a query;

select, from a plurality of language models with diverse compute capabilities, a small language model (SLM) on the user device based on data available on the user device, query parameters of the query, a required accuracy for the query, and available network connectivity;

retrieving user context and industry context relevant to the query from the data on the user device, wherein the user context and the industry context is automatically generated from collected data and provided in vectorized databases for storage on the user device;

use the SLM on the user device to provide a response to the query tailored to the user context and the industry context; and

provide the response to a user.

12 . The user device of claim 11 , wherein the plurality of language models include the SLM at the user device, a medium language model at a device of a back office edge in communication with the user device, and a large language model (LLM) at a device of a cloud network in communication with the user device.

13 . The user device of claim 12 , wherein the SLM has less compute as compared to the medium language model and the LLM, and the medium language model has more compute as compared to the SLM and less compute as compared to the LLM.

14 . The user device of claim 11 , wherein the processor is further operable to select, from the plurality of language models, a medium language model on a device at a back office edge in communication with the user device to provide the response to the query in response to determining that the SLM is unable to provide the response to the query within a latency threshold, wherein the medium language model includes more compute as compared to the SLM.

15 . The user device of claim 11 , wherein the processor is further operable to select, from the plurality of language models, a large language model (LLM) at a cloud network in communication with the user device to provide the response to the query in response to determining that the SLM is unable to provide an accurate response to the query, wherein the LLM includes more compute as compared to the SLM.

16 . The user device of claim 11 , wherein the processor is further operable to use, by the SLM, the user context and the industry context stored on the user device to provide the response to the query.

17 . The user device of claim 16 , wherein the context includes user context with data specific to the user obtained from images of a location of the user and the response is tailored to the user using the data specific to the user.

18 . The user device of claim 16 , wherein the industry context includes data specific to an industry that is related to the user.

19 . The user device of claim 16 , wherein the processor is further operable to periodically receive the context from a device at a back office edge in communication with the user device or devices on a cloud network in communication with the user device, wherein the context is obtained from external sources and information specific to the user.

20 . The user device of claim 16 , wherein the context is automatically generated by a device at a back office edge in communication with the user device or devices on a cloud network in communication with the user device in response to sensory data obtained by sensors at a location; and the context is provided to the user device for storage in vectorized databases.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2024
From: CHANDRA, RANVEER
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 068780/0494 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2024
From: ABDOLLAHIAN NOGHABI, SHADI; NUNES, LEONARDO DE OLIVEIRA; CROWN, ALEXANDER STEVEN; BENARA, VINAMRA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 067338/0825 →
Continuity (1)
Related Publication 20250348349A1 · Nov 13, 2025
References Cited (27)
US 9064495B1 · Torok · 2015 [cited by examiner]
US 9430465B2 · Waibel · 2016 [cited by examiner]
US 9514747B1 · Bisani · 2016 [cited by examiner]
US 10621282B1 · Selfridge · 2020 [cited by examiner]
US 11238849B1 · Mimassi · 2022 [cited by examiner]
US 11314941B2 · Aly · 2022 [cited by examiner]
US 11521066B2 · Lee · 2022 [cited by examiner]
US 12155535B2 · Pack · 2024 [cited by examiner]
US 12206552B2 · Guim Bernat · 2025 [cited by examiner]
US 12265788B1 · Fieldman · 2025 [cited by examiner]
US 12321794B2 · Faonte · 2025 [cited by examiner]
US 20180211668A1 · Willett · 2018 [cited by examiner]
US 20230164030A1 · Ning · 2023 [cited by examiner]
US 20230325670A1 · Clemons · 2023 [cited by examiner]
US 20240203127A1 · Kamani · 2024 [cited by examiner]
US 20240330699A1 · Yao · 2024 [cited by examiner]
US 20250086952A1 · Wu · 2025 [cited by examiner]
US 20250094878A1 · Rivlin · 2025 [cited by examiner]
US 20250140245A1 · Lee · 2025 [cited by examiner]
US 20250147811A1 · Khosrowpour · 2025 [cited by examiner]
US 20250156487A1 · Nguyen · 2025 [cited by examiner]
WO WO2025093339A1 · 2025 [cited by examiner]
Minrui Xu et al. “Unleashing the Power of Edge-Cloud Generative AI in Mobile Networks: A Survey of AIGC Services”, Oct. 31, 2023, 43 pages. (Year: 2023). [cited by examiner]
Lubomir Bulej et al. “Managing latency in edge-cloud environment”, The Journal of Systems & Software 172 (2021) 110872, 15 pages. (Year: 2021). [cited by examiner]
Jinke Ren et al. “Collaborative Cloud and Edge Computing for Latency Minimization”, IEEE Transactions on Vehicular Technology, vol. 68, No. 5, May 2019, 14 pages. (Year: 2019). [cited by examiner]
Extended European search report received for European Application No. 25172623.8, mailed on Oct. 20, 2025, 7 Pages. [cited by applicant]
Pravinkrishnan, et al., “An Overview of Chatbots using ML Algorithms in Agricultural Domain”, International Journal of Computer Applications, vol. 184, No. 11, May 1, 2022, pp. 15-22. [cited by applicant]