IP Library › Granted Patent US 12,287,816
Granted Patent B1
US 12,287,816 · App. 18/385,408 · Granted Apr 29, 2025

Reducing latency by processing parts of a language model query in parallel

Inventors: Sayan Dev Pathak (Kirkland, WA); Osama Abuelsorour (Menlo Park, CA); Christopher Hakan Basoglu (Everett, WA); Harini Kesavamoorthy (Bellevue, WA); Girish Milind Mahajan (Redmond, WA); Salman Mohammad Quazi (Mountain View, CA); Valeriy Viktorovich Kirshin (Kirkland, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F16/3329G06F16/3344
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,287,816
App. No.
18/385,408
Granted
Apr 29, 2025
Kind
B1
Abstract

A technique partitions a user's original query into plural smaller component queries, each of which has a common part and an instance-specific part. The technique distributes the component queries to plural processor instances of a processor. The plural processor instances transform the respective component queries into query-component responses by acting in parallel, independent of each other. The technique generates a final response based on the query-component responses, e.g., by assembling the component-query responses into the final response. The technique reduces latency because the processor instances work on parts of the user's original query at the same time, rather than as a single stream of consecutive tokens. The plural processor instances have access to a shared cache memory, and utilize relevant data that has been computed in response to previous queries.

Claims (51)

1. A method for processing a query using a machine-trained language model, comprising:

receiving an original query;

generating component queries based on the original query, the component queries having a same common part, and the component queries having different respective instance-specific parts;

distributing the component queries to respective processor instances, the processor instances being instances of one or more processors, each processor instance executing an instance of the machine-trained language model,

the processor instances generating respective component-query responses in parallel based on the plural component queries, and based on intermediate results previously generated by the machine-trained language model and stored in a cache memory, the previously generated intermediate results including key-value information used in performing an attention operation in the machine-trained language model;

receiving the component-query responses;

generating a final response based on the component-query responses; and

generating output information based on the final response.

2. The method of claim 1 ,

wherein the original query includes a question and source text, the source text serving as context for use by the language model in answering the question, and

wherein the instance-specific parts are associated with respective selected source parts of the source text.

3. The method of claim 2 , wherein the selected source parts are a subset of the source text that are collectively less than an entirety of the source text.

4. The method of claim 2 , wherein the source text includes plural consecutive source parts, and wherein the selected source parts include at least a contiguous portion of the plural consecutive source parts.

5. The method of claim 2 , wherein the selected source parts are automatically selected based on a determination that the selected source parts have a greatest relevance to the question.

6. The method of claim 1 , further comprising assigning a common query identifier to the component queries, wherein the component-query responses are associated with the common query identifier.

7. The method of claim 1 , further comprising determining whether it is appropriate to partition the original query into the component queries by determining whether the original query includes a predetermined key term and/or a predetermined semantic concept, and/or matches a predetermined structure.

8. The method of claim 1 , wherein the generating a final response comprises assembling the component-query responses into the final response in an order in which the component queries appear in the original query.

9. The method of claim 1 , wherein the generating a final response comprises comparing information imparted by at least two of the component-query responses, and generating an output result that expresses a result of the comparing.

10. The method of claim 1 , further comprising detecting a predetermined key term and/or a predetermined semantic concept in the original query and/or the component query responses, wherein the generating a final response is controlled based on the key term and/or semantic concept that has been detected by the detecting.

11. The method of claim 1 , further comprising: determining that at least one of the component-query responses includes invocation information associated with a particular function; invoking the particular function; and receiving supplemental data in response to the invoking.

12. The method of claim 1 , wherein the method further comprises:

automatically generating additional component queries based on at least one of the component-query responses, the additional component queries having a child relationship with respect to said at least one of the component-query responses; and

instructing the plural processor instances to generate another set of component-query responses based on the additional component queries.

13. The method of claim 12 , wherein the additional component queries include supplemental data obtained in response to invoking a particular function, the particular function being invoked in response to invocation information provided by the component-query responses.

14. The method of claim 1 , further comprising selecting the one or more processors from among a group of candidate processors based on a determination that the original query contains tokens that have been previously processed by the one or more processors.

15. A computing system for processing a query using a machine-trained language model, comprising:

an instruction data store for storing computer-readable instructions; and

a processing system for executing the computer-readable instructions in the data store, to perform operations including:

receiving an original query that includes a question and source text;

determining that the original query is capable of being partitioned into component queries;

in response to the determining, generating the component queries based on the question and the source text, the component queries having a same common part that expresses the question, and the component queries expressing respective selected source parts of the source text;

distributing the component queries to respective processor instances, the processor instances being instances of one or more processors, each processor instance executing an instance of the machine-trained language model,

the processor instances generating respective component-query responses in parallel based on the plural component queries, and based on previously-generated intermediary results stored in a shared cache memory of the one or more processors,

the previously-generated intermediate results including key-value information used in performing an attention operation in the machine-trained language model;

receiving the component-query responses;

generating a final response based on the component-query responses; and

generating output information based on the final response.

16. The computing system of claim 15 , wherein the determining that the query is capable of being partitioned involves determining whether the original query includes a predetermined key term and/or a predetermined semantic concept, and/or matches a predetermined structure.

17. A computer-readable storage medium for storing computer-readable instructions, a processing system executing the computer-readable instructions to perform operations, the operations comprising each of:

receiving an original query that includes a question and a separate source text that accompanies the question and is separate therefrom,

the question containing a reference to the source text,

the source text including plural source parts that serve as context for use by a machine-trained language model in answering the question;

upon determining that the original query is severable into independent parts, generating component queries based on the original query, the component queries having a same common part, and the component queries expressing respective selected source parts of the plural source parts;

distributing the component queries to respective processor instances, the processor instances being instances of one or more processors, each processor instance executing an instance of the machine-trained language model,

the processor instances auto-regressively generating respective component-query responses in parallel based on the plural component queries;

receiving the component-query responses;

generating a final response based on the component-query responses; and

generating output information based on the final response.

18. The computer-readable storage medium of claim 17 , wherein the selected source parts are a subset of the source text that are collectively less than an entirety of the source text.

19. The computer-readable storage medium of claim 18 , wherein the operations further include automatically selecting the source parts from the source text based on a determination that the source parts have a greatest relevance to the question.

20. The computer-readable storage medium of claim 17 , wherein the processor instances generate the respective component-query responses based, in part, on intermediate results previously generated by the machine-trained language model and stored in a cache memory, the previously generated intermediate results including key-value information used in performing an attention operation in the machine-trained language model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: PATHAK, SAYAN DEV; ABUELSOROUR, OSAMA; BASOGLU, CHRISTOPHER HAKAN; KESAVAMOORTHY, HARINI; MAHAJAN, GIRISH MILIND; QUAZI, SALMAN MOHAMMAD; KIRSHIN, VALERIY VIKTOROVICH
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065398/0530 →
References Cited (30)
US 9934261B2 · Kuno · 2018 [cited by examiner]
US 11308106B1 · Muralimanohar · 2022 [cited by examiner]
US 20110228668A1 · Pillai · 2011 [cited by examiner]
US 20130125099A1 · Budiu · 2013 [cited by examiner]
US 20150160817A1 · Hwang · 2015 [cited by examiner]
US 20160171050A1 · Das · 2016 [cited by examiner]
US 20210217408A1 · Hakkani-Tur et al. · 2021 [cited by applicant]
US 20220101189A1 · Ben-Itzhak · 2022 [cited by examiner]
US 20230177401A1 · Yu · 2023 [cited by examiner]
US 20230359619A1 · Sharan · 2023 [cited by examiner]
US 20230385275A1 · Obeidi · 2023 [cited by examiner]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv, arXiv:1810.04805v2 [cs.CL], May 24, 2019, 16 pages. [cited by applicant]
Vaswani, et al., “Attention Is All You Need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017, 11 pages. [cited by applicant]
Radford, et al., “Improving Language Understanding by Generative Pre-Training,” available at http://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf, OpenAI, San Francisco, Californi… [cited by applicant]
Touvron, et al., “LLaMA: Open and Efficient Foundation Language Models,” arXiv, arXiv:2302.13971v1 [cs.CL], Feb. 27, 2023, 27 pages. [cited by applicant]
Scao, et al., “BLOOM: A 176B-Parameter Open-Access Multilingual Language Model,” arXiv, arXiv:2211.05100v2 [cs.CL], Dec. 11, 2022, 62 pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners,” arXiv, arXiv:2005.14165v4 [cs.CL], Jul. 22, 2020, 75 pages. [cited by applicant]
Legenzoff, Derek “Function calling is now available in Azure OpenAI Service,” available at read://https_techcommunity.microsoft.com/?url=https%3A%2F%2Ftechcommunity.microsoft.com%2Ft5%2Fazure-ai-services-blog%2Ffunction… [cited by applicant]
“Extending LLM Context Length,” available at https://github.com/abacusai/Long-Context, Github, accessed on Oct. 26, 2023, 12 pages. [cited by applicant]
“LangChain,” available at https://en.wikipedia.org/wiki/LangChainWikipedia article, accessed on Oct. 26, 2023, 3 pages. [cited by applicant]
“OpenAPI Specification,” available at https://en.wikipedia.org/wiki/OpenAPI_Specification, Wikipedia article, accessed on Oct. 26, 2023, 5 pages. [cited by applicant]
“How to use function calling with Azure OpenAI Service (Preview),” available at https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/function-calling, Microsoft Azure documentation, Jul. 20, 2023, 8 pages. [cited by applicant]
Gupta, “Compression of Deep Learning Models for Text: A Survey,” arXiv, arXiv:2008.05221v4 [cs. CL], Jun. 13, 2021, 53 pages. [cited by applicant]
Chen, et al., “Stabilized In-Context Learning with Pre-trained Language Models for Few Shot Dialogue State Tracking,” arXiv, arXiv:2302.05932v1 [cs.CL], Feb. 12, 2023, 14 pages. [cited by applicant]
Santra, et al., “Frugal Prompting for Dialog Models,” arXiv, arXiv:2305.14919v1 [cs.CL], May 24, 2023, 22 pages. [cited by applicant]
Jacobs, et al., “DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models,” arXiv, arXiv:2309.14509v1 [cs.LG], Sep. 25, 2023, 9 pages. [cited by applicant]
Li, et al., “LightSeq: Sequence Level Parallelism for Distributed Training of Long Context Transformers,” arXiv, arXiv:2310.03294v1 [cs.LG], Oct. 5, 2023, 14 pages. [cited by applicant]
Li, et al., “Sequence Parallelism: Making 4D Parallelism Possible,” arXiv, arXiv:2105.13120v1 [cs.LG], May 26, 2021, 12 pages. [cited by applicant]
PCT Search Report and Written Opinion for PCT/US2024/049691, identified mailing date of Dec. 17, 2024, 14 pages. [cited by applicant]
U.S. Appl. No. 18/401,060, filed Dec. 29, 2023. [cited by applicant]