IP Library › Granted Patent US 12,619,821
Granted Patent B2
US 12,619,821 · App. 18/395,189 · Granted May 5, 2026

Expediting generative token production using speculative sampling, added guidance, and language models of different capacities

Inventors: Ayyoob Imanigooghari (Munich, DE); Mohsen Fayyaz (Berlin, DE); Eric Chris Wolfgang Sommerlade (Oxford, GB)
Assignee: Microsoft Technology Licensing, LLC
G06F40/284G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,821
App. No.
18/395,189
Granted
May 5, 2026
Kind
B2
Abstract

A technique accelerates the generative production of tokens using a target language model that operates in cooperation with a draft language model. The target language model is more capable, but slower, compared to the draft language model. In operation, the draft language model transforms prompt tokens into draft tokens. The target language model edits the draft tokens, e.g., by selecting zero, one, or more of the draft tokens, and by also predicting a next token to follow the draft token(s) (if any) that are selected. Further, the target language model produces guidance vector information. In a subsequent cycle, the draft language model uses the guidance vector information to produce an updated set of set of draft tokens. The guidance vector information informs the draft language model of the embedding space being used by the target language model. This achieves a more effective cooperative relation between the two models.

Claims (47)

1 . A method for accelerating generation of output tokens using a target language model, which operates in cooperation with a draft language model, comprising:

receiving, by the target language model, a set of draft tokens produced by the draft language model based on, at least in part, prompt tokens provided to the draft language model;

wherein the target language model and the draft language model are two different neural networks, and wherein the target language model has more parameters and is more accurate compared to the draft language model, and wherein the target language model consumes more memory and processing resources compared to the draft language model, and wherein the target language model is slower in operation compared to the draft language model;

producing, using a first head neural network of the target language model, one or more target output tokens based on the prompt tokens and the set of draft tokens,

the one or more target output tokens including zero, one, or more draft tokens chosen from among the set of draft tokens, and an additional target output token which is predicted by the target language model to follow the zero, one, or more draft tokens that are selected;

generating, using a second head neural network of the target language model, guidance vector information based on the prompt tokens and the set of draft tokens, the second head neural network being different than the first head neural network; and

forwarding the one or more target output tokens and the guidance vector information to the draft language model, a combination of the prompt tokens, the one or more target output tokens, and the guidance vector information being used by the draft language model as input tokens to be transformed into updated draft tokens.

2 . The method of claim 1 , wherein the target language model applies an attention operation to input tokens that are input to the target language model, and the draft language model applies an attention operation to input tokens that are input to the draft language model.

3 . The method of claim 1 , wherein a server system implements the target language model and a local system implements the local language model, the server system being accessible to the local system via a computer network.

4 . The method of claim 1 , wherein both the target language model and the draft language model are implemented by a same system.

5 . The method of claim 1 , wherein the producing accepts a particular draft token of the zero, one, or more draft tokens upon determining that a probability generated by the target language model for the particular draft token is greater than a probability generated by the draft language model for the particular draft token.

6 . The method of claim 1 , wherein the target language model operates by:

transforming the prompt tokens and the set of draft tokens to hidden state information using a base target language model;

transforming the hidden state information to output token probability information using the first head neural network, on basis of which the one or more target output tokens are produced; and

transforming the hidden state information to the guidance vector information using the second head neural network.

7 . The method of claim 1 , further comprising producing the one or more target output tokens in a single forward pass of the target language model.

8 . The method of claim 1 , wherein the draft tokens in the set of draft tokens are produced auto-regressively by the draft language model.

9 . The method of claim 1 ,

further comprising transforming, using the draft language model, the combination of the prompt tokens, the one or more target output tokens, and the guidance vector information to the updated draft tokens,

wherein the target language model has been trained based on a first loss measure that depends on a difference between first ground-truth information and the one or more target output tokens, and a second loss measure that depends on a difference between second ground-truth information and the updated draft tokens.

10 . The method of claim 9 , wherein the draft language model has also been trained based on the second loss measure.

11 . The method of claim 9 , wherein the first ground-truth information is text that is manually specified by a human reviewer as being correct, and wherein the second ground-truth information is text auto-regressively generated by the target language model.

12 . The method of claim 1 , further comprising switching to a mode in which the target language model is asked by the draft language model to generate an instance of guidance vector information for initial prompt tokens, and wherein generation of output tokens thereafter takes place based on the guidance vector information using the draft language model independent of interaction with the target language model.

13 . A computing system for using a draft language model to accelerate generation of output tokens using a target language model, comprising:

an instruction data store for storing computer-readable instructions; and

a processing system for executing the computer-readable instructions in the data store, to perform operations including:

receiving a set of target output tokens produced by the target language model, and guidance vector information produced by the target language model, wherein a combination of the prompt tokens, the set of target output tokens, and the guidance vector information comprise input tokens;

transforming the input tokens into draft tokens,

wherein the target language model and the draft language model are two different neural networks, and wherein the draft language model has fewer parameters and is less accurate compared to the target language model, and wherein the draft language model consumes less memory and processing resources compared to the target language model, and wherein the draft language model is faster in operation compared to the target language model; and

sending the draft tokens to the target language model, for use by the target language model in producing an updated set of target output tokens using a first head neural network and updated guidance vector information using a second head neural network that is different than the first head neural network.

14 . The computing system of claim 13 , wherein the updated set of target output tokens are produced by the target language model by selecting from among the draft tokens.

15 . The computing system of claim 13 , wherein the operations further comprise switching to a mode in which the target language model is asked by the draft language model to generate an instance of guidance vector information for initial prompt tokens, and wherein generation of output tokens thereafter takes place based on the guidance vector information using the draft language model independent of interaction with the target language model.

16 . A computer-readable storage medium for storing computer-readable instructions, a processing system executing the computer-readable instructions to perform operations, the operations comprising each of:

receiving, by a target language model, a set of draft tokens produced by a draft language model based on, at least in part, prompt tokens provided to the draft language model;

wherein the target language model and the draft language model are two different neural networks, and wherein the target language model has more parameters and is more accurate compared to the draft language model, and wherein the target language model consumes more memory and processing resources compared to the draft language model, and wherein the target language model is slower in operation compared to the draft language model;

producing, using a first head neural network of the target language model, one or more target output tokens based on the prompt tokens and the set of draft tokens,

the one or more target output tokens including zero, one, or more draft tokens chosen from among the set of draft tokens, and an additional target output token which is predicted by the target language model to follow the zero, one, or more draft tokens that are selected;

generating, using a second head neural network of the target language model, guidance vector information based on the prompt tokens and the set of draft tokens, wherein the second head neural network is different than the first head neural network, and wherein a combination of the prompt tokens, the set of target output tokens, and the guidance vector information comprise input tokens; and

transforming, using the draft language model, the input tokens to updated draft tokens,

the target language model has been trained based on a first loss measure that depends on a difference between first ground-truth information and the one or more target output tokens, and a second loss measure that depends on a difference between second ground-truth information and the updated draft tokens.

17 . The computer-readable storage medium of claim 16 , wherein the draft language model has also been trained based on the second loss measure.

18 . The computer-readable storage medium of claim 16 , wherein the first ground-truth information is text that is manually specified by a human reviewer as being correct, and wherein the second ground-truth information is text auto-regressively generated by the target language model.

19 . The method of claim 12 , wherein the mode is selected automatically based on one or more factors, as expressed in one or more input signals.

20 . The computer-readable storage medium of claim 16 , wherein the target language model operates by:

transforming the prompt tokens and the set of draft tokens to hidden state information using a base target language model;

transforming the hidden state information to output token probability information using the first head neural network, on basis of which the one or more target output tokens are produced; and

transforming the hidden state information to the guidance vector information using the second head neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2023
From: IMANIGOOGHARI, AYYOOB; FAYYAZ, MOHSEN; SOMMERLADE, ERIC CHRIS WOLFGANG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065947/0236 →
Continuity (1)
Related Publication 20250209271A1 · Jun 26, 2025
References Cited (84)
US 20240320433A1 · Lott · 2024 [cited by examiner]
US 20250053748A1 · Fayyaz et al. · 2025 [cited by applicant]
US 20250086187A1 · Fayyaz et al. · 2025 [cited by applicant]
US 20250209271A1 · Imanigooghari et al. · 2025 [cited by applicant]
US 20250299026A1 · Imanigooghari et al. · 2025 [cited by applicant]
Chen, Charlie, et al. “Accelerating large language model decoding with speculative sampling.” arXiv preprint arXiv:2302.01318 (2023). (Year: 2023). [cited by examiner]
Liu Xiaoxuan et al: “Online Speculative Decoding”,, Oct. 17, 2023 (Oct. 17, 2023), XP093248186, Retrieved from the Internet: URL:https://arxiv.org/pdf/2310.07177v2 (Year: 2023). [cited by examiner]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv, arXiv:1810.04805v2 [cs.CL], May 24, 2019, 16 pages. [cited by applicant]
Scao, et al., “BLOOM: A 176B-Parameter Open-Access Multilingual Language Model,” arXiv, arXiv:2211.05100v2 [cs.CL], Dec. 11, 2022, 62 pages. [cited by applicant]
Vaswani, et al., “Attention Is All You Need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017, 11 pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners,” arXiv, arXiv:2005.14165v4 [cs.CL], Jul. 22, 2020, 75 pages. [cited by applicant]
“Introducing The World's Largest Open Multilingual Language Model: BLOOM,” available at https://bigscience.huggingface.co/blog/bloom, accessed on Feb. 13, 2023, 2 pages. [cited by applicant]
Houlsby, et al., “Parameter-Efficient Transfer Learning for NLP,” arXiv, arXiv:1902.00751v2 [cs.LG], Jun. 13, 2019, 13 pages. [cited by applicant]
Leviathan, et al., “Fast Inference from Transformers via Speculative Decoding,” in Proceedings of the 40th International Conference on Machine Learning, PMLR 202, Jul. 2023, 13 pages. [cited by applicant]
Chen, et al., “Accelerating Large Language Model Decoding with Speculative Sampling,” arXiv, arXiv:2302.01318v1 [cs.CL], Feb. 2, 2023, 11 pages. [cited by applicant]
Hu, et al., “LoRA: Low-Rank Adaptation of Large Language Models,” in Proceedings of 10th International Conference on Learning Representations, Apr. 2022, 13 pages. [cited by applicant]
Rafailov, et al., “Direct Preference Optimization: Your Language Model is Secretly a Reward Model,” arXiv, arXiv:2305.18290v1 [cs.LG], May 29, 2023, 26 pages. [cited by applicant]
Banino, et al., “PonderNet: Learning to Ponder,” in 8th ICML Workshop on Automated Machine Learning (2021), 2021, 16 pages. [cited by applicant]
Lester, Brian, “Guiding Frozen Language Models with Learned Soft Prompts,” available at https://ai.googleblog.com/2022/02/guiding-frozen-language-models-with.html, Google Research Blogs, Feb. 10, 2022, 5 pages. [cited by applicant]
Lester, et al., “The Power of Scale for Parameter-Efficient Prompt Tuning,” arXiv, arXiv:2104.08691v2 [cs.CL], Sep. 2, 2021, 15 pages. [cited by applicant]
Rao, et al., “DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification,” in 35th Conference on Neural Information Processing Systems (NeurIPS 2021), 2021, 13 pages. [cited by applicant]
Hu, at al., “LoRA: Low-Rank Adaptation of Large Language Models,” arXiv, arXiv:2106.09685v2 [cs.CL], Oct. 16, 2021, 26 pages. [cited by applicant]
Radford, et al., “Improving Language Understanding by Generative Pre-Training,” available at https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf, OpenAI, San Francisco, Californ… [cited by applicant]
Touvron, et al., “LLaMA: Open and Efficient Foundation Language Models,” arXiv, arXiv:2302.13971v1 [cs.CL], Feb. 27, 2023, 27 pages. [cited by applicant]
Fayyaz, et al., “Compressing Information Provided to a Machine-Trained Model Using Abstract Tokens,” U.S. Appl. No. 18/232,485, filed Aug. 10, 2023, 62 pages. [cited by applicant]
Fayyaz, et al., “Executing a Client Model Using a Task Prompt Produced by a Main System,” U.S. Appl. No. 18/244,229, filed Sep. 9, 2023, 52 pages. [cited by applicant]
Khoshnoodi, et al., “A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models,” arXiv, arXiv:2405.13019v2 [cs.CL], Mar. 24, 2024, 27 pages. [cited by applicant]
International Search Report and Written Opinion for PCT Application No. PCT/US2024/055295, mailing date Feb. 20, 2025, 15 pages. [cited by applicant]
Leviathan, et al., “Fast Inference from Transformers via Speculative Decoding,” arXiv, arXiv:2211.17192v1 [cs.LG] Nov. 30, 2022, 12 pages. [cited by applicant]
Liu, et al., “Online Speculative Decoding,” arXiv, arXiv:2310.07177v2 [cs.AI], Oct. 17, 2023, 14 pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners,” in 34th Conference on Neural Information Processing Systems (NeurIPS 2020), 2020, 25 pages. [cited by applicant]
Bubeck, et al., “Sparks of Artificial General Intelligence: Early experiments with GPT-4,” arXiv, arXiv:2303.12712v5 [cs.CL], Apr. 13, 2023, 155 pages. [cited by applicant]
Wei, et al., “Emergent Abilities of Large Language Models,” arXiv, arXiv:2206.07682v2 [cs.CL], Oct. 26, 2022, 30 pages. [cited by applicant]
Dehghani, et al., “The Efficiency Misnomer,” arXiv, arXiv:2110.12894v2 [cs.LG], Mar. 16, 2022, 16 pages. [cited by applicant]
Frantar, et al., “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers,” arXiv, arXiv:2210.17323v2 [cs.LG], Mar. 22, 2023, 16 pages. [cited by applicant]
Gale, et al., “The State of Sparsity in Deep Neural Networks,” arXiv, arXiv:1902.09574v1 [cs.LG], Feb. 25, 2019, 15 pages. [cited by applicant]
Ghazvininejad, et al., “Mask-Predict: Parallel Decoding of Conditional Masked Language Models,” arXiv, arXiv:1904.09324v2 [cs.CL], Sep. 4, 2019, 10 pages. [cited by applicant]
Goyal, et al., “POWER-BERT: Accelerating BERT Inference via Progressive Word-vector Elimination,” in Proceedings of the 37th International Conference on Machine Learning, Online, PMLR 119, 2020, 10 pages. [cited by applicant]
Gu, et al., “Non-Autoregressive Neural Machine Translation,” arXiv, arXiv:1711.02281v2 [cs.CL], Mar. 9, 2018, 13 pages. [cited by applicant]
Gu, et al., “Levenshtein Transformer,” in 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 11 pages. [cited by applicant]
Gunasekar, et al., “Textbooks Are All You Need,” arXiv, arXiv:2306.11644v2 [cs.CL], Oct. 2, 2023, 26 pages. [cited by applicant]
Guo, et al., “Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine Translation,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, pp. 376-… [cited by applicant]
Han, et al., “Dynamic Neural Networks: A Survey,” arXiv, arXiv:2102.04906v4 [cs.CV], Dec. 2, 2021, 20 pages. [cited by applicant]
Hinton, et al., “Distilling the Knowledge in a Neural Network,” arXiv, arXiv:1503.02531v1 [stat.ML], Mar. 9, 2015, 9 pages. [cited by applicant]
Holtzman, et al., “The Curious Case of Neural Text Degeneration,” arXiv, arXiv:1904.09751v1 [cs.CL], Apr. 22, 2019, 11 pages. [cited by applicant]
Iandola, et al., “SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,” arXiv, arXiv:2006.11316v1 [cs.CL], Jun. 19, 2020, 19 pages. [cited by applicant]
Jaszczur, et al., “Sparse is Enough in Scaling Transformers,” arXiv, arXiv:2111.12763v1 [cs.LG], Nov. 24, 2021, 22 pages. [cited by applicant]
Jaszczur, et al., “Sparse is Enough in Scaling Transformers,” in 35th Conference on Neural Information Processing Systems (NeurIPS 2021), 2021, 13 pages. [cited by applicant]
Jiang, et al., “Mistral 7B,” arXiv, arXiv:2310.06825v1 [cs.CL], Oct. 10, 2023, 9 pages. [cited by applicant]
Gante, Joao, “Assisted Generation: a new direction toward low-latency text generation,” available at https://huggingface.co/blog/assisted-generation, Hugging Face, May 11, 2023, 16 pages. [cited by applicant]
Kim, et al., “I-BERT: Integer-only BERT Quantization,” arXiv, arXiv:2101.01321v3 [cs.CL], Jun. 8, 2021, 15 pages. [cited by applicant]
Kim, et al., “Speculative Decoding with Big Little Decoder,” arXiv, arXiv:2302.07863v4 [cs.CL], Oct. 12, 2023, 21 pages. [cited by applicant]
Kitaev, et al., “Reformer: The Efficient Transformer,” in International Conference on Learning Representations, 2020, 12 pages. [cited by applicant]
Lan, et al., “Albert: A Lite BERT for Self-supervised Learning of Language Representations,” arXiv, arXiv:1909.11942v6 [cs.CL], Feb. 9, 2020, 17 pages. [cited by applicant]
Langley, Pat, “Crafting Papers on Machine Learning,” available at https://icml.cc/Conferences/2002/craft.html, in ICML '00: Proceedings of the Seventeenth International Conference on Machine Learning, Jun. 2000, 7 pages. [cited by applicant]
Lee, et al., “Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative Refinement,” arXiv, arXiv:1802.06901v3 [cs.LG], Aug. 27, 2018, 11 pages. [cited by applicant]
Welleck, et al., “Non-Monotonic Sequential Text Generation,” in Proceedings of the 36th International Conference on Machine Learning, PMLR 97, 2019, 11 pages. [cited by applicant]
Li, et al., “Textbooks Are All You Need II: phi-1.5 technical report,” arxiv, arXiv:2309.05463v1 [cs.CL], Sep. 11, 2023, 16 pages. [cited by applicant]
Li, et al, “Hint-Based Training for Non-Autoregressive Machine Translation,” arXiv, arXiv:1909.06708v1 [cs.CL], Sep. 15, 2019, 9 pages. [cited by applicant]
Michel, et al., “Are Sixteen Heads Really Better than One?,” arXiv, arXiv:1905.10650v3 [cs.CL], Nov. 4, 2019, 13 pages. [cited by applicant]
Saha, “Can Language Models Teach Weaker Agents? Teacher Explanations Improve Students via Personalization,” arXiv, arXiv:2306.09299v2 [cs.CL], Nov. 14, 2023, 2023, 23 pages. [cited by applicant]
Sanh, et al., “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” arXiv, arXiv:1910.01108v4 [cs.CL], Mar. 1, 2020, 5 pages. [cited by applicant]
Schuster, et al., “Consistent Accelerated Inference via Confident Adaptive Transformers,” arXiv, arXiv:2104.08803v2 [cs.CL], Sep. 9, 2021, 18 pages. [cited by applicant]
Schwartz, et al., “The Right Tool for the Job: Matching Model and Instance Complexities,” arXiv, arXiv:2004.07453v2 [cs.CL], May 9, 2020, 12 pages. [cited by applicant]
Shao, et al., “Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, 2020, pp. 198-205. [cited by applicant]
Shazeer, Noam, “Fast Transformer Decoding: One Write-Head is All You Need,” arXiv, arXiv:1911.02150v1 [cs.NE], Nov. 6, 2019, 9 pages. [cited by applicant]
So, et al., “Primer: Searching for Efficient Transformers for Language Modeling,” in 35th Conference on Neural Information Processing Systems (NeurIPS 2021), 2021, 13 pages. [cited by applicant]
Stern, et al., “Blockwise Parallel Decoding for Deep Autoregressive Models,” in 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), 2018, 10 pages. [cited by applicant]
Stern, et al., “Insertion Transformer: Flexible Sequence Generation via Insertion Operations,” in Proceedings of the 36 th International Conference on Machine Learning, PMLR 97, 2019, 10 pages. [cited by applicant]
Sukhbaatar, et al., “Adaptive Attention Span in Transformers,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 2019, pp. 331-335. [cited by applicant]
Zhang, et al., “OPT: Open Pre-trained Transformer Language Models,” arXiv, arXiv:2205.01068v4 [cs.CL], Jun. 21, 2022, 30 pages. [cited by applicant]
Wellman, et al., “Theory of mind for learning and teaching: the nature and role of explanation,” in Cognitive Development, vol. 19, Issue 4, Oct. 2004, pp. 479-497. [cited by applicant]
Sun, et al., “Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Con… [cited by applicant]
Sun, et al., “Fast Structured Decoding for Sequence Models,” in 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 11 pages. [cited by applicant]
Sun, et al., “MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices,” arXiv, arXiv:2004.02984v2 [cs.CL], Apr. 14, 2020, 13 pages. [cited by applicant]
Touvron, et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models,” arXiv, arXiv:2307.09288v2 [cs.CL], Jul. 19, 2023, 23 pages. [cited by applicant]
Wang, et al., “Non-Autoregressive Machine Translation with Auxiliary Regularization,” in AAAI'19/IAAI'19/EAAl'19: Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Ap… [cited by applicant]
Wei, et al., “Imitation Learning for Non-Autoregressive Neural Machine Translation,” arXiv, arXiv:1906.02041v2 [cs.CL], Jul. 18, 2019, 9 pages. [cited by applicant]
Chen, et al., “Punica: Multi-Tenant LoRA Serving,” arXiv, arXiv:2310.18547v1 [cs.DC], Oct. 28, Oct. 2023, 13 pages. [cited by applicant]
Leviathan, et al., “Fast Inference from Transformers via Speculative Decoding,” arXiv, arXiv:2211.17192v2, May 18, 2023, 13 pages. [cited by applicant]
Xia, “Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation,” arXiv, arXiv:2203.16487v6 [cs.CL], Oct. 30, 2023, 17 pages. [cited by applicant]
Xia, et al., “Speculative Decoding: Lossless Speedup of Autoregressive Translation with Generalized Aggressive Decoding,” arXiv, arXiv:2203.16487v5 [cs.CL], Oct. 16, 2023, 24 pages. [cited by applicant]
U.S. Appl. No. 18/232,485, filed Aug. 10, 2023. [cited by applicant]
U.S. Appl. No. 18/244,229, filed Sep. 9, 2023. [cited by applicant]