IP Library Granted Patent US 12,547,872
Granted Patent B2
US 12,547,872 · App. 17/657,606 · Granted Feb 10, 2026

Machine learning model processing based on perplexity

Inventors: Bita Darvish Rouhani (Bellevue, WA); Douglas Christopher Burger (Bellevue, WA); Eric S Chung (Woodinville, WA)
Assignee: Microsoft Technology Licensing, LLC
G06N3/042G06N3/0495G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,547,872
App. No.
17/657,606
Granted
Feb 10, 2026
Kind
B2
Abstract

A method for operating a machine learning model is presented. The machine learning model includes a plurality of sequential transformer blocks. The method comprises receiving input data at a transformer block and processing the input data via a mixture of experts layer. At an auxiliary classifier, a measure of perplexity of the processed input data is determined. Based on the determined measure of perplexity, one or more experts in a downstream transformer block that will subsequently process the input data are indicated. Weight matrices are then fetched for the indicated one or more experts.

Claims (44)

1 . A method for operating a machine learning model including a plurality of sequential transformer blocks, comprising:

receiving input data at a transformer block;

processing at least a portion of the input data via a mixture of experts layer;

at an auxiliary classifier, determining a measure of perplexity of the processed input data;

based on the determined measure of perplexity, indicating one or more experts of a group of experts in a downstream transformer block that will subsequently process at least a portion of the processed input data;

fetching sparsified weight matrices for the indicated one or more experts of the group of experts to generate one or more sparsified experts, wherein a level of sparsity for the sparsified weight matrices is based on the determined measure of perplexity; and

further processing at least a portion of the processed input data via the one or more sparsified experts to generate a prediction.

2 . The method of claim 1 , wherein the downstream transformer block is a next transformer block.

3 . The method of claim 1 , further comprising:

indicating that the processed input data is likely to bypass one or more transformer blocks of the plurality of sequential transformer blocks.

4 . The method of claim 1 , further comprising:

performing a top k selection to select one of the one or more experts of the group of experts that will subsequently process the input data.

5 . The method of claim 1 , wherein the measure of perplexity is determined for a single input data shard.

6 . The method of claim 1 , wherein the measure of perplexity is determined for a group of input data shards.

7 . The method of claim 1 , wherein the measure of perplexity is a loss function.

8 . The method of claim 7 , wherein the loss function is a cross-entropy loss function.

9 . A method for operating a machine learning model, comprising:

at a mixture of experts layer, receiving input data comprising a plurality of input data shards;

sorting the input data shards into batches based on common modalities;

fetching sparsified weight matrices for one or more selected neural network experts of a group of neural network experts, to generate one or more sparsified selected neural network experts, the selected neural network experts trained in modalities represented in the batches;

scheduling each batch for processing by one or more sparsified selected neural network experts trained in a relevant modality; and

processing at least a portion of each batch via respective sparsified selected neural network experts to generate a prediction.

10 . The method of claim 9 , further comprising:

maintaining the fetched sparsified weight matrices for each sparsified selected neural network expert at a node based on a relevant batch processing schedule.

11 . The method of claim 10 , further comprising:

unloading the fetched sparsified weight matrices from the node following processing of a batch; and

fetching sparsified weights for a different selected neural network expert to be loaded onto the node.

12 . The method of claim 9 , wherein scheduling each batch for processing by one or more sparsified selected neural network experts trained in a relevant modality is performed by a reinforcement learning agent.

13 . The method of claim 12 , wherein the reinforcement learning agent is trained in load-balancing.

14 . A computing system, comprising:

one or more processors; and

a storage machine having instructions stored thereon executable by the one or more processors to instantiate a machine learning model, comprising:

a plurality of sequential transformer blocks configured to receive input data, each transformer block comprising:

a mixture of experts layer configured to process the input data; and

an auxiliary classifier configured to determining a measure of perplexity of the processed input data; and

wherein the one or more processors are configured to:

based on the determined measure of perplexity, indicate one or more experts of a group of experts in a downstream transformer block that will subsequently process at least a portion of the processed input data;

fetch sparsified weight matrices for the indicated one or more experts of the group of experts to generate one or more sparsified experts, wherein a level of sparsity for the sparsified weight matrices is based on the determined measure of perplexity; and

further process at least a portion of the processed input data via the one or more sparsified experts to generate a prediction.

15 . The computing system of claim 14 , wherein the one or more processors are further configured to:

indicate that the processed input data is likely to bypass one or more transformer blocks of the plurality of sequential transformer blocks.

16 . The computing system of claim 14 , wherein the measure of perplexity is determined for a group of input data shards.

17 . The computing system of claim 14 , wherein the measure of perplexity is a loss function.

18 . The method of claim 1 , wherein sparsified weight matrices fetched for experts indicated for processing processed input data with a low level of perplexity have a greater level of sparsity as compared to sparsified weight matrices fetched for experts indicated for processing processed input data with a high level of perplexity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2022
From: DARVISH ROUHANI, BITA; BURGER, DOUGLAS CHRISTOPHER; CHUNG, ERIC S
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 059465/0885 →
Continuity (1)
Related Publication 20230316043A1 · Oct 5, 2023
References Cited (19)
US 20190251423A1 · Shazeer · 2019 [cited by examiner]
US 20230316042A1 · Darvish Rouhani et al. · 2023 [cited by applicant]
Shazeer et al.“Outrageously Large Neural Networks:The Sparsely-Gated Mixture-Of-Experts Layer” dated Jan. 23, 2017 and retrieved from arXiv: 1701.06538v1 (Year: 2017). [cited by examiner]
Fedus, et al., “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”, In Repository of arXiv:2101.03961v1, Jan. 11, 2021, 31 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/010644”, Mailed Date: Apr. 20, 2023, 17 Pages. [cited by applicant]
Rajbhandari, et al., “DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale”, In Repository of arXiv:2201.05596v1, Jan. 14, 2022, 31 Pages. [cited by applicant]
Riquelme, et al., “Scaling Vision with Sparse Mixture of Experts”, In Repository of arXiv:2106.05974v1, Jun. 10, 2021, 43 Pages. [cited by applicant]
Lepikhin, et al., “GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding”, In Repository of arXiv:2006.16668v1, Jun. 30, 2020, 35 Pages. [cited by applicant]
Liao, et al., “Doubly Sparse: Sparse Mixture of Sparse Experts for Efficient Softmax Inference”, In Repository of arXiv:1901.10668v1, Jan. 30, 2019, 10 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/010639”, Mailed Date: May 12, 2023, 12 Pages. [cited by applicant]
Nie, et al., “EvoMoE: An Evolutional Mixture-of-Experts Training Framework via Dense-To-Sparse Gate”, arXiv:2112.14397v2, Oct. 9, 2022, 14 pages. [cited by applicant]
Non-Final Office Action mailed on Mar. 25, 2025, in U.S. Appl. No. 17/657,604 34 Pages. [cited by applicant]
Wang, et al., “Deep Mixture of Experts via Shallow Embedding”, Proceedings of Machine Learning Research, vol. 115, 2020, 11 Pages. [cited by applicant]
You, et al., “SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Expert”, Interspeech, Aug. 30, 2021, pp. 2077-2081. [cited by applicant]
Communication pursuant to Article 94(3) EPC Received for European Application No. 23705122.2, mailed on Sep. 5, 2025, 19 pages. [cited by applicant]
Final Office Action mailed on Oct. 16, 2025, in U.S. Appl. No. 17/657,604 39 Pages. [cited by applicant]
Goyal, et al., “A Multimodal Mixture-Of-Experts Model for Dynamic Emotion Prediction in Movies,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 2822-2826. [cited by applicant]
Yuksel, et al., “Twenty Years of Mixture of Experts”, In Journal of IEEE Transactions on Neural Networks and Learning Systems, vol. 23, Issue No. 8, Aug. 2012, pp. 1177-1193. [cited by applicant]
Communication Pursuant to Article 94(3) Received for European Application No. 23705118.0, mailed on Nov. 11, 2025, 06 pages. [cited by applicant]