IP Library › Granted Patent US 12,481,530
Granted Patent B2
US 12,481,530 · App. 18/072,735 · Granted Nov 25, 2025

Performing computing tasks using decoupled models for different data types

Inventors: Eric Chris Wolfgang Sommerlade (Oxford, GB); Mohsen Fayyaz (Berlin, DE); Nazuk Jain (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F9/5011
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,530
App. No.
18/072,735
Granted
Nov 25, 2025
Kind
B2
Abstract

A technique executes tasks using a data store of machine-trained models. The data store specifically includes a subset of encoder-type machine-trained models for converting input data items having different input data types into respective embeddings in a vector space, and a subset of decoder-type machine-trained models for converting embeddings in the same vector space into data items having respective different output data types. When executing a particular task that involves one or more data types, the technique selects one or more machine-trained models that match those data types. In some implementations, the technique provides a clipboard store for storing embeddings produced by the encoder-type machine-trained models and consumable by the decoder-type machine-trained models. The technique includes provisions for ensuring that any decoder-type machine-model is capable of processing embeddings produced by different versions of the encoder-type machine-trained models.

Claims (52)

1 . A computer-implemented method comprising:

receiving an input data item having a particular input data type;

executing at least part of a task that includes an instruction of pasting the input data item into a target item having a particular output data type that is different than the input data type, comprising:

executing a particular encoder-type machine-trained model to convert the input data item having the particular input data type to a particular embedding,

wherein the particular encoder-type machine-trained model is one of a subset of encoder-type machine-trained models that map input data items having different input data types to respective embeddings in a single shared vector space;

storing the particular embedding in a clipboard store;

executing a particular decoder-type machine-trained model that is associated with the particular output data type, to convert the particular embedding to an output data item of the particular output data type of the target item,

wherein the particular decoder-type machine-trained model is one of a subset of decoder-type machine-trained models that map the embeddings in the same single shared vector space to respective output data items having different output data types; and

pasting the output data item into the target item.

2 . The method of claim 1 , wherein the method is executed, at least in part, by a control system of a computing system.

3 . The method of claim 2 , wherein the control system also manages versions of the subset of encoder-type machine-trained models and the subset of decoder-type machine-trained models.

4 . The method of claim 1 , wherein the embeddings output by the encoder-type machine-trained models and input to the decoder-type machine-trained models are distributed-representation vectors mapped to the single shared vector space, and wherein a distance between any two vectors in the vector space reflects an extent of similarity between the two vectors.

5 . The method of claim 4 , wherein the encoder-type machine-trained models and the decoder-type machine-trained models have been trained to produce or consume embeddings in the single shared vector space.

6 . The method of claim 1 , wherein the clipboard store also includes a predecessor embedding produced by an earlier version of the particular encoder-type machine-trained model for the input data item, prior to generating the particular embedding.

7 . The method of claim 6 , wherein the particular embedding includes a base part that matches information in the predecessor embedding, and another part that includes information that is not present in the predecessor embedding.

8 . The method of claim 1 ,

wherein the method further includes generating output information for presentation in a user interface presentation that represents contents of the clipboard store, and

wherein the output information includes an image that conveys semantic contents of the particular embedding, for presentation in the user interface presentation together with a representation of the particular embedding.

9 . The method of claim 1 , wherein a given decoder-type machine-trained model in the subset of decoder-type machine-trained models operates on two or more input data items, the two or more input data items including a particular embedding that expresses semantic content of a particular input data item.

10 . The method of claim 9 , wherein another of the two or more input data items is an image mask that identifies a portion of the particular input data item.

11 . The method of claim 1 , further comprising:

generating a supplemental item in addition to the particular embedding; and

storing the supplemental item in the clipboard store along with the particular embedding.

12 . The method of claim 11 , wherein the supplemental item is an instance of randomly-generated noise information.

13 . The method of claim 11 , wherein the particular decoder-type machine-trained produces the output data item based on the particular embedding and the supplemental item.

14 . The method of claim 1 , wherein at least one of the decoder-type machine-trained models of the subset of decoder-type machine-trained models receives two or more input data items from the clipboard store, and maps the two more input data items into a single output data item.

15 . The method of claim 1 , further comprising storing the input data item in the clipboard store prior to producing the particular embedding.

16 . The method of claim 1 , wherein at least one entry in the clipboard store provides an identifier for an embedding and an indication of a source data item from which the embedding originated.

17 . A computing system, comprising:

a processing system for executing machine-readable instructions; and

a storage device for storing the machine-readable instructions,

the storage device also storing a set of machine-trained models, the set of machine-trained models including a subset of encoder-type machine-trained models that map input data items having different input data types to respective embeddings in a single shared vector space, and a subset of decoder-type machine-trained models that map the embeddings in the single shared vector space to respective output data items having different output data types, wherein any of the subset of encoder-type machine-trained models is capable of producing an embedding that is consumable by any of the subset of subset of decoder-type machine-trained models,

the processing system executing the machine-readable instructions for:

receiving an input data item having a particular input data type;

executing at least part of a task that includes an instruction of pasting the input data item into a target item having a particular output data type that is different than the input data type, comprising:

executing a particular encoder-type machine-trained model of the subset of decoder-type machine-trained models to convert the input data item having the particular input data type to a particular embedding;

storing the particular embedding in a clipboard store;

executing a particular decoder-type machine-trained model, of the subset of decoder-type machine-trained models, that is associated with the particular output data type, to convert the particular embedding to an output data item of the particular output data type of the target item; and

pasting the output data item into the target item.

18 . The computing system of claim 17 , wherein the embeddings output by the encoder-type machine-trained models and input to the decoder-type machine-trained models are distributed-representation vectors mapped to the single shared vector space, and wherein a distance between any two vectors in the vector space reflects an extent of similarity between the two vectors.

19 . A non-transitory computer-readable storage medium for storing computer-readable instructions, wherein a processing system executing the computer-readable instructions performs operations comprising:

receiving an input data item having a particular input data type;

executing at least part of a task that includes an instruction of pasting the input data item into a target item having a particular output data type that is different than the input data type, comprising:

executing a particular encoder-type machine-trained model to convert the input data item having the particular input data type to a particular embedding,

wherein the particular encoder-type machine-trained model is one of a subset of encoder-type machine-trained models stored in the computer-readable storage medium that map input data items having different input data types to respective embeddings in a single shared vector space;

storing the particular embedding in a clipboard store;

executing a particular decoder-type machine-trained model that is associated with the particular output data type, to convert the particular embedding to an output data item of the particular output data type of the target item,

wherein the particular decoder-type machine-trained model is one of a subset of decoder-type machine-trained models stored in the computer-readable storage medium that map the embeddings in the same single shared vector space to respective output data items having different output data types; and

pasting the output data item into the target item.

20 . The non-transitory computer-readable storage medium of claim 19 ,

wherein the clipboard store also includes a predecessor embedding produced by an earlier version of the particular encoder-type machine-trained model for the input data item, prior to generating the particular embedding, and

wherein the particular embedding includes a base part that matches information in the predecessor embedding, and another part that includes information that is not present in the predecessor embedding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2022
From: SOMMERLADE, ERIC CHRIS WOLFGANG; FAYYAZ, MOHSEN; JAIN, NAZUK
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061931/0721 →
Continuity (1)
Related Publication 20240184629A1 · Jun 6, 2024
References Cited (71)
US 11734242B1 · Bouyarmane · 2023 [cited by examiner]
US 20020147694A1 · Dempsey · 2002 [cited by examiner]
US 20020147754A1 · Dempsey · 2002 [cited by examiner]
US 20180307745A1 · Bachrach · 2018 [cited by examiner]
US 20190005374A1 · Shankar · 2019 [cited by examiner]
US 20190197121A1 · Jeon · 2019 [cited by examiner]
US 20190318040A1 · Chaudhury · 2019 [cited by examiner]
US 20190370616A1 · Eser · 2019 [cited by examiner]
US 20190371476A1 · Hernandez · 2019 [cited by examiner]
US 20200074246A1 · Goyal · 2020 [cited by examiner]
US 20200082807A1 · Kim · 2020 [cited by examiner]
US 20200184316A1 · Kavukcuoglu · 2020 [cited by examiner]
US 20200234725A1 · Garbacea · 2020 [cited by examiner]
US 20200311542A1 · Wang · 2020 [cited by examiner]
US 20200327963A1 · Ui · 2020 [cited by examiner]
US 20200342643A1 · Gouws · 2020 [cited by examiner]
US 20200372621A1 · Naruniec · 2020 [cited by examiner]
US 20200394998A1 · Kim · 2020 [cited by examiner]
US 20210056980A1 · Gfeller · 2021 [cited by examiner]
US 20210081673A1 · Lai · 2021 [cited by examiner]
US 20210109966A1 · Ayush · 2021 [cited by examiner]
US 20210133539A1 · Srivastava · 2021 [cited by examiner]
US 20210142783A1 · Kim · 2021 [cited by examiner]
US 20210287138A1 · Chang · 2021 [cited by examiner]
US 20210326751A1 · Liu · 2021 [cited by examiner]
US 20210390400A1 · Yao · 2021 [cited by examiner]
US 20210394784A1 · Blaiotta · 2021 [cited by examiner]
US 20220067408A1 · Sheu · 2022 [cited by examiner]
US 20220198715A1 · Wilczynski · 2022 [cited by examiner]
US 20220207362A1 · Meyerson · 2022 [cited by examiner]
US 20220230625A1 · Zhu · 2022 [cited by examiner]
US 20220237389A1 · Dognin · 2022 [cited by examiner]
US 20220262002A1 · Wang · 2022 [cited by examiner]
US 20220284551A1 · Neofytou · 2022 [cited by examiner]
US 20220319500A1 · Lee · 2022 [cited by examiner]
US 20220351041A1 · Lee · 2022 [cited by examiner]
US 20220368356A1 · Luo · 2022 [cited by examiner]
US 20220414429A1 · Torrado · 2022 [cited by examiner]
US 20230019211A1 · Wang · 2023 [cited by examiner]
US 20230082079A1 · Douillard · 2023 [cited by examiner]
US 20230153611A1 · Mathur · 2023 [cited by examiner]
US 20230162758A1 · Borgstrom · 2023 [cited by examiner]
US 20230204721A1 · Meyer · 2023 [cited by examiner]
US 20230206898A1 · Stanton · 2023 [cited by examiner]
US 20230252268A1 · Giovannini · 2023 [cited by examiner]
US 20240020048A1 · Bert · 2024 [cited by examiner]
US 20240104719A1 · Gulsun · 2024 [cited by examiner]
US 20240135683A1 · Li · 2024 [cited by examiner]
US 20240152749A1 · Shanahan · 2024 [cited by examiner]
Molino, et al., “Ludwig: a type-based declarative deep learning toolbox,” arXiv, arXiv:1909.07930v1 [cs.LG], Sep. 17, 2019, 15 pages. [cited by applicant]
Molino, et al., “COTA: Improving the Speed and Accuracy of Customer Support through Ranking and Deep Networks,” arXiv, arXiv:1807.01337v1 [cs.LG], Jul. 3, 2018, 9 pages. [cited by applicant]
Lendave, Vijaysinh, “LUDWIG: a Type-Based Declarative Deep Learning Toolbox,” available at https://web.archive.org/web/20211229132112/https://analyticsindiamag.com/ludwig-a-type-based-declarative-deep-learning-toolbox/,… [cited by applicant]
PCT Search Report and Written Report in PCT/US2023/035554, mailing date Feb. 12, 2024, 15 pages. [cited by applicant]
Xu, et al., “AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks,” arXiv, Cornell University, arXiv:1711.10485v1 [cs.CV], Nov. 28, 2017, 9 pages. [cited by applicant]
Agnese, et al., “A Survey and Taxonomy of AdversarialNeural Networks for Text-to-Image Synthesis,” arXiv, Cornell University, arXiv:1910.09399v1 [cs.CV], Oct. 21, 2019, 26 pages. [cited by applicant]
Dosovitskiy, et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” arXiv, Cornell University, arXiv:2010.11929v2cs.CV], Jun. 3, 2021, 22 pages. [cited by applicant]
Ramesh, et al., “Zero-Shot Text-to-Image Generation,” arXiv, Cornell University, arXiv:2102.12092v2 [cs.CV], Feb. 26, 2021, 20 pages. [cited by applicant]
Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision,” arXiv, Cornell University, arXiv:2103.00020v1 [cs.CV], Feb. 26, 2021 48 pages. [cited by applicant]
Nichol, et al., “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models,” arXiv, Cornell University, arXiv:2112.10741v3 [cs.CV], Mar. 8, 2022, 20 pages. [cited by applicant]
Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models,” arXiv, Cornell University, arXiv:2112.10752v2 [cs.CV], Apr. 13, 2022, 45 pages. [cited by applicant]
Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents,” arXiv, Cornell University, arXiv:2204.06125v1 [cs.CV], Apr. 13, 2022, 27 pages. [cited by applicant]
Luo, Calvin, “Understanding Diffusion Models: a Unified Perspective,” arXiv, Cornell University, arXiv:2208.11970v1 [cs.LG], Aug. 25, 2022, 23 pages. [cited by applicant]
Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” arXiv, Cornell University, arXiv:2205.11487v1 [cs.CV], May 23, 2022, 46 pages. [cited by applicant]
O'Connor, Ryan, “How Imagen Actually Works,” available at https://www.assemblyai.com/blog/how-imagen-actually-works/, Assembly AI, Jun. 23, 2022, 19 pages. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv, Cornell University, arXiv:1810.04805v2 [cs.CL], May 24, 2019, 16 pages. [cited by applicant]
Vaswani, et al., “Attention Is All You Need,” arXiv, Cornell University, arXiv:1706.03762v5 [cs.CL], Dec. 6, 2017 15 pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners,” arXiv, Cornell University, arXiv:2005.14165v4 [cs.CL], Jul. 22, 2020, 75 pages. [cited by applicant]
Royal Skies, “Stable Diffusion Seeds: (EXPLAINED !!),” available at https://www.youtube.com/watch?v=bu2E122YYmo, printout of a YouTube page, playback time 0:00, accessed on Nov. 30, 2022, 1 page. [cited by applicant]
Redmond, et al., “You Only Look Once: Unified, Real-Time Object Detection,” arXiv, Cornell University, arXiv:1506.02640v5 [cs.CV], May 9, 2016, 10 pages. [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition,” arXiv, Cornell University, arXiv:1512.03385v1 [cs.CV], Dec. 10, 2015, 12 pages. [cited by applicant]
International Preliminary Report on Patentability for PCT Application No. PCT/US23/035554, mailed on Jun. 12, 2025, 9 pages. [cited by applicant]