IP Library › Granted Patent US 12,314,347
Granted Patent B2
US 12,314,347 · App. 17/703,552 · Granted May 27, 2025

Method and system of retrieving multimodal assets

Inventors: Adit Krishnan (Mountain View, CA); Ji Li (San Jose, CA); Amit Srivastava (San Jose, CA)
Assignee: Microsoft Technology Licensing, LLC
G06F18/256G06F16/24556G06F16/24578G06F16/248G06F16/9032G06F16/9038G06F18/21355G06F18/24147G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,347
App. No.
17/703,552
Granted
May 27, 2025
Kind
B2
Abstract

A system and method and for retrieving one or more one or more multimodal assets includes receiving a search query for searching for one or more multimodal assets from among a plurality of candidate multimodal assets, encoding the search query into one or more query embedding representations via a trained query representation machine-learning (ML) model, comparing, via a matching unit, the one or more query embedding representations to a plurality of multimodal tensor representations, each of the plurality of multimodal tensor representations being a representation of one of the plurality of candidate multimodal assets, and identifying, based on the comparison, at least one of the plurality of the candidate multimodal assets as a search result for the search query, and providing the at least one of the plurality of the candidate multimodal assets for display as the search result.

Claims (86)

1. A data processing system comprising:

a processor; and

a memory in communication with the processor, the memory comprising executable instructions that, when executed by the processor, cause the data processing system to perform:

receiving, via a query representation model, a search query for searching for one or more multimodal assets from among a plurality of candidate multimodal assets, wherein the one or more multimodal assets and the search query each includes multimodal content containing two or more different types of content including graphic or image content;

parsing, via the query representation model, the search query including the multimodal content;

identifying, based on the parsing, a first content type and a second content type in the search query, the second content type being a graphic or image content type;

transmitting the first content type to a first representation model to generate a first set of vector embeddings;

transmitting the second content type to a second representation model to generate a second set of vector embeddings;

transmitting the first and second sets of vector embeddings to a tensor generation unit to generate tensors based on the first and second sets of vector embeddings and to output a query tensor representation;

comparing, via a matching unit, the query tensor representation to a plurality of multimodal tensor representations, each of the plurality of multimodal tensor representations being a representation of one of the plurality of candidate multimodal assets; and

identifying, based on the comparing, at least one of the plurality of the candidate multimodal assets as a search result for the search query; and

providing the at least one of the plurality of the candidate multimodal assets for display as the search result.

2. The data processing system of claim 1 , wherein the graphic or image content comprises at least one of an image element, an icon element, a GIF element, and an illustration element.

3. The data processing system of claim 1 , wherein the multimodal tensor representation for each one of the plurality of candidate multimodal assets includes a vector embedding representation for each of a plurality of elements in each of the plurality of candidate multimodal assets.

4. The data processing system of claim 3 , wherein the executable instructions, when executed by the processor, further cause the data processing system to perform:

providing the plurality of candidate multimodal assets to a trained asset representation machine-learning (ML) model to generate the vector embedding representation for each of the plurality of elements;

receiving the vector embedding representation for each of the plurality of elements as an output from the trained asset representation ML model;

generating the multimodal tensor representation for each one of the plurality of candidate multimodal assets asserts from the received vector embedding representations; and

generating a summarized embedding representation for the multimodal tensor representation.

5. The data processing system of claim 4 , wherein generating the summarized embedding representation for the multimodal tensor representation includes creating one aggregate embedding representation for each different modality in each one of the plurality of candidate multimodal assets.

6. The data processing system of claim 4 , wherein the executable instructions, when executed by the processor, further cause the data processing system to perform:

generating a summarized query embedding representation for the query tensor representation;

applying an Approximate Nearest Neighbor (ANN) technique to the summarized query embedding representation to generate an indexed query representation; and

applying the ANN technique to a summarized tensor representation for each one of the plurality of candidate assets,

wherein comparing the one or more query embedding representations to the plurality of multimodal tensor representations includes:

applying an ANN search technique to the summarized tensor representation for each one of the plurality of candidate assets, based on the summarized query embedding representation, to generate a smaller candidate set of multimodal assets from among the plurality of candidate multimodal assets;

performing a tensor-to-tensor comparison between the query tensor and one or more tensors associated with the smaller candidate set of multimodal assets;

generating a similarity score between the query tensor and each of the multimodal assets in the smaller candidate set of multimodal assets; and

identifying the at least one of the plurality of multimodal assets as the search result based on the similarity score.

7. The data processing system of claim 4 , wherein the summarized embedding representation for the multimodal tensor representation has a fixed size.

8. A method for retrieving one or more multimodal assets comprising:

receiving, via a query representation model, a search query for searching for the one or more multimodal assets from among a plurality of candidate multimodal assets, wherein the one or more multimodal assets and the search query each includes multimodal content containing two or more different types of content including graphic or image content;

parsing, via the query representation model, the search query including the multimodal content;

identifying, based on the parsing, a first content type and a second content type in the search query, the second content type being a graphic or image content type;

transmitting the first content type to a first representation model to generate a first set of vector embeddings;

transmitting the second content type to a second representation model to generate a second set of vector embeddings;

transmitting the first and second sets of vector embeddings to a tensor generation unit to generate tensors based on the first and second sets of vector embeddings and to output a query tensor representation;

comparing, via a matching unit, the query tensor representation to a plurality of multimodal tensor representations, each of the plurality of multimodal tensor representations being a representation of one of the plurality of candidate multimodal asset; and

identifying, based on the comparing, at least one of the plurality of the candidate multimodal assets as a search result for the search query; and

providing the at least one of the plurality of the candidate multimodal assets for display as the search result.

9. The method of claim 8 , wherein the graphic or image content comprises at least one of an image element, an icon element, a GIF element, and an illustration element.

10. The method of claim 8 , wherein a multimodal tensor representation for each one of the plurality of candidate multimodal assets includes a vector embedding representation for each of a plurality of elements in each of the plurality of candidate multimodal assets.

11. The method of claim 10 , further comprising:

providing the plurality of candidate multimodal assets to a trained asset representation machine-learning (ML) model to generate the vector embedding representation for each of the plurality of elements;

receiving the vector embedding representation for each of the plurality of elements as an output from the trained asset representation ML model;

generating the multimodal tensor representation for each one of the plurality of candidate multimodal assets from the received vector embedding representations; and

generating a summarized embedding representation for the multimodal tensor representation.

12. The method of claim 11 , wherein generating the summarized embedding representation for the multimodal tensor representation includes creating one aggregate embedding representation for each different modality in each one of the plurality of candidate multimodal assets.

13. The method of claim 11 , further comprising:

generating a summarized query embedding representation for the query tensor representation;

applying an Approximate Nearest Neighbor (ANN) technique to the summarized query embedding representation to generate an indexed query representation; and

applying the ANN technique to a summarized tensor representation for each one of the plurality of candidate assets,

wherein comparing the one or more query embedding representations to the plurality of multimodal tensor representations includes:

applying an ANN search technique to the summarized tensor representation for each one of the plurality of candidate assets, based on the summarized query embedding representation, to generate a smaller candidate set of multimodal assets from among the plurality of candidate multimodal assets;

performing a tensor-to-tensor comparison between the query tensor and one or more tensors associated with the smaller candidate set of multimodal assets;

generating a similarity score between the query tensor and each of the multimodal assets in the smaller candidate set of multimodal assets; and

identifying the at least one of the plurality of multimodal assets as the search result based on the similarity score.

14. The method of claim 11 , wherein the summarized embedding representation for the multimodal tensor representation has a fixed size.

15. A non-transitory computer readable medium on which are stored instructions that, when executed, cause a programmable device to perform:

receiving, via a query representation model, a search query for searching for one or more multimodal assets from among a plurality of candidate multimodal assets, wherein the one or more multimodal asset and the search query each includes multimodal content containing two or more different types of content including graphic or image content;

parsing, via the query representation model, the search query including the multimodal content;

identifying, based on the parsing, a first content type and a second content type in the search query, the second content type being a graphic or image content type;

transmitting the first content type to a first representation model to generate a first set of vector embeddings;

transmitting the second content type to a second representation model to generate a second set of vector embeddings;

transmitting the first and second sets of vector embeddings to a tensor generation unit to generate tensors based on the first and second sets of vector embeddings and to output a query tensor representation;

comparing, via a matching unit, the query tensor representation to a plurality of multimodal tensor representations, each of the plurality of multimodal tensor representations being a representation of one of the plurality of candidate multimodal asset; and

identifying, based on the comparing, at least one of the plurality of the candidate multimodal assets as a search result for the search query; and

providing the at least one of the plurality of the candidate multimodal assets for display as the search result.

16. The non-transitory computer readable medium of claim 15 , wherein the graphic or image content comprises at least one of an image element, an icon element, a GIF element, and an illustration element.

17. The non-transitory computer readable medium of claim 15 , wherein a multimodal tensor representation for each one of the plurality of candidate multimodal assets includes a vector embedding representation for each of a plurality of elements in each of the plurality of candidate multimodal assets.

18. The non-transitory computer readable medium of claim 17 , further comprising instructions for:

providing the plurality of candidate multimodal assets to a trained asset representation machine-learning (ML) model to generate the vector embedding representation for each of the plurality of elements;

receiving the vector embedding representation for each of the plurality of elements as an output from the trained asset representation ML model;

generating the multimodal tensor representation for each one of the plurality of candidate multimodal assets from the received vector embedding representations; and

generating a summarized embedding representation for the multimodal tensor representation.

19. The non-transitory computer readable medium of claim 18 , wherein generating the summarized embedding representation for the multimodal tensor representation includes creating one aggregate embedding representation for each different modality in each one of the plurality of candidate multimodal assets.

20. The non-transitory computer readable medium of claim 18 , further comprising:

generating a summarized query embedding representation for the query tensor representation;

applying an Approximate Nearest Neighbor (ANN) technique to the summarized query embedding representation to generate an indexed query representation; and

applying the ANN technique to a summarized tensor representation for each one of the plurality of candidate assets,

wherein comparing the one or more query embedding representations to the plurality of multimodal tensor representations includes:

applying an ANN search technique to the summarized tensor representation for each one of the plurality of candidate assets, based on the summarized query embedding representation, to generate a smaller candidate set of multimodal assets from among the plurality of candidate multimodal assets;

performing a tensor-to-tensor comparison between the query tensor and one or more tensors associated with the smaller candidate set of multimodal assets;

generating a similarity score between the query tensor and each of the multimodal assets in the smaller candidate set of multimodal assets; and

identifying the at least one of the plurality of multimodal assets as the search result based on the similarity score.

21. The non-transitory computer readable medium of claim 18 , wherein the summarized embedding representation for the multimodal tensor representation has a fixed size.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2023
From: KRISHNAN, ADIT; LI, JI; SRIVASTAVA, AMIT
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062874/0080 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2022
From: KRISHNAN, ADIT; LI, JI; SRIVASTAVA, AMIT
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 059393/0001 →
Continuity (1)
Related Publication 20230306087A1 · Sep 28, 2023
References Cited (64)
US 8250613B2 · Faulkner et al. · 2012 [cited by applicant]
US 10782456B2 · Schürmann · 2020 [cited by applicant]
US 10783456B2 · Strope et al. · 2020 [cited by applicant]
US 11003856B2 · Kiros et al. · 2021 [cited by applicant]
US 11416534B2 · Wang · 2022 [cited by applicant]
US 11533495B2 · Jain et al. · 2022 [cited by applicant]
US 11620331B2 · Kislyuk · 2023 [cited by examiner]
US 11768837B1 · Newman · 2023 [cited by applicant]
US 20050257240A1 · Faulkner et al. · 2005 [cited by applicant]
US 20110047226A1 · Gabriel et al. · 2011 [cited by applicant]
US 20110307425A1 · Wang et al. · 2011 [cited by applicant]
US 20130166543A1 · MacDonald et al. · 2013 [cited by applicant]
US 20130166587A1 · Berry · 2013 [cited by applicant]
US 20150034357A1 · Dower et al. · 2015 [cited by applicant]
US 20150067541A1 · Owens et al. · 2015 [cited by applicant]
US 20150278226A1 · Franks · 2015 [cited by applicant]
US 20150347357A1 · Maughan et al. · 2015 [cited by applicant]
US 20160042050A1 · Chen et al. · 2016 [cited by applicant]
US 20160092447A1 · Venkataraman et al. · 2016 [cited by applicant]
US 20170006357A1 · Obara · 2017 [cited by applicant]
US 20170098283A1 · Rajan et al. · 2017 [cited by applicant]
US 20170337265A1 · Garrett et al. · 2017 [cited by applicant]
US 20180068023A1 · Douze · 2018 [cited by examiner]
US 20180113888A1 · Peña Muñoz · 2018 [cited by applicant]
US 20190007755A1 · Obara · 2019 [cited by applicant]
US 20190065492A1 · Cheng · 2019 [cited by examiner]
US 20190102397A1 · Hornkvist · 2019 [cited by applicant]
US 20190163766A1 · Gulati · 2019 [cited by applicant]
US 20190258713A1 · Kiros et al. · 2019 [cited by applicant]
US 20190258722A1 · Guo et al. · 2019 [cited by applicant]
US 20190303402A1 · Berry · 2019 [cited by applicant]
US 20190347556A1 · Yim et al. · 2019 [cited by applicant]
US 20200413154A1 · Obara · 2020 [cited by applicant]
US 20210133264A1 · Tiwari · 2021 [cited by applicant]
US 20210191925A1 · Sianez · 2021 [cited by applicant]
US 20210264203A1 · Fuxman et al. · 2021 [cited by applicant]
US 20210295822A1 · Tomkins · 2021 [cited by applicant]
US 20220075961A1 · Cavallari · 2022 [cited by applicant]
US 20220138170A1 · Misiewicz · 2022 [cited by applicant]
US 20220156298A1 · Mahmoud · 2022 [cited by examiner]
US 20220245706A1 · Chaidaroon et al. · 2022 [cited by applicant]
US 20230169110A1 · Li et al. · 2023 [cited by applicant]
US 20230244727A1 · Liu · 2023 [cited by applicant]
US 20230325391A1 · Li et al. · 2023 [cited by applicant]
US 20240248901A1 · Krishnan · 2024 [cited by applicant]
EP 3794836A1 · 2021 [cited by applicant]
WO 2020051249A1 · 2020 [cited by applicant]
“Non Final Office Action Issued in U.S. Appl. No. 17/538,880”, Mailed Date: Jun. 13, 2023, 15 Pages. [cited by applicant]
Final Office Action mailed on Jan. 8, 2024, in U.S. Appl. No. 17/716,653, 28 pages. [cited by applicant]
U.S. Appl. No. 18/158,121, filed Jan. 23, 2023. [cited by applicant]
Non-Final Office Action mailed on May 21, 2024, in U.S. Appl. No. 17/716,653, 38 pages. [cited by applicant]
Notice of Allowance mailed on Mar. 18, 2024, in U.S. Appl. No. 17/538,880, 10 pages. [cited by applicant]
Hassan, et al., “Multi-Modal Information Integration for Document Retrieval”, In Proceedings of 12th International Conference On Document Analysis And Recognition, Aug. 25, 2013, pp. 1200-1204. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/010536”, Mailed Date: Apr. 26, 2023, 10 Pages. [cited by applicant]
Chi, et al., “Zero-Shot Cross-Media Embedding Learning With Dual Adversarial Distribution Network”, In Journal of IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, Issue 4, Apr. 3, 2020, pp. 1173-… [cited by applicant]
Lin, et al., “Learning Cross-Aligned Latent Embeddings for Zero-Shot Cross-Modal Retrieval”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, Issue 7, Feb. 7, 2020, pp. 11515-11522. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/011566”, Mailed Date: May 25, 2023, 10 Pages. [cited by applicant]
“Non Final Office Action Issued in U.S. Appl. No. 17/716,653”, Mailed Date: Jul. 14, 2023, 21 Pages. [cited by applicant]
Final Office Action mailed on Nov. 9, 2023, in U.S. Appl. No. 17/538,880, 13 pages. [cited by applicant]
“Application as Filed in U.S. Appl. No. 17/538,880”, Filed Date: Nov. 30, 2021, 41 Pages [cited by applicant]
U.S. Appl. No. 17/716,653, filed Apr. 8, 2022. [cited by applicant]
Non-Final Office Action mailed on Jul. 24, 2024, in U.S. Appl. No. 18/158,121, 74 pages. [cited by applicant]
Notice of Allowance mailed on Sep. 30, 2024, in U.S. Appl. No. 17/716,653, 10 pages. [cited by applicant]
Notice of Allowance mailed on Feb. 12, 2025, in U.S. Appl. No. 18/158,121, 15 pages. [cited by applicant]