IP Library › Granted Patent US 11,836,175
Granted Patent B1
US 11,836,175 · App. 17/853,273 · Granted Dec 5, 2023

Systems and methods for semantic search via focused summarizations

Inventors: Itzik Malkiel (Ramat Gan, IL); Noam Koenigstein (Tel Aviv, IL); Oren Barkan (Tel Aviv, IL); Jonathan Ephrath (Tel Aviv, IL); Yonathan Weill (Tel Aviv, IL); Nir Nice (Tel Aviv, IL)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G06F16/3347G06F16/345
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,836,175
App. No.
17/853,273
Granted
Dec 5, 2023
Kind
B1
Abstract

Semantic search techniques via focused summarizations are described. For example, a search query is received for a text-based content item in a data set comprising a plurality of text-based content items. A first feature vector representative of the search query is obtained. A respective semantic similarity score is determined between the first feature vector and each of a plurality of second feature vectors. Each of the second feature vectors is representative of a machine-generated summarization of a respective text-based content item. The machine-generated summarization comprises a plurality of multi-word fragments that are selected from the respective text-based content item via a transformer-based machine learning model. A search result is provided responsive to the search query. The search result comprises a subset of the plurality of text-based content items associated with a respective second feature vector having a semantic similarity score that has a predetermined relationship with a predetermined threshold value.

Claims (81)

1. A system, comprising:

at least one processor circuit; and

at least one memory that stores program code that, when executed by the at least one processor circuit, performs operations, the operations comprising:

receiving a search query for a first text-based content item in a data set comprising a first plurality of text-based content items;

obtaining a first feature vector representative of the search query;

determining a respective semantic similarity score between the first feature vector and each of a plurality of second feature vectors generated by a transformer-based machine learning model, each of the second feature vectors representative of a machine-generated summarization of a respective first text-based content item of the first plurality of text-based content items, the machine-generated summarization comprising a first plurality of multi-word fragments that are selected from the respective first text-based content item, each machine-generated summarization generated by:

extracting a second plurality of multi-word fragments of text from the respective first text-based content item;

determining importance scores for the second plurality of multi-word fragments based on a similarity matrix;

ranking the second plurality of multi-word fragments based on the importance scores;

selecting a subset of multi-word fragments from the second plurality of multi-word fragments having an N highest importance scores; and

generating the summarization based on sorting the subset; and

providing a search result comprising a subset of the first plurality of text-based content items associated with a respective second feature vector having a semantic similarity score that has a predetermined relationship with a predetermined threshold value.

2. The system of claim 1 , the operations further comprising:

providing, as a first input to the transformer-based machine learning model, first pairs of multi-word fragments of text, wherein each of the first pairs are selected from a same text-based content item of a training data set comprising a second plurality of text-based content items;

providing, as a second input to the transformer-based machine learning model, second pairs of multi-word fragments of text, wherein, for each second pair, a first multi-word fragment of text of the second pair is from a first text-based content item of the training data set and a second multi-word fragment of text of the second pair is from a second text-based content item of the training data set that is different than the first text-based content item; and

training the transformer-based machine learning model based on the first pairs and the second pairs such that the transformer-based machine learning model generates a feature vector representative of a multi-word fragment of text provided as an input thereto.

3. The system of claim 2 , wherein each of the first pairs are from within a same paragraph of the same text-based content item of the training data set.

4. The system of claim 2 , wherein the summarization of the respective first text-based content item is further generated by:

providing a representation of each multi-word fragment of the second plurality of multi-word fragments as an input to the transformer-based machine learning model, the transformer-based machine learning model generating a feature vector representation for each multi-word fragment of the second plurality of multi-word fragments;

generating the similarity matrix to store, for each of the plurality of multi-word fragment pairs of the second plurality of multi-word fragments, a semantic similarity score representative of a semantic similarity between the multi-word fragment pair based on the feature vector representations generated for the multi-word fragment pair;

sorting the subset of multi-word fragments in an order in which each multi-word fragment of the subset of multi-word fragments appears in the respective first text-based content item; and

wherein N is a positive integer.

5. The system of claim 4 , wherein the semantic similarity score for a respective multi-word fragment pair of the second plurality of multi-word fragment pairs is based on a cosine similarity between the respective feature vectors representative of the respective multi-word fragment pair that are generated by the transformer-based machine learning model.

6. The system of claim 4 , wherein the second feature vector representative of the summarization of the respective first text-based content item is generated by:

for each multi-word fragment of the sorted subset:

providing a representation of the multi-word fragment of the sorted subset as an input to the transformer-based machine learning model, the transformer-based machine learning model generating a feature vector representation of the multi-word fragment of the sorted subset; and

combining each feature vector generated by the transformer-based machine learning model for the sorted subset to generate the second feature vector representative of the summarization of the respective first text-based content item.

7. The system of claim 2 , wherein obtaining the first feature vector representative of the search query comprises:

providing a representation of the search query as an input to the transformer-based machine learning model, the transformer-based machine learning model generating the first feature vector based on the representation of the search query.

8. A method, comprising:

receiving a search query for a first text-based content item in a data set comprising a first plurality of text-based content items;

obtaining a first feature vector representative of the search query;

determining a respective semantic similarity score between the first feature vector and each of a plurality of second feature vectors generated by a transformer-based machine learning model, each of the second feature vectors representative of a machine-generated summarization of a respective first text-based content item of the first plurality of text-based content items, the machine-generated summarization comprising a first plurality of multi-word fragments that are selected from the respective first text-based content item, each machine-generated summarization generated by:

extracting a second plurality of multi-word fragments of text from the respective first text-based content item;

determining importance scores for the second plurality of multi-word fragments based on a similarity matrix;

ranking the second plurality of multi-word fragments based on the importance scores;

selecting a subset of multi-word fragments from the second plurality of multi-word fragments having an N highest importance scores; and

generating the summarization based on sorting the subset; and

providing a search result comprising a subset of the first plurality of text-based content items associated with a respective second feature vector having a semantic similarity score that has a predetermined relationship with a predetermined threshold value.

9. The method of claim 8 , further comprising:

providing, as a first input to the transformer-based machine learning model, first pairs of multi-word fragments of text, wherein each of the first pairs are selected from a same text-based content item of a training data set comprising a second plurality of text-based content items;

providing, as a second input to the transformer-based machine learning model, second pairs of multi-word fragments of text, wherein, for each second pair, a first multi-word fragment of text of the second pair is from a first text-based content item of the training data set and a second multi-word fragment of text of the second pair is from a second text-based content item of the training data set that is different than the first text-based content item; and

training the transformer-based machine learning model based on the first pairs and the second pairs such that the transformer-based machine learning model generates a feature vector representative of a multi-word fragment of text provided as an input thereto.

10. The method of claim 9 , wherein each of the first pairs are from within a same paragraph of the same text-based content item of the training data set.

11. The method of claim 9 , wherein the summarization of the respective first text-based content item is further generated by:

providing a representation of each multi-word fragment of the second plurality of multi-word fragments as an input to the transformer-based machine learning model, the transformer-based machine learning model generating a feature vector representation for each multi-word fragment of the second plurality of multi-word fragments;

generating the similarity matrix to store, for each of the plurality of multi-word fragment pairs of the second plurality of multi-word fragments, a semantic similarity score representative of a semantic similarity between the multi-word fragment pair based on the feature vector representations generated for the multi-word fragment pair;

sorting the subset of multi-word fragments in an order in which each multi-word fragment of the subset of multi-word fragments appears in the respective first text-based content item; and

wherein N is a positive integer.

12. The method of claim 11 , wherein the semantic similarity score for a respective multi-word fragment pair of the second plurality of multi-word fragment pairs is based on a cosine similarity between the respective feature vectors representative of the respective multi-word fragment pair that are generated by the transformer-based machine learning model.

13. The method of claim 11 , wherein the second feature vector representative of the summarization of the respective first text-based content item is generated by:

for each multi-word fragment of the sorted subset:

providing a representation of the multi-word fragment of the sorted subset as an input to the transformer-based machine learning model, the transformer-based machine learning model generating a feature vector representation of the multi-word fragment of the sorted subset; and

combining each feature vector generated by the transformer-based machine learning model for the sorted subset to generate the second feature vector representative of the summarization of the respective first text-based content item.

14. The method of claim 9 , wherein obtaining the first feature vector representative of the search query comprises:

providing a representation of the search query as an input to the transformer-based machine learning model, the transformer-based machine learning model generating the first feature vector based on the representation of the search query.

15. A computer-readable storage medium having program instructions recorded thereon that, when executed by at least one processing circuit, perform a method, the method comprising:

receiving a search query for a first text-based content item in a data set comprising a first plurality of text-based content items;

obtaining a first feature vector representative of the search query;

determining a respective semantic similarity score between the first feature vector and each of a plurality of second feature vectors generated by a transformer-based machine learning model, each of the second feature vectors representative of a machine-generated summarization of a respective first text-based content item of the first plurality of text-based content items, the machine-generated summarization comprising a first plurality of multi-word fragments that are selected from the respective first text-based content item, each machine-generated summarization generated by:

extracting a second plurality of multi-word fragments of text from the respective first text-based content item;

determining importance scores for the second plurality of multi-word fragments based on a similarity matrix;

ranking the second plurality of multi-word fragments based on the importance scores;

selecting a subset of multi-word fragments from the second plurality of multi-word fragments having an N highest importance scores; and

generating the summarization based on sorting the subset; and

providing a search result comprising a subset of the first plurality of text-based content items associated with a respective second feature vector having a semantic similarity score that has a predetermined relationship with a predetermined threshold value.

16. The computer-readable storage medium of claim 15 , the method further comprising:

providing, as a first input to the transformer-based machine learning model, first pairs of multi-word fragments of text, wherein each of the first pairs are selected from a same text-based content item of a training data set comprising a second plurality of text-based content items;

providing, as a second input to the transformer-based machine learning model, second pairs of multi-word fragments of text, wherein, for each second pair, a first multi-word fragment of text of the second pair is from a first text-based content item of the training data set and a second multi-word fragment of text of the second pair is from a second text-based content item of the training data set that is different than the first text-based content item; and

training the transformer-based machine learning model based on the first pairs and the second pairs such that the transformer-based machine learning model generates a feature vector representative of a multi-word fragment of text provided as an input thereto.

17. The computer-readable storage medium of claim 16 , wherein each of the first pairs are from within a same paragraph of the same text-based content item of the training data set.

18. The computer-readable storage medium of claim 16 , wherein the summarization of the respective first text-based content item is further generated by:

providing a representation of each multi-word fragment of the second plurality of multi-word fragments as an input to the transformer-based machine learning model, the transformer-based machine learning model generating a feature vector representation for each multi-word fragment of the second plurality of multi-word fragments;

generating the similarity matrix to store, for each of the plurality of multi-word fragment pairs of the second plurality of multi-word fragments, a semantic similarity score representative of a semantic similarity between the multi-word fragment pair based on the feature vector representations generated for the multi-word fragment pair;

sorting the subset of multi-word fragments in an order in which each multi-word fragment of the subset of multi-word fragments appears in the respective first text-based content item; and

wherein N is a positive integer.

19. The computer-readable storage medium of claim 18 , wherein the semantic similarity score for a respective multi-word fragment pair of the second plurality of multi-word fragment pairs is based on a cosine similarity between the respective feature vectors representative of the respective multi-word fragment pair that are generated by the transformer-based machine learning model.

20. The computer-readable storage medium of claim 18 , wherein the second feature vector representative of the summarization of the respective first text-based content item is generated by:

for each multi-word fragment of the sorted subset:

providing a representation of the multi-word fragment of the sorted subset as an input to the transformer-based machine learning model, the transformer-based machine learning model generating a feature vector representation of the multi-word fragment of the sorted subset; and

combining each feature vector generated by the transformer-based machine learning model for the sorted subset to generate the second feature vector representative of the summarization of the respective first text-based content item.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2022
From: MALKIEL, ITZIK; KOENIGSTEIN, NOAM; BARKAN, OREN; EPHRATH, JONATHAN; WEILL, YONATHAN; NICE, NIR
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 060357/0292 →
Cited By (2)
US 12,271,411 US 12,505,141