IP Library Granted Patent US 10,148,777
Granted Patent B2
US 10,148,777 · App. 15/162,190 · Granted Dec 4, 2018

Entity based search retrieval and ranking

Inventors: Dhruv Arya (Sunnyvale, CA); Abhimanyu Lad (San Mateo, CA); Shakti Dhirendraji Sinha (Sunnyvale, CA); Satya Pradeep Kanduri (Mountain View, CA)
Assignee: Microsoft Technology Licensing, LLC
H04L67/22G06F17/30672G06N99/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,148,777
App. No.
15/162,190
Granted
Dec 4, 2018
Kind
B2
Abstract

In an example embodiment, one or more query terms are obtained. Then, for each of the one or more query terms, a standardized entity taxonomy is searched to locate a standardized entity that most closely matches the query term, with the standardized entity taxonomy comprising an entity identification for each of a plurality of different standardized entities. A confidence score is then calculated for the query term-standardized entity pair for the standardized entity that most closely matches the query term, and the query term is tagged with the entity identification corresponding to the standardized entity that most closely matches the query term and the calculated confidence score.

Claims (39)

1. A computer-implemented method, comprising:

obtaining one or more query terms from a search query input by a user;

for each of the one or more query terms:

searching a standardized entity taxonomy to locate a standardized entity that most closely matches the query term, the standardized entity taxonomy comprising an entity identification for each of a plurality of different standardized entities;

calculating a confidence score for the query term-standardized entity pair for the standardized entity that most closely matches the query term;

tagging each query term in the search query with the entity identification corresponding to the standardized entity that most closely matches the query term and the calculated confidence score; and

augmenting the search query with the entity identification corresponding to the standardized entity that most closely matches the query term along with an OR operator between the query term and the entity identification corresponding to the standardized entity.

2. The method of claim 1 , wherein the method further comprises, for each of the one or more query terms, augmenting the search query with a standardized entity and corresponding entity identification for a standardized entity indicated as a synonym for the query term.

3. The method of claim 1 , wherein the method further comprises:

eliminating any standardized entity from the search query that has a corresponding confidence score that does not transgress a preset threshold.

4. The method of claim 1 , wherein the confidence score indicates a statistical likelihood that a user specifying the query term in a search query would have, under ideal circumstances, also entered the corresponding standardized entity in the search query.

5. The method of claim 4 , wherein the confidence score is calculated by using a confidence score model trained via a machine learning algorithm based on member profiles and member activities in a social networking service.

6. The method of claim 5 , wherein the confidence score model is trained based on a statistical analysis of how often users who specify the query term in a search query click on a subsequent result containing the corresponding standardized entity.

7. The method of claim 5 , wherein the confidence score model is trained based on a statistical analysis of how often member profiles listing the query term also list the standardized entity.

8. The method of claim 1 , further comprising utilizing the tagged one or more query terms during indexing of a new document containing the one or more query terms.

9. The method of claim 1 , further comprising utilizing the tagged one or more query terms in ranking search results upon execution of a search query containing the one or more query terms.

10. A system comprising:

a non-transitory computer-readable medium having instructions stored thereon, which, when executed by a processor, cause the system to:

obtain one or more query terms from a search query input by a user;

for each of the one or more query terms:

search a standardized entity taxonomy to locate a standardized entity that most closely matches the query term, the standardized entity taxonomy comprising an entity identification for each of a plurality of different standardized entities;

calculate a confidence score for the query term-standardized entity pair for the standardized entity that most closely matches the query term;

tag each query term in the search query with the entity identification corresponding to the standardized entity that most closely matches the query term and the calculated confidence score; and

augment the search query with the entity identification corresponding to the standardized entity that most closely matches the query term along with an OR operator between the query term and the entity identification corresponding to the standardized entity.

11. The system of claim 10 , wherein the instructions, when executed by the processor, further cause the system to, for each of the one or more query terms, augment the search query with a standardized entity and corresponding entity identification for a standardized entity indicated as a synonym for the query term.

12. The system of claim 10 , wherein the instructions, when executed by the processor, further cause the system to:

eliminate any standardized entity from the search query that has a corresponding confidence score that does not transgress a preset threshold.

13. A non-transitory machine-readable storage medium comprising instructions, which when implemented by one or more machines, cause the one or more machines to perform operations comprising:

obtaining one or more query terms from a search query input by a user;

for each of the one or more query terms:

searching a standardized entity taxonomy to locate a standardized entity that most closely matches the query term, the standardized entity taxonomy comprising an entity identification for each of a plurality of different standardized entities;

calculating a confidence score for the query term-standardized entity pair for the standardized entity that most closely matches the query term;

tagging each query term in the search query with the entity identification corresponding to the standardized entity that most closely matches the query term and the calculated confidence score; and

augmenting the search query with the entity identification corresponding to the standardized entity that most closely matches the query term along with an OR operator between the query term and the entity identification corresponding to the standardized entity.

14. The non-transitory machine-readable storage medium of claim 13 , wherein the instructions further cause the one or more machines to perform operations comprising, for each of the one or more query terms, augmenting the search query with a standardized entity and corresponding entity identification for a standardized entity indicated as a synonym for the query term.

15. The non-transitory machine-readable storage medium of claim 13 , wherein the instructions further cause the one or more machines to perform operations comprising:

eliminating any standardized entity from the search query that has a corresponding confidence score that does not transgress a preset threshold.

16. The non-transitory machine-readable storage medium of claim 13 , wherein the confidence score indicates a statistical likelihood that a user specifying the query term in a search query would have, under ideal circumstances, also entered the corresponding standardized entity in the search query.

17. The non-transitory machine-readable storage medium of claim 16 , wherein the confidence score is calculated by using a confidence score model trained via a machine learning algorithm based on member profiles and member activities in a social networking service.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE RE-FILE THE EXECUTEDASSIGNMENT WITH SIGNATURES AND DATES PREVIOUSLY RECORDED ON REEL 038688 FRAME 0873. ASSIGNOR(S) HEREBY CONFIRMS THE THE ASSIGNMENT. Recorded Dec 4, 2018
From: ARYA, DHRUV; LAD, ABHIMANYU; SINHA, SHAKTI DHIRENDRAJI; KANDURI, SATYA PRADEEP
To: LINKEDIN CORPORATION
Reel/Frame 047716/0498 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2017
From: LINKEDIN CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 044746/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2016
From: ARYA, DHRUV; LAD, ABHIMANYU; SINHA, SHAKTI DHIRENDRAJI; KANDURI, SATYA PRADEEP
To: LINKEDIN CORPORATION
Reel/Frame 038688/0873 →
Continuity (1)
Related Publication 20170337202A1 · Nov 23, 2017