IP Library › Granted Patent US 8,868,567
Granted Patent B2
US 8,868,567 · App. 13/019,696 · Granted Oct 21, 2014

Information retrieval using subject-aware document ranker

Inventors: Girish Kumar (Kirkland, WA); Alfian Tan (Issaqua, WA); Nicholas Eric Craswell (Seattle, WA)
Assignee: Microsoft Corporation
G06F17/2785G06F17/3053G06F17/279
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,868,567
App. No.
13/019,696
Granted
Oct 21, 2014
Kind
B2
Abstract

Subject matter described herein is related to determining a document score, which suggests a relevance of a document (e.g., webpage) to a search query. For example, a search query is received that is comprised of one or more terms, which represent a subject. An equivalent subject is identified that is semantically similar to the subject. The document score is determined by accounting for both a subject frequency and an equivalent-subject frequency.

Claims (41)

1. A computer-storage medium storing computer-executable instructions that, when executed, perform a method of determining a document score, which suggests a relevance of a document to a search query, the method comprising:

receiving the search query;

parsing the search query into a first n-gram having a first weight and a second n-gram having a second weight, wherein the first weight quantifies an importance of the first n-gram to the search query and the second weight quantifies an importance of the second n-gram to the search query;

determining that the first weight and the second weight satisfy a threshold weight criterion, wherein an n-gram is not used to determine the document score when a weight of the n-gram does not satisfy the threshold weight criterion;

identifying a first equivalent subject that is semantically similar to the first n-gram and a second equivalent subject that is semantically similar to the second n-gram, wherein the first n-gram and the first equivalent subject comprise a first subject group and the second g-gram and the second equivalent subject comprise a second subject group; and

determining the document score of the document,

wherein the document score is comprised of a first subject-group score and a second subject-group score, and

wherein both subject-group scores are calculated using both a subject frequency, which includes a number of times a respective n-gram is found in the document, and an equivalent-subject frequency, which includes a number of times a respective equivalent-subject is found in the document.

2. The computer-storage medium of claim 1 , wherein the first weight is a function of both a frequency of each term of the first n-gram in a corpus and a frequency of the first n-gram in the corpus.

3. The computer-storage medium of claim 1 , wherein the second weight is reduced when the second n-gram is deemed part of a larger n-gram.

4. The computer-storage medium of claim 1 , wherein identifying the equivalent subject includes referencing an equivalent-subject datastore, which maintains a listing of equivalent subjects that have been identified as semantically similar to the subject.

5. The computer-storage medium of claim 1 , wherein, when calculating the subject-group score, the equivalent-subject frequency is weighted based on a confidence that the equivalent subject is semantically similar to the subject.

6. The computer-storage medium of claim 1 ,

wherein the subject-group score is calculated by applying a saturation function to a subject-group frequency, and

wherein the subject-group frequency includes both the subject frequency and the equivalent-subject frequency.

7. The computer-storage medium of claim 6 , wherein the subject frequency and the equivalent subject frequency are weighted.

8. The computer-storage medium of claim 6 , wherein the saturation function includes a sum of a customizable parameter and the subject-group frequency.

9. The computer-storage medium of claim 1 , wherein the document score is calculated using a plurality of subject-group scores that correspond to a plurality of subjects included in the search query.

10. The computer-storage medium of claim 1 , wherein the document score of the document is comprised of a sum of the plurality of subject-group scores.

11. A method of determining a document score, which suggests a relevance of a document to a search query, the method comprising:

receiving the search query;

parsing the search query into a plurality of n-grams including a first n-gram, which includes a first weight quantifying an importance of the first n-gram to the search query;

determining that the first weight satisfies a threshold weight criterion;

identifying a first equivalent subject that is semantically similar to the first n-gram and that forms a subject group with the first n-gram, wherein the first equivalent subject is associated with an equivalent-subject score quantifying a confidence that the first equivalent subject and the first n-gram identify a same subject;

determining a first-subject-group frequency comprised of both a first-subject frequency, which includes a number of times the first n-gram is found in the document, and a first-equivalent-subject frequency, which includes a number of times the first equivalent subject is found in the document; and

calculating the document score of the document,

wherein the document score is comprised of a first-subject-group score, and

wherein the first-subject-group score is calculated by combining the first-equivalent-subject frequency with the equivalent-subject score, such that the first-equivalent-subject frequency is weighted based on a confidence that the first equivalent subject is semantically similar to the first n-gram.

12. The method of claim 11 , wherein the first-subject frequency and the first-equivalent-subject frequency are weighted based on respective locations within the document at which the first n-gram subject and first-equivalent subject are found.

13. The method of claim 11 , wherein identifying an equivalent subject includes referencing an equivalent-subject datastore, which maintains a listing of equivalent subjects that have been identified as semantically similar to the first subject and the second subject.

14. The method of claim 11 , wherein the method further comprises determining a second-subject-grouping score of another n-gram parsed from the search query, and wherein the document score is calculated by adding the first-subject-grouping score and the second-subject-grouping score.

15. A computer system that determines a document score, which suggests a relevance of a document to a search query, the computer system including a processor coupled to computer-storage media, which includes computer-readable instructions executed by the processor to provide:

a search-query receiver that receives the search query

a subject identifier that leverages the processor to parse the search query into a plurality of n-grams including a first n-gram, wherein the first n-gram includes a weight that quantifies a relevance of the first n-gram to the search query and that is deemed to satisfy a threshold weight criterion;

an equivalent-subject identifier that references an equivalent-subject datastore to identify an equivalent subject, which is semantically similar to the n-gram, wherein the n-gram and the equivalent subject comprise a subject group and wherein the equivalent subject includes an equivalent-subject score quantifying a confidence that the equivalent subject and the first n-gram identify a same subject; and

a document ranker that leverages the processor to calculate the document score of the document by;

(1) determining a subject frequency, which includes a number of times the subject is found in the document, and an equivalent-subject frequency, which includes a number of times the equivalent subject is found in the document, and

(2) calculating a weighted equivalent-subject frequency by multiplying the equivalent-subject frequency by the equivalent-subject score.

16. The computer system of claim 15 ,

wherein calculating the document score further comprises calculating a subject-group score; and

wherein the document ranker calculates the subject-group score by calculating a sum of the weighted equivalent-subject frequency and a weighted subject frequency, which is weighted based on a location of the n-gram in the document.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034544/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2011
From: KUMAR, GIRISH; TAN, ALFIAN; CRASWELL, NICHOLAS ERIC
To: MICROSOFT CORPORATION
Reel/Frame 025734/0683 →
Continuity (1)
Related Publication 20120197905A1 · Aug 2, 2012