IP Library Granted Patent US 9,990,421
Granted Patent B2
US 9,990,421 · App. 15/423,138 · Granted Jun 5, 2018

Phrase-based searching in an information retrieval system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,990,421
App. No.
15/423,138
Granted
Jun 5, 2018
Kind
B2
Abstract

An information retrieval system uses phrases to index, retrieve, organize and describe documents. Phrases are identified that predict the presence of other phrases in documents. Documents are the indexed according to their included phrases. Related phrases and phrase extensions are also identified. Phrases in a query are identified and used to retrieve and rank documents. Phrases are also used to cluster documents in the search results, create document descriptions, and eliminate duplicate documents from the search results, and from the index.

Claims (45)

1. A computer-implemented method comprising:

obtaining, from a phrase-based index for an Internet search engine, a list of documents from a collection of documents available via the Internet that contain a first phrase, the first phrase being relevant to a query;

for each document in the list:

determining, using related phrase information stored in the index for each document in the list of documents, whether the document includes one or more related phrases of the first phrase, where each related phrase has an actual co-occurrence rate of the related phrase and the first phrase in the document collection that exceeds an expected co-occurrence rate of the related phrase and the first phrase in the document collection;

ranking the documents in the list based on a quantity of related phrases determined for each document, so that documents with more related phrases are ranked higher than documents with fewer related phrases; and

selecting at least some of the highest-ranked documents to include in a result to the query.

2. The method of claim 1 , wherein determining whether the document includes one or more related phrases of the first phrase includes:

accessing a posting list for the first phrase, the posting list including, for each document identified in the posting list, an indication of the quantity of related phrases present in the document.

3. The method of claim 1 , wherein a document with a low frequency of query terms but a plurality of related phrases for the first phrase ranks higher than a document with a higher frequency of query terms but with no related phrases.

4. The method of claim 1 , further comprising:

storing the related phrases for a first phrase with respect to a document in a bit vector, wherein a bit of the bit vector is set for each related phrase of the first phrase that is present in the document, and a bit of the vector is unset for each related phrase of the first phrase that is not present in the document, wherein the bit vector has a numerical value.

5. The method of claim 1 , wherein the list is a first list and the method further comprises:

obtaining a second list of documents containing related phrases of the first phrase but lacking the first phrase; and

ranking the second list of documents in descending order of a ratio of the actual co-occurrence rate and the expected co-occurrence rate; and

selecting at least one of the highest ranking documents in the second list to include in a result to the query.

6. The method of claim 5 , further comprising:

removing documents relating to multiple topics from the second list prior to ranking.

7. The method of claim 1 , the list being a first list and the method further comprising:

obtaining a second list of documents identified by an inlink from a referring document, the inlink having anchor text that contains the first phrase; and

for each document in the second list, determining a ranking score that ranking the including the second list of documents in the ranking.

8. The method of claim 1 , where the expected co-occurrence rate of the related phrase and the first phrase is a function of a number of documents in the document collection that include the first phrase and a number of documents in the document collection that include the related phrase, and the actual co-occurrence rate being a function of a number of times the first phrase appears within a threshold number of words of the related phrase in the document collection.

9. The method of claim 8 , wherein the threshold is about 100.

10. A system for selecting documents from a document collection in response to a query, the system comprising:

one or more memory devices configured store executable instructions; and

one or more processors configured to execute the stored instructions to cause the system to:

obtain, from a phrase-based index for an Internet search engine, a list of documents from a collection of documents available via the Internet that contain a first phrase, the first phrase being relevant to a query,

for each document in the list: determine, using related phrase information stored in the index for each document in the list of documents, whether the document includes one or more related phrases of the first phrase, where each related phrase has an actual co-occurrence rate of the related phrase and the first phrase in the document collection that exceeds an expected co-occurrence rate of the related phrase and the first phrase in the document collection,

rank the documents in the list based on a quantity of related phrases determined for each document, so that documents with more related phrases are ranked higher than documents with fewer related phrases, and

select at least some of the highest-ranked documents to include in a result to the query.

11. The system of claim 10 , wherein determining whether the document includes one or more related phrases of the first phrase includes:

accessing a posting list for the first phrase, the posting list including, for each document identified in the posting list, an indication of the quantity of related phrases present in the document.

12. The system of claim 10 , wherein a document with a low frequency of query terms but a plurality of related phrases for the first phrase ranks higher than a document with a higher frequency of query terms but with no related phrases.

13. The system of claim 10 , wherein the one or more processors configured to execute the stored instructions further cause the system to:

store the related phrases for a first phrase with respect to a document in a bit vector, wherein a bit of the bit vector is set for each related phrase of the first phrase that is present in the document, and a bit of the vector is unset for each related phrase of the first phrase that is not present in the document, wherein the bit vector has a numerical value.

14. The system of claim 10 , wherein the list is a first list and the one or more processors configured to execute the stored instructions further cause the system to:

obtain a second list of documents containing related phrases of the first phrase but lacking the first phrase; and

rank the second list of documents in descending order of a ratio of the actual co-occurrence rate and the expected co-occurrence rate; and

select at least one of the highest ranking documents in the second list to include in a result to the query.

15. The system of claim 14 , wherein the one or more processors configured to execute the stored instructions further cause the system to:

remove documents relating to multiple topics from the second list prior to ranking.

16. The system of claim 10 , the list being a first list and the one or more processors configured to execute the stored instructions further cause the system to:

obtain a second list of documents identified by an inlink from a referring document, the inlink having anchor text that contains the first phrase; and

for each document in the second list, determine a ranking score that ranking the including the second list of documents in the ranking.

17. The system of claim 10 , where the expected co-occurrence rate of the related phrase and the first phrase is a function of a number of documents in the document collection that include the first phrase and a number of documents in the document collection that include the related phrase, and the actual co-occurrence rate being a function of a number of times the first phrase appears within a threshold number of words of the related phrase in the document collection.

18. The system of claim 17 , wherein the threshold is about 100.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2018
From: PATTERSON, ANNA L.
To: GOOGLE INC.
Reel/Frame 045605/0727 →
CHANGE OF NAME Recorded Dec 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044695/0115 →