IP Library › Granted Patent US 7,584,175
Granted Patent B2
US 7,584,175 · App. 10/900,075 · Granted Sep 1, 2009

Phrase-based generation of document descriptions

Assignee: Google Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,584,175
App. No.
10/900,075
Granted
Sep 1, 2009
Kind
B2
Abstract

An information retrieval system uses phrases to index, retrieve, organize and describe documents. Phrases are identified that predict the presence of other phrases in documents. Documents are the indexed according to their included phrases. Related phrases and phrase extensions are also identified. Phrases in a query are identified and used to retrieve and rank documents. Phrases are also used to cluster documents in the search results, create document descriptions, and eliminate duplicate documents from the search results, and from the index.

Claims (33)

1. A method of automatically generating a description of a document, the method comprising:

retrieving a document in response to a query, the query comprising a query phrase, the document including a plurality of sentences;

calculating, by operation of a processor adapted to manipulate data within a computer system, for sentences of the document, a first count that includes a measure of a number of instances in which the query phrase occurs in the sentences;

calculating, by operation of a processor adapted to manipulate data within a computer system, for sentences of the document, a second count that includes a measure of a number of instances in which any of one or more related phrases of the query phrase occurs in the sentences;

calculating, by operation of a processor adapted to manipulate data within a computer system, for sentences of the document, a third count that includes a measure of a number of instances in which any of one or more phrase extensions of the query phrase occurs in the sentences, wherein a phrase extension is a super-sequence that begins with the query phrase;

selecting one or more of the sentences of the document based on their respective first, second and third counts; and

forming a description of the document from the selected sentences, wherein a phrase g j is a related phrase of another phrase g k where an information gain of g j with respect to g k exceeds a predetermined threshold, the information gain being a function of both actual and expected co-occurrence rates of g j and g k .

2. The method of claim 1 , wherein selecting the one or more of the sentences of the document based on their respective first, second and third counts, comprises:

sorting the sentences of the document in declining order of their respective counts; and

selecting a number of the sentences of the document having the highest counts.

3. The method of claim 2 , wherein sorting the sentences of the document in declining order of their respective counts, comprises:

the first count of the query phrase constituting a primary sort key;

the second count constituting a secondary sort key; and

the third count constituting a tertiary sort key.

4. The method claim 1 , wherein forming a description of the document from the selected sentences, comprises concatenating the selected sentences to form a block of text.

5. The method claim 1 , wherein identifying related phrases of the query phrase in a sentence comprises reading a related phrase bit vector of the query phrase, the related phrase bit vector having an ordered set of bits, each bit indicating whether a corresponding related phrase is present in a given document.

6. A tangible computer readable storage medium storing a computer program executable by a processor for automatically generating a description of a document, the operations of the computer program comprising:

retrieving a document in response to a query, the query comprising a query phase, the document including a plurality of sentences;

calculating for sentences of the document, a first count that includes a measure of a number of instances in which the query phrase occurs in the sentences;

calculating, for sentences of the document, a second count that includes a measure of a number of instances in which any of one or more related phrases of the query phrase occurs in the sentences;

calculating, for sentences of the document, a third count that includes a measure of a number of instances in which any of one or more phrase extensions of the query phrase occurs in the sentences, wherein a phrase extension is a super-sequence that begins with the query phrase;

selecting one or more of the sentences of the document based on their respective first, second and third counts; and

forming a description of the document from the selected sentences, wherein a phrase g j is a related phrase of another phrase g k where an information gain of g j with respect to g k exceeds a predetermined threshold, the information gain being a function of both actual and expected co-occurrence rates of g j and g k .

7. A computer implemented system for automatically generating a description of a document, comprising:

a document retrieval system executed by a computer and adapted to retrieve a document in response to a query, the query including a query phase, the document including a plurality of sentences; and

a document description system executed by a computer and adapted to:

calculate, for sentences of the document, a first count that includes a measure of a number of instances in which the query phrase occurs in the sentences;

select one or more of the sentences of the document based on their respective counts, the selecting further comprising:

sorting sentences of the document in declining order of their respective counts, the first count constituting a primary sort key, a second count that includes a measure of a number of instances in which any of one or more related phrases of the query phrase occurs in the respective sentences of the document constituting a secondary sort key, and a third count that includes a measure of a number of instances in which any of one or more phrase extensions of the query phrase occurs in the respective sentences of the document constituting a tertiary sort key, and

selecting a number of the sentences of the document having the highest counts; and

form a description of the document from the selected sentences.

8. The system of claim 7 , wherein generating a document description comprising selected sentences of the document, comprises concatenating the selected sentences to form a block of text.

9. The system of claim 7 , wherein a phrase g j is a related phrase of another phrase g k where an information gain of g j with respect to g k exceeds a predetermined threshold, the information gain being a function of both actual and expected co-occurrence rates of g j and g k .

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044101/0610 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2004
From: PATTERSON, ANN ALYNN
To: GOOGLE, INC.
Reel/Frame 015429/0512 →
Continuity (1)
Related Publication 20060020571A1 · Jan 26, 2006