IP Library Granted Patent US 8,756,212
Granted Patent B2
US 8,756,212 · App. 12/497,887 · Granted Jun 17, 2014

Techniques for web site integration

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,756,212
App. No.
12/497,887
Granted
Jun 17, 2014
Kind
B2
Abstract

Disclosed is a method and device for finding documents, such as Web pages, for presentation to a user, automatically or in response to a user expression of interest, which documents are part of a Web site being accessed by the user, and which documents relate to a document, such as a Web page, being accessed in the Web site. The method takes advantage of information retrieval techniques. The method generates the search query to use to find documents by reference to the text of the document in the Web site being accessed by the user. The method further uses a weighting function to weigh the terms used in the search query.

Claims (41)

1. A method, comprising:

providing, by operation of a computer system, a first document from a web site to a user, wherein the first document includes a plurality of terms;

automatically, by operation of the computer system, generating a search query from terms included in the first document, wherein the search query includes multiple terms from the first document, and wherein generating the search query comprises:

determining a respective first ratio for each of the multiple terms in the search query from a number of occurrences of the term in the first document and a total number of term occurrences in the first document,

determining a respective second ratio for each of the multiple terms in the search query from a number of occurrences of the term in the web site and a total number of term occurrences in the web site,

computing a respective weight for each of the multiple terms from the first ratio for the term and the second ratio for the term, and

assigning the respective weight for each of the multiple terms in the search query to the term;

using the search query to determine a respective score for each of a plurality of documents in the web site, wherein the respective score for each document is based upon occurrences in the document of terms in the search query and on the respective weights assigned to the terms in the search query; and

identifying a set of documents from the plurality of documents in the web site based on the respective scores.

2. The method of claim 1 , wherein generating the search query, using the search query, and identifying the set of documents are performed automatically in response to a user accessing the first document.

3. The method of claim 1 , wherein generating the search query and using the search query are performed in response to a request from a user.

4. The method of claim 1 , wherein the first document is a web page in the web site and the plurality of documents are web pages in the web site.

5. The method of claim 1 , further comprising:

outputting results associated with the identified set of documents.

6. The method of claim 1 , wherein determining respective scores for each of the plurality of documents comprises assigning scores using compressed document surrogates associated with the plurality of documents, wherein a particular compressed document surrogate is associated with a particular document among the plurality of documents, and wherein the particular compressed document surrogate comprises data representing counts of the occurrences of at least a subset of terms in the particular document.

7. The method of claim 6 , wherein determining respective scores comprises:

determining a score for a first one of a plurality of compressed document surrogates based on how often at least one term occurs in the first document compared to how often the at least one term occurs in the first compressed document surrogate.

8. The method of claim 6 wherein identifying a set of documents from the plurality of documents in the web site based on the respective scores comprises identifying a set of documents from the plurality of documents in the web site associated with scores greater than a predetermined threshold.

9. The method of claim 1 wherein generating the search query includes generating the search query while presenting the first document to the user.

10. The method of claim 1 , wherein computing the respective weight for each of the multiple terms from the first ratio for the term and the second ratio for the term comprises:

computing the weight for the term by computing a logarithm of a ratio between the first ratio for the term and the second ratio for the term.

11. A system, comprising:

a computer system and non-transitory media containing a computer program, the computer program programming the computer system to perform operations comprising:

providing, by operation of a computer system, a first document from a web site to a user, wherein the first document includes a plurality of terms;

automatically, by operation of the computer system, generating a search query from terms included in the first document, wherein the search query includes multiple terms from the first document, and wherein generating the search query comprises:

determining a respective first ratio for each of the multiple terms in the search query from a number of occurrences of the term in the first document and a total number of term occurrences in the first document,

determining a respective second ratio for each of the multiple terms in the search query from a number of occurrences of the term in the web site and a total number of term occurrences in the web site,

computing a respective weight for each of the multiple terms from the first ratio for the term and the second ratio for the term, and

assigning the respective weight for each of the multiple terms in the search query to the term;

using the search query to determine a respective score for each of a plurality of documents in the web site, wherein the respective score for each document is based upon occurrences in the document of terms in the search query and on the respective weights assigned to the terms in the search query; and

identifying a set of documents from the plurality of documents in the web site based on the respective scores.

12. The system of claim 11 , wherein the first document comprises a web page in the web site and the plurality of documents comprises other web pages in the web site.

13. The system of claim 11 , wherein determining respective scores comprises determining a score for a first one of a plurality of compressed document surrogates based on how often at least one term occurs in the first document compared to how often the at least one term occurs in the first compressed document surrogate.

14. The system of claim 11 , wherein generating the search query, using the search query, and identifying the set of documents are performed automatically in response to the user accessing the first document.

15. The system of claim 11 , wherein generating the search query and using the search query are performed in response to a request from a user.

16. The system of claim 11 wherein the operations further comprise:

outputting results associated with the identified set of documents.

17. The system of claim 11 wherein determining respective scores for each of the plurality of documents comprises assigning scores using compressed document surrogates associated with the plurality of documents, wherein a particular compressed document surrogate is associated with a particular document among the plurality of documents, and wherein the particular compressed document surrogate comprises data representing counts of the occurrences of at least a subset of terms in the particular document.

18. The system of claim 11 wherein generating the search query includes generating the search query while presenting the first document to the user.

19. The system of claim 11 , wherein computing the respective weight for each of the multiple terms from the first ratio for the term and the second ratio for the term comprises:

computing the weight for the term by computing a logarithm of a ratio between the first ratio for the term and the second ratio for the term.

Assignments (5)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2011
From: CHIPALKATTI, RENU; GETCHIUS, JEFFREY; PONTE, JAY
To: GTE LABORATORIES INCORPORATED
Reel/Frame 025637/0729 →
CHANGE OF NAME Recorded Jan 13, 2011
From: GTE LABORATORIES INCORPORATED
To: VERIZON LABORATORIES INC.
Reel/Frame 025637/0769 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2010
From: VERIZON PATENT AND LICENSING INC.
To: GOOGLE INC.
Reel/Frame 025328/0910 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2009
From: VERIZON LABORATORIES INC.
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 023574/0995 →