IP Library Granted Patent US 7,519,588
Granted Patent B2
US 7,519,588 · App. 11/452,709 · Granted Apr 14, 2009

Keyword characterization and application

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,519,588
App. No.
11/452,709
Granted
Apr 14, 2009
Kind
B2
Abstract

Methods, apparatuses, and articles for receiving a collection of documents and/or objects determined to be potentially relevant to a keyword, and processing the collection of documents and/or objects to extract one or more keyword characterizations for use as proxies for the keyword, are described herein. In various embodiments, the one or more keyword characterizations may be used to compute a measure of keyword similarity for the keyword, facilitate keyword behavior modeling of the keyword, and/or find one or more advertisements.

Claims (61)

1. A method comprising:

receiving, by a computing device, a first collection of documents and/or objects determined by a first process to be relevant to a keyword, wherein the first process comprises searching a multiplicity of documents and/or objects;

processing, by the computing device, the first collection of documents and/or objects to extract one or more keyword characterizations from within at least one of the documents and/or objects of the first collection, wherein the processing comprises generating, by the computing device, a spectrum of n-grams to characterize the one or more keywords, where n is an integer equal to or greater than 1; and

receiving, by the computing device, a second collection of documents and/or objects determined by a second process to be relevant to the one or more keyword characterizations, wherein the second process comprises using at least one of the one or more keyword characterizations as proxies for the keyword.

2. The method of claim 1 , wherein a selected one of the first and the second collection of documents and/or objects is received as search results produced by a search engine from a search, based on the keyword, of a selected one of a database, a corpus of information, or a World Wide Web.

3. The method of claim 1 , wherein a selected one of the first and the second collection of documents and/or objects comprises at least one of: web pages determined to be potentially relevant to the keyword, documents from an electronic information corpus, and data objects including at least one of images, video files, audio files, executable applications, and abstractions of physical objects.

4. The method of claim 1 , wherein the processing further comprises

extracting, by the computing device, noun phrases, proper nouns, or named entities from the first collection of documents and/or objects and aggregating the noun phrases, proper nouns, or named entities;

determining, by the computing device, links to and/or from a web page of the first collection of documents and/or objects;

calculating, by the computing device, a distance to a set of websites or data resources, wherein the distance is a number of link traversals required to get between a search results page of the first collection of documents and/or objects and one of the set of websites or data resources;

determining, by the computing device, a distance metric from a word of the keyword to representations of a range of core word senses; and

determining, by the computing device, a web page of the first collection of documents and/or objects.

5. The method of claim 1 , wherein the generating the spectrum of n-grams comprises determining, by the computing device, a frequency of occurrence of each of the plurality of n-grams and normalizing the frequency of occurrence of each of the plurality of n-grams relative to a reference corpus.

6. The method of claim 1 , further comprising computing, by the computing device, a measure of keyword similarity for the keyword, based at least on the one or more keyword characterizations for use as proxies for the keyword.

7. The method of claim 6 , wherein the processing comprises generating, by the computing device, a spectrum of n-grams, and the measure of keyword similarity is computed by taking a dot product of the spectrum of n-grams and weighing each n-gram based on an inverse frequency of that n-gram.

8. The method of claim 6 , wherein the measure of keyword similarity is computed using a Bayesian classifier, wherein the keyword or one of the one or more keyword characterizations is treated as a document, and another keyword or one of the one or more keyword characterizations is treated as a category.

9. The method of claim 1 , further comprising facilitating, by the computing device, keyword behavior modeling of the keyword, based at least on the one or more keyword characterizations for use as proxies for the keyword.

10. The method of claim 9 , wherein the one or more keyword characterizations are input into models of keyword click-through and revenue-generating properties of search advertisements.

11. The method of claim 9 , wherein the keyword behavior modeling includes at least one of a neural network and a backward propagation system.

12. The method of claim 1 , further comprising filtering, by the computing device, a plurality of keywords, based at least on the one or more keyword characterizations.

13. The method of claim 1 , further comprising finding, by the computing device, one or more advertisements, by a search engine, based at least on the one or more keyword characterizations for use as proxies for the keyword.

14. The method of claim 13 , further comprising finding, by the computing device, a topic most relevant to the one or more keyword characterizations, and finding the one or more advertisements based at least in part on the topic.

15. The method of claim 13 , wherein the one or more advertisements are relevant to a domain name.

16. The method of claim 1 , further comprising

processing, by the computing device, the second collection of documents and/or objects to extract an additional one or more keyword characterizations to be merged with the one or more keyword characterizations for use as proxies for the keyword.

17. An apparatus comprising:

a processor; and

a generator, operated by the processor and adapted to

receive a first collection of documents and/or objects determined by a first process to be relevant to a keyword, wherein the first process comprises searching a multiplicity of documents and/or objects,

process the collection of documents and/or objects to extract one or more keyword characterizations from within at least one of the documents and/or objects of the first collection, and

receive a second collection of documents and/or objects determined by a second process to be relevant to the one or more keyword characterizations, wherein the second process comprises using the one or more keyword characterizations as proxies for the keyword;

wherein said process the collection of documents and/or objects comprises generation of a spectrum of n-grams to characterize the one or more keywords, where n is an integer equal to or greater than 1.

18. The apparatus of claim 17 , wherein a selected one of the first and the second collection of documents and/or objects is received as search results produced by a search engine from a search, based on the keyword, of a selected one of a database, a corpus of information, or a World Wide Web.

19. The apparatus of claim 17 , wherein a selected one of the first and the second collection of documents and/or objects comprises at least one of: web pages determined to be potentially relevant to the keyword, documents from an electronic information corpus, and data objects including at least one of images, video files, audio files, executable applications, and abstractions of physical objects.

20. The apparatus of claim 17 , wherein the generator is adapted to process a selected one of the first and the second collection of documents and/or objects, and the processing further comprises:

extracting noun phrases, proper nouns, or named entities from the selected one collection of documents and/or objects and aggregating the noun phrases, proper nouns, or named entities;

determining links to and/or from a web page of the selected one collection of documents and/or objects;

calculating a distance to a set of websites or data resources, wherein the distance is a number of link traversals required to get between a search results page of the selected one collection of documents and/or objects and one of the set of websites or data resources;

determining a distance metric from a word of the keyword to representations of a range of core word senses; and

determining a web page of the selected one collection of documents and/or objects.

21. The apparatus of claim 17 , wherein the apparatus further comprises a computing engine adapted to compute a measure of keyword similarity for the keyword, based at least on the one or more keyword characterizations to be used as proxies for the keyword.

22. The apparatus of claim 17 , wherein the apparatus further comprises a modeler adapted to facilitate keyword behavior modeling of the keyword, based at least on the one or more keyword characterizations to be used as proxies for the keyword.

23. The apparatus of claim 17 , wherein the apparatus further comprises a filter adapted to filter a plurality of keywords, based at least on the one or more keyword characterizations.

24. The apparatus of claim 17 , wherein the apparatus further comprises a search engine adapted to find one or more advertisements, based at least on the one or more keyword characterizations to be used as proxies for the keyword.

25. The apparatus of claim 17 , wherein the generator is included in a keyword search engine of the apparatus.

26. An article of manufacture comprising:

a storage medium; and

a plurality of programming instructions designed to program an apparatus and enable the apparatus to

receive a collection of documents and/or objects determined by a first process to be relevant to a keyword, wherein the first process comprises searching a multiplicity of documents and/or objects; and

process the collection of documents and/or objects to extract one or more keyword characterizations from within at least one of the documents and/or objects of the first collection, the one or more keyword characterizations to be used as proxies for the keyword in a second process, wherein the second process comprises searching a multiplicity of documents and/or objects;

wherein process comprises generation of a spectrum of n-grams to characterize the one or more keywords, where n is an integer equal to or greater than 1.

27. The article of claim 26 , wherein the collection of documents and/or objects comprise at least one of: web pages determined to be potentially relevant to the keyword, documents from an electronic information corpus, and data objects including at least one of images, video files, audio files, executable applications, and abstractions of physical objects.

28. The article of claim 26 , wherein the programming instructions are further designed to enable the apparatus to process the collection of documents and/or objects, and the processing further comprises:

extracting noun phrases, proper nouns, or named entities from the collection of documents and/or objects and aggregating the noun phrases, proper nouns, or named entities;

determining links to and/or from a web page of the collection of documents and/or objects;

calculating a distance to a set of websites or data resources, wherein the distance is a number of link traversals required to get between a search results page of the collection of documents and/or objects and one of the set of websites or data resources;

determining a distance metric from a word of the keyword to representations of a range of core word senses; and

determining a web page of the collection of documents and/or objects.

29. The article of claim 26 , wherein the programming instructions are further designed to enable the apparatus to compute a measure of keyword similarity for the keyword, based at least on the one or more keyword characterizations to be used as proxies for the keyword.

30. The article of claim 26 , wherein the programming instructions are further designed to enable the apparatus to facilitate keyword behavior modeling of the keyword, based at least on the one or more keyword characterizations to be used as proxies for the keyword.

31. The article of claim 26 , wherein the programming instructions are further designed to enable the apparatus to find one or more advertisements, based at least on the one or more keyword characterizations to be used as proxies for the keyword.

Assignments (3)
CHANGE OF NAME Recorded Mar 6, 2019
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 048525/0042 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2012
From: EFFICIENT FRONTIER, INC.
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 027702/0156 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2006
From: MASON, ZACHARY
To: EFFICIENT FRONTIER
Reel/Frame 017974/0719 →