IP Library Granted Patent US 8,856,145
Granted Patent B2
US 8,856,145 · App. 11/639,849 · Granted Oct 7, 2014

System and method for determining concepts in a content item using context

Inventors: Jignashu Parikh (Gujarat, IN); John Thrall (San Francisco, CA)
Assignee: Yahoo! Inc.
G06F17/30613G06F17/30864
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,856,145
App. No.
11/639,849
Granted
Oct 7, 2014
Kind
B2
Abstract

The present invention is directed towards systems and methods for indexing one or more items of content. The method of the present invention comprises extracting one or more items of text from a given item of content. The one or more items of extracted text are tokenized into one or more concepts. One or more related concepts associated with the one or more concepts are identified. A support score is generated for the one or more concepts, and the item of content is index with the one or more concepts and the one or more associated support scores.

Claims (71)

1. A method implemented on at least one machine having at least one processor, storage, and a communication platform connected to a network for indexing one or more items of content, the method comprising:

extracting, by the at least one processor, one or more items of text from a given item of content;

tokenizing, by the at least one processor, the one or more extracted items of text into one or more concepts based on past queries submitted by one or more users;

identifying one or more related concepts associated with the one or more concepts;

obtaining, by the at least one processor, a support score for the individual one or more concepts based on whether one or more of the one or more concepts appear in the given item of content and/or whether one or more of the one or more related concepts appear in the given item of content; and

generating an index, the index comprising the given item of content associated with the one or more concepts and corresponding support scores for the individual one or more concepts;

receiving a search query;

identifying, based on the index, a set of items of content responsive to the search query, wherein individual items of content in the set are indexed with one or more concepts that are related to the search query;

obtaining, for each individual item of content in the set, a sum of support scores associated with the one or more concepts that are related to the search query; and

providing the set, wherein the items of content in the set are sorted based on the sum of support scores.

2. The method of claim 1 , wherein an item of content comprises a web page.

3. The method of claim 1 , wherein an item of content comprises a document.

4. The method of claim 1 , wherein an item of content comprises a video file.

5. The method of claim 1 , wherein an item of content comprises an audio file.

6. The method of claim 1 , wherein extracting one or more items of text comprises extracting one or more terms from the given item of content.

7. The method of claim 1 , wherein extracting one or more items of text comprises extracting one or more items of metadata describing the given item of content.

8. The method of claim 1 , wherein tokenizing the one or more items of text into one or more concepts comprises tokenizing the one or more items of text into one or more keywords or phrases.

9. The method of claim 8 , wherein tokenizing the one or more items of text into one or more keywords or phrases comprises:

identifying a frequency with which one or more keywords or phrases appear in a corpus of content items; and

tokenizing the one or more items of extracted text into the one or more keywords or phrases that appear with the greatest frequency in the corpus of content items.

10. The method of claim 9 , wherein identifying the frequency with which one or more keywords or phrases appear in the corpus of content items comprises identifying a frequency with which one or more keywords or phrases appear in the corpus of content items through use of ngram detection.

11. The method of claim 9 , wherein identifying the frequency with which one or more keywords or phrases appear in the corpus of content items comprises identifying a frequency with which one or more keywords or phrases appear in the corpus of content items as logically related units.

12. The method of claim 9 , wherein identifying the frequency with which one or more keywords or phrases appear in the corpus of content items comprises identifying a frequency with which one or more keywords or phrases appear in one or more query logs.

13. The method of claim 1 , wherein identifying one or more related concepts associated with the one or more concepts comprises identifying one or more keywords or phrases associated with the one or more concepts.

14. The method of claim 13 , wherein identifying one or more related concepts comprises identifying one or more query refinement keywords or phrases associated with a given concept.

15. The method of claim 13 , wherein identifying one or more related concepts comprises identifying one or more keywords or phrases submitted by a user during a query session.

16. The method of claim 13 , wherein identifying one or more related concepts comprises identifying one or more frequently co-occurring keywords or phrases in a corpus of content items.

17. The method of claim 13 , wherein identifying one or more related concepts comprises identifying one or more keywords or phrases associated with a given concept as specified by a human editor.

18. The method of claim 1 , wherein obtaining a support score for the individual one or more concepts comprises:

identifying a frequency with which the one or more concepts appear in the item of content.

19. The method of claim 1 , the method further comprising:

identifying one or more dominant concepts from the one or more concepts, wherein the one or more dominant concepts have corresponding support scores exceeding a support score threshold; and

wherein the index comprises the given item of content associated with the one or more dominant concepts and the corresponding support scores.

20. The method of claim 1 , wherein providing the set comprises transmitting the set to a client device.

21. A system comprising a processor coupled to a memory for indexing one or more items of content, the system comprising:

a text extractor operative to extract one or more items of text from an item of content;

a concept dictionary operative to maintain concepts;

a context dictionary operative to maintain related concepts associated with the concepts maintained in the concept dictionary; and

an aboutness extractor operative to:

tokenize the one or more extracted items of text into one or more concepts maintained in the concept dictionary based on past queries submitted by one or more users;

identify one or more related concepts associated with the one or more concepts in the item of content based on the context dictionary;

obtain a support score for the individual one or more concepts based on whether one or more of the one or more concepts appear in the item of content and/or whether one or more of the one or more related concepts appear in the item of content;

generate an index, the index comprising the item of content associated with the one or more concepts and corresponding support scores for the individual one or more concepts;

receive a search query;

identify, based on the index, a set of items of content responsive to the search query, wherein individual items of content in the set are indexed with one or more concepts that are related to the search query;

obtain, for each individual item of content in the set, a sum of support scores associated with the one or more concepts that are related to the search query; and

provide the set, wherein the items of content in the set are sorted based on the sum of support scores.

22. The system of claim 21 , wherein an item of content comprises a web page.

23. The system of claim 21 , wherein an item of content comprises a document.

24. The system of claim 21 , wherein an item of content comprises a video file.

25. The system of claim 21 , wherein an item of content comprises an audio file.

26. The system of claim 21 , wherein the text extractor is operative to extract one or more terms included in the item of content.

27. The system of claim 21 , wherein the text extractor is operative to extract one or more items of metadata associated with the item of content.

28. The system of claim 21 , wherein the concept dictionary is operative to maintain the concepts comprising keywords or phrases.

29. The system of claim 28 , wherein the concept dictionary is operative to maintain one or more keywords or phrases frequently appearing in a corpus of content items.

30. The system of claim 28 , wherein the concept dictionary is operative to maintain one or more keywords or phrases frequently appearing in one or more query logs.

31. The system of claim 21 , wherein the context dictionary is operative to maintain the related concepts comprising keywords or phrases associated with the concepts maintained in the concept dictionary.

32. The system of claim 31 , wherein the context dictionary is operative to maintain one or more query refinement keywords or phrases associated with a given concept maintained in the concept dictionary.

33. The system of claim 31 , wherein the context dictionary is operative to maintain one or more keywords or phrases submitted by a user during a query session.

34. The system of claim 31 , wherein the context dictionary is operative to maintain one or more frequently co-occurring keywords or phrases appearing in a corpus of content items.

35. The system of claim 31 , wherein the context dictionary is operative to maintain one or more keywords or phrases associated with a given concept as specified by a human editor.

36. The system of claim 21 , wherein the aboutness extractor is operative to:

identify one or more dominant concept from the one or more concepts, wherein the one or more dominant concepts have corresponding support scores exceeding a support score threshold; and

wherein the index comprises the item of content associated with the one or more dominant concepts and the corresponding support scores.

37. The system of clam 21 , comprising a data store.

38. The system of claim 37 , wherein the data store is operative to maintain information indicating the reliability of one or more items of content.

39. The system of claim 37 , wherein the data store is operative to maintain information indicating the reliability of one or more concepts maintained in the concept dictionary.

40. The system of claim 37 , wherein the aboutness extractor is operative to obtain the support score using the information maintained in the data store.

41. The system of claim 21 , comprising a dictionary manager.

42. The system of claim 41 , wherein the dictionary manager is operative to provide updated information associated with at least one of the concepts to the concept dictionary.

43. The system of claim 41 , wherein the dictionary manager is operative to provide updated information associated with at least one of the related concepts to the context dictionary.

Assignments (9)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNOR NAME PREVIOUSLY RECORDED AT REEL: 052853 FRAME: 0153. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 29, 2021
From: R2 SOLUTIONS LLC
To: STARBOARD VALUE INTERMEDIATE FUND LP, AS COLLATERAL AGENT
Reel/Frame 056832/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED ON REEL 053654 FRAME 0254. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST GRANTED PURSUANT TO THE PATENT SECURITY AGREEMENT PREVIOUSLY RECORDED. Recorded Dec 30, 2020
From: STARBOARD VALUE INTERMEDIATE FUND LP
To: R2 SOLUTIONS LLC
Reel/Frame 054981/0377 →
RELEASE OF SECURITY INTEREST IN PATENTS Recorded Jul 8, 2020
From: STARBOARD VALUE INTERMEDIATE FUND LP
To: ACACIA RESEARCH GROUP LLC; AMERICAN VEHICULAR SCIENCES LLC; BONUTTI SKELETAL INNOVATIONS LLC; CELLULAR COMMUNICATIONS EQUIPMENT LLC; INNOVATIVE DISPLAY TECHNOLOGIES LLC; LIFEPORT SCIENCES LLC; LIMESTONE MEMORY SYSTEMS LLC; MOBILE ENHANCEMENT SOLUTIONS LLC; MONARCH NETWORKING SOLUTIONS LLC; NEXUS DISPLAY TECHNOLOGIES LLC; PARTHENON UNIFIED MEMORY ARCHITECTURE LLC; R2 SOLUTIONS LLC; SAINT LAWRENCE COMMUNICATIONS LLC; STINGRAY IP SOLUTIONS LLC; SUPER INTERCONNECT TECHNOLOGIES LLC; TELECONFERENCE SYSTEMS LLC; UNIFICATION TECHNOLOGIES LLC
Reel/Frame 053654/0254 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2020
From: EXCALIBUR IP, LLC
To: R2 SOLUTIONS LLC
Reel/Frame 053459/0059 →
PATENT SECURITY AGREEMENT Recorded Jun 5, 2020
From: ACACIA RESEARCH GROUP LLC; AMERICAN VEHICULAR SCIENCES LLC; BONUTTI SKELETAL INNOVATIONS LLC; CELLULAR COMMUNICATIONS EQUIPMENT LLC; INNOVATIVE DISPLAY TECHNOLOGIES LLC; LIFEPORT SCIENCES LLC; LIMESTONE MEMORY SYSTEMS LLC; MERTON ACQUISITION HOLDCO LLC; MOBILE ENHANCEMENT SOLUTIONS LLC; MONARCH NETWORKING SOLUTIONS LLC; NEXUS DISPLAY TECHNOLOGIES LLC; PARTHENON UNIFIED MEMORY ARCHITECTURE LLC; R2 SOLUTIONS LLC; SAINT LAWRENCE COMMUNICATIONS LLC; STINGRAY IP SOLUTIONS LLC; SUPER INTERCONNECT TECHNOLOGIES LLC; TELECONFERENCE SYSTEMS LLC; UNIFICATION TECHNOLOGIES LLC
To: STARBOARD VALUE INTERMEDIATE FUND LP, AS COLLATERAL AGENT
Reel/Frame 052853/0153 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2016
From: YAHOO! INC.
To: EXCALIBUR IP, LLC
Reel/Frame 038950/0592 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2016
From: EXCALIBUR IP, LLC
To: YAHOO! INC.
Reel/Frame 038951/0295 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2016
From: YAHOO! INC.
To: EXCALIBUR IP, LLC
Reel/Frame 038383/0466 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2006
From: PARIKH, JIGNASHU; THRALL, JOHN
To: YAHOO! INC.
Reel/Frame 018692/0283 →
Priority Claims (1)
IN 1396/CHE/2006 · Aug 4, 2006 · national
Continuity (1)
Related Publication 20080033982A1 · Feb 7, 2008