IP Library › Granted Patent US 9,720,979
Granted Patent B2
US 9,720,979 · App. 14/504,420 · Granted Aug 1, 2017

Method and system of identifying relevant content snippets that include additional information

Inventors: Krishna Kishore Dhara (Hyderabad, IN); Anil Jwalanna (Cupertino, CA)
G06F17/3053H04L67/1002
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,720,979
App. No.
14/504,420
Filed
Oct 2, 2014
Granted
Aug 1, 2017
Kind
B2
Art Unit
2455
USPC
709/203
Abstract

In one exemplary aspect, a method includes the step of obtaining a content of a content block. The content is represented as a content vector. A query is received. The query is represented as a query vector. A hierarchical sliding similarity and dissimilarity is determined for matching the content vector and the query vector, this step can include the steps of: determining a similarity measure and a dissimilarity measure for each content vector element with respect to the query vector; identifying a strong match over a sliding window of sub-terms of each content vector element; computing a sub-similarity score and a sub-dissimilarity score for each level of the convent vector element; determining a final similarity score as a combination of the strong match of some sub-vectors at different levels; and determining a final dissimilarity score as a combination of the strong match of some sub-vectors at different levels.

Claims (44)

1. A method of information retrieval comprising:

obtaining a content of a content block;

representing the content as a content vector;

receiving a query;

representing the query as a query vector;

determining a hierarchical sliding similarity and dissimilarity for matching the content vector and the query vector, and wherein determining a hierarchical sliding similarity for matching the content vector and the query vector:

determining a similarity measure and a dissimilarity measure for each content vector element with respect to the query vector, wherein each content vector element is provided a unique sub-similarity similarity measure and a unique sub-dissimilarity measure;

for each unique sub-similarity similarity measure and each unique sub-dissimilarity measure of each content vector element, identifying a strong match over a sliding window of sub-terms of each content vector element, wherein the strong match comprises a strong similarity match and a strong dissimilarity match;

computing a sub-similarity score and a sub-dissimilarity score for each level of the convent vector element;

determining a final similarity score as a combination of the strong match of some sub-vectors at different levels; and

determining a final dissimilarity score as a combination of the strong match of some sub-vectors at different levels.

2. The method of claim 1 , further comprising:

returning at least one content block based on a combination of the final similarity score and the final dissimilarity score.

3. The method of claim 1 , wherein a content block comprises a structured digital information or an unstructured digital information, and wherein the content block comprises a set of text, a digital image, a video file or an audio file.

4. The method of claim 1 , wherein the vector format of the content comprises one or more feature vectors representing sub-terms of the content block.

5. The method of claim 1 , wherein the query vector comprises a vector model of a set of search words, a highlighted section of text or a microblog post.

6. The method of claim 1 , wherein determining the hierarchal sliding similarity does not determine a similarity or near-similarity based on the entire content vector.

7. The method of claim 1 , wherein the sub-terms of each content vector element comprise a sub-vector of the content vector.

8. The method of claim 1 , wherein the dissimilarity measure is computed as a compliment of a partial similarity between the query vector and the content vector and using a weighted contextual information derived from the query.

9. A computerized system of information retrieval comprising:

a hardware processor configured to execute instructions;

a memory including instructions when executed on the processor, causes the processor to perform operations that:

obtain a content of a content block;

represent the content as a content vector; receive a query;

represent the query as a query vector;

determine a hierarchical sliding similarity and dissimilarity for matching the content vector and the query vector, and wherein determining a hierarchical sliding similarity for matching the content vector and the query vector:

determine a similarity measure for each content vector element with respect to the query vector, wherein each content vector element is provided a unique sub-similarity similarity measure;

for each unique sub-similarity similarity measure of each content vector element, identify a strong similarity match over a sliding window of sub-terms of each content vector element;

compute a sub-similarity score for each level of the convent vector element;

determine a final similarity score as a combination of the strong match of some sub-vectors at different levels;

determine a dissimilarity measure for each content vector element with respect to the query vector, wherein each content vector element is provided a unique sub-dissimilarity measure;

for each unique sub-dissimilarity measure of each content vector element, identify a strong dissimilarity match over a sliding window of sub-terms of each content vector element;

compute a sub-dissimilarity score for each level of the convent vector element; and

determine a final dissimilarity score as a combination of the strong match, of some sub-vectors at different levels.

10. The computerized system of claim 9 , wherein the memory further includes instructions when executed on the processor, causes the processor to perform operations that:

return at least one content block based on a combination of the final similarity score and the final dissimilarity score.

11. The computerized system of claim 10 , wherein the returned at least one content block has a property of strong similarity with the query vector and a strong dissimilarity with the query vector.

12. The computerized system of claim 9 , wherein a content block comprises a structured digital information or an unstructured digital information, and wherein the content block comprises a set of text, a digital image, a video file or an audio file.

13. The computerized system of claim 9 , wherein the vector format of the content comprises one or more feature vectors representing sub-terms of the content block.

14. The computerized system of claim 9 , wherein the query vector comprises a vector model of a set of search words, a highlighted section of text or a microblog post.

15. The computerized system of claim 9 , wherein determining the hierarchal sliding similarity does not determine a similarity or near-similarity based on the entire content vector.

16. The computerized system of claim 9 , wherein the sub-terms of each content vector element comprise a sub-vector of the content vector.

17. The computerized system of claim 9 , wherein the similarity measure relates to a first aspect of the content vector element and the dissimilarity measure relates to a second content vector element.

18. The computerized system of claim 17 , wherein the first content vector element comprises a digital image element and the second content vector element comprises a participant identifier element.

Continuity (4)
Continuation In Part 13915327 · Jun 11, 2013
Provisional Application 61663169 · Jun 22, 2012
Provisional Application 61773083 · Mar 5, 2013
Related Publication 20150120720A1 · Apr 30, 2015