IP Library Granted Patent US 8,869,208
Granted Patent B2
US 8,869,208 · App. 13/467,339 · Granted Oct 21, 2014

Computing similarity between media programs

Inventors: Grzegorz Glowaty (Przemysl, PL); Michal Brzozowski (Warszawa, PL); Marcin Wielgus (Krakow, PL)
Assignee: Google Inc.
G06F17/30817
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,869,208
App. No.
13/467,339
Granted
Oct 21, 2014
Kind
B2
Abstract

System and method are provided to associate or compare media programs. A method includes: obtaining, using at least one processing circuit, first metadata for a first media program and second metadata for a second media program, wherein the first metadata are organized into a plurality of first fields, and the second metadata are organized into a plurality of second fields; extracting, using at least one processing circuit, a plurality of first tokens from one of the plurality of the first fields and a plurality of second tokens from one of the plurality of second fields; assigning a weight factor to each of the first and second tokens; cross-correlating the first and second tokens between the plurality of first fields and the plurality of second fields; and calculating a similarity score between the first and second media programs based on the cross-correlating.

Claims (27)

1. A computer-implemented method of associating media programs as part of a smart television platform, the method comprising:

obtaining, using at least one processing circuit, first metadata for a first media program comprising a plurality of first fields and second metadata for a second media program comprising a plurality of second fields;

extracting, using at least one processing circuit, a plurality of first tokens from each of the plurality of the first fields and a plurality of second tokens from each of the plurality of second fields;

calculating, for each of the first tokens, a first weight factor and calculating, for each of the second tokens, a second weight factor, wherein the first weight factor and the second weight factor are based on the difference between: the ratio of the logarithm of the number of token occurrences to the logarithm of the sum of one and the total number of possible occurrences, and one;

generating, for each of the first tokens, a first weighted token based on one of the first weight factors;

generating, for each of the second tokens, a second weighted token based on one of the second weight factors;

generating, for each of the first fields, a first vector comprising a plurality of the first weighted tokens;

generating, for each of the second fields, a second vector comprising a plurality of the second weighted tokens.

calculating a plurality of correlations indicative of similarities between the first fields and the second fields, wherein each of the plurality of correlations is a dot product of one of the first vectors and one of the second vectors that are selected based on a field type;

determining a similarity score between the first metadata and the second metadata by combining the plurality of correlations; and

associating the first and second media programs based on the determined similarity score.

2. The method of claim 1 , wherein the plurality of first tokens comprises text included in the plurality of the fields.

3. The method of claim 1 , wherein at least one of the first tokens comprising a combination of text keywords included in the field.

4. The method of claim 1 , further comprises calculating a term frequency for each of the plurality of first tokens equal to the sum of the logarithm of the number of the token's occurrences in a field and 1.

5. The method of claim 1 further comprising identifying a word cluster as a token.

6. The method of claim 1 wherein the plurality of first tokens are extracted from each of the plurality of first fields using a probabilistic topic model.

7. The method of claim 1 further comprising identifying a token having a frequency of occurrence exceeding a predetermined threshold and removing the identified token from the plurality of first tokens.

8. The method of claim 1 , further comprises multiplying a term frequency and an inverse document frequency to generate the first weight factor for each of the first tokens.

9. The method of claim 1 further comprising normalizing each of the first vectors to have a length of 1.

10. A computer-implemented method of associating media programs as part of a smart television platform, the method comprising:

(A) obtaining, using at least one processing circuit, first metadata for a first media program comprising a plurality of first fields and second metadata for a second media program comprising a plurality of second fields;

(B) extracting, using at least one processing circuit, a plurality of first tokens from each of the plurality of the first fields and a plurality of second tokens from each of the plurality of second fields;

(C) calculating a term frequency (TF) for each of the plurality of extracted tokens, and the TF representing a frequency of occurrence of the token within the field;

(D) calculating an inverse document frequency (IDF) for each token and the IDF representing a frequency of occurrence of the token across a plurality of the fields, wherein the IDF is calculated as the difference between: the ratio of the logarithm of the number of token occurrences to the logarithm of the sum of 1 and the total number of possible occurrences; and 1;

(E) combining the TF and the IDF for each token score to generate a weight factor for each token;

(F) determining a similarity score between a first vector comprising the plurality of weighted first tokens and a second vector comprising the plurality of weighted second tokens; and

(G) associating the first and the second media programs responsive to the similarity score.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044277/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2012
From: GLOWATY, GRZEGORZ; BRZOZOWSKI, MICHAL; WIELGUS, MARCIN
To: GOOGLE INC.
Reel/Frame 028181/0225 →
Continuity (2)
Provisional Application 61553221 · Oct 30, 2011
Related Publication 20130111526A1 · May 2, 2013