IP Library Granted Patent US 10,970,327
Granted Patent B2
US 10,970,327 · App. 16/248,627 · Granted Apr 6, 2021

Selecting balanced clusters of descriptive vectors

Inventors: Aneesh Vartakavi (Emeryville, CA); Peter C. DiMaria (Berkeley, CA); Markus K. Cremer (Orinda, CA); Phillip Popp (Oakland, CA)
Assignee: GRACENOTE, INC.
G06F16/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,970,327
App. No.
16/248,627
Granted
Apr 6, 2021
Kind
B2
Abstract

A clustering machine can cluster descriptive vectors in a balanced manner. The clustering machine calculates distances between pairs of descriptive vectors and generates clusters of vectors arranged in a hierarchy. The clustering machine determines centroid vectors of the clusters, such that each cluster is represented by its corresponding centroid vector. The clustering machine calculates a sum of inter-cluster vector distances between pairs of centroid vectors, as well as a sum of intra-cluster vector distances between pairs of vectors in the clusters. The clustering machine calculates multiple scores of the hierarchy by varying a scalar and calculating a separate score for each scalar. The calculation of each score is based on the two sums previously calculated for the hierarchy. The clustering machine may select or otherwise identify a balanced subset of the hierarchy by finding an extremum in the calculated scores.

Claims (92)

1. A method comprising:

accessing, by one or more processors, descriptive vectors that describe items, each descriptive vector comprising one or more values indicative of an extent to which one or more characteristics are present in a respective item of the items;

determining, by the one or more processors, one or more vector distances between one or more pairs of the descriptive vectors;

generating, by the one or more processors, a hierarchy of vector clusters by clustering the descriptive vectors into the vector clusters based on the determined one or more vector distances;

determining, by the one or more processors, centroid vectors of the vector clusters in the hierarchy, each centroid vector corresponding to a respective vector cluster in the hierarchy;

summing, by the one or more processors, one or more inter-cluster vector distances between one or more pairs of the centroid vectors;

summing, by the one or more processors, for each of the vector clusters, one or more intra-cluster vector distances between one or more pairs of descriptive vectors;

determining, by the one or more processors, a plurality of scores of the hierarchy by applying a plurality of weightings to the summed inter-cluster vector distances and the summed intra-cluster vector distances, wherein each score of the plurality of scores corresponds to a respective weighting of the plurality of weightings, and wherein a particular weighting of the plurality of weightings corresponds to an extreme score of the plurality of scores; and

selecting, by the one or more processors, a subset of the vector clusters in the hierarchy based on the weighting that corresponds to the extreme score.

2. The method of claim 1 , further comprising:

accessing the items prior to accessing the descriptive vectors, each of the items including respective media content; and

determining the descriptive vectors by generating a respective descriptive vector for each of the items, wherein generating each respective descriptive vector comprises analyzing the respective media content in a corresponding item to be described.

3. The method of claim 2 , wherein:

the accessed items are media items;

the method further comprises normalizing the media items by at least one of:

omitting duplicate media items, omitting non-original media items, omitting media items released on compilation albums, omitting media items recorded at live performances, or retaining media items recorded in studios; and

determining the descriptive vectors comprises generating a respective descriptive vector for each of the normalized media items.

4. The method of claim 1 , wherein:

determining the one or more vector distances between the one or more pairs of the descriptive vectors comprises determining the one or more vector distances between the one or more pairs of the descriptive vectors based on correlations among the descriptive vectors.

5. The method of claim 1 , wherein:

determining the one or more vector distances between the one or more pairs of the descriptive vectors comprises calculating one or more quadratic-chi histogram distances between the one or more pairs of the descriptive vectors.

6. The method of claim 1 , wherein:

clustering the descriptive vectors comprises clustering the descriptive vectors according to an agglomerative hierarchical clustering algorithm.

7. The method of claim 6 , wherein:

the agglomerative hierarchical clustering algorithm includes a complete-linkage clustering algorithm.

8. The method of claim 1 , wherein:

determining each score of the plurality of scores comprises:

selecting a scalar between zero and unity;

multiplying the scalar by the summed intra-cluster vector distances to obtain a first multiplicative product;

multiplying the summed inter-cluster vector distances by the scalar subtracted from unity to obtain a second multiplicative product; and

adding the first multiplicative product to the second multiplicative product to obtain the score.

9. The method of claim 8 , wherein:

the items are media items released in a set of albums by a same artist; and

selecting the scalar comprises selecting the scalar based on a count of albums in the set of albums by the same artist.

10. The method of claim 1 , further comprising:

modifying the selected subset of the vector clusters in the hierarchy, wherein modifying the selected subset comprises:

calculating weights of vector clusters in the selected subset, a first calculated weight corresponding to a first vector cluster in the selected subset; and

removing the first vector cluster from the selected subset based on the first calculated weight failing to exceed a threshold percentile of the calculated weights of the vector clusters in the selected subset.

11. The method of claim 10 , wherein:

calculating the weights of the vector clusters in the selected subset comprises calculating the weights of the vector clusters in the selected subset based on sizes of the vector clusters in the selected subset, the first calculated weight being calculated based on a count of descriptive vectors in the first vector cluster within the selected subset.

12. The method of claim 10 , wherein:

calculating the weights of the vector clusters in the selected subset comprises calculating the weights of the vector clusters in the selected subset based on average popularity scores of the vector clusters in the selected subset, the first calculated weight being calculated based on an average of a group of popularity scores that correspond to a group of items described by at least some descriptive vectors in the first vector cluster within the selected subset.

13. The method of claim 10 , wherein:

calculating the weights of the vector clusters in the selected subset comprises calculating the weights of the vector clusters in the selected subset based on values of most dominant dimensions of the descriptive vectors in the vector clusters in the selected sub set,

the first vector cluster having a first centroid vector among the centroid vectors,

the first calculated weight being calculated based on a ratio of a most dominant value of a most dominant dimension in the first centroid vector of the first vector cluster to a sum of less dominant values of less dominant dimensions in the first centroid vector of the first vector cluster.

14. The method of claim 1 , further comprising:

generating labels that identify the vector clusters in the selected subset of the hierarchy,

a first label of the generated labels identifying a first vector cluster in the selected subset,

the first vector cluster having a first centroid vector among the centroid vectors,

wherein generating the first label comprises:

determining a set of most dominant dimensions in the first centroid vector of the first vector cluster, the set of most dominant dimensions having most dominant values in the first centroid vector;

accessing a database that maps the set of most dominant dimensions to corresponding textual descriptors; and

incorporating the textual descriptors into the first label.

15. The method of claim 1 , wherein:

the descriptive vectors that describe the items are mood vectors that describe media items all recorded by a same artist, each mood vector indicating an extent to which one or more emotions are perceivable in a respective media item of the media items;

the hierarchy of vector clusters is a nested hierarchy of mood clusters that group the mood vectors; and

the selected subset of the mood clusters corresponds to a tier among multiple tiers of the nested hierarchy, the centroid vectors of the selected mood clusters describing and representing the same artist.

16. The method of claim 1 , wherein:

the items described by the descriptive vectors have a common source;

the selected subset of the vector clusters is representative of the common source of the items; and

the method further comprises:

storing identifiers of centroid vectors of vector clusters in the selected subset, the identifiers being stored with a contemporary timestamp in an evolutionary history of items attributed to the common source.

17. The method of claim 1 , wherein:

the items described by the descriptive vectors are sourced from multiple sources that include a first source and a second source;

the selected subset of the vector clusters has a first portion that is representative of the first source of the items and has a second portion that is representative of the second source of the items; and

the method further comprises:

determining that the first source represented by the first portion of the selected subset is distinct from the second source; and

causing presentation of a notification that the first and second sources are different.

18. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

accessing descriptive vectors that describe items, each descriptive vector comprising one or more values indicative of an extent to which one or more characteristics are present in a respective item of the items;

determining one or more vector distances between one or more pairs of the descriptive vectors;

generating a hierarchy of vector clusters by clustering the descriptive vectors into the vector clusters based on the determined one or more vector distances;

determining centroid vectors of the vector clusters in the hierarchy, each centroid vector corresponding to a respective vector cluster in the hierarchy;

summing one or more inter-cluster vector distances between one or more pairs of the centroid vectors;

summing, for each of the vector clusters, one or more intra-cluster vector distances between one or more pairs of descriptive vectors;

determining a plurality of scores of the hierarchy by applying a plurality of weightings to the summed inter-cluster vector distances and the summed intra-cluster vector distances, wherein each score of the plurality of scores corresponds to a respective weighting of the plurality of weightings, and wherein a particular weighting of the plurality of weightings corresponds to an extreme score of the plurality of scores; and

selecting a subset of the vector clusters in the hierarchy based on the weighting that corresponds to the extreme score.

19. The non-transitory machine-readable storage medium of claim 18 , wherein:

selecting the subset of the vector clusters in the hierarchy comprises determining that the weighting that corresponds to the extreme score corresponds to a minimum score among the plurality of scores; and

the selected subset of the vector clusters corresponds to a tier among multiple tiers of the hierarchy.

20. A system comprising:

one or more processors; and

a memory storing instructions that, when executed by at least one processor among the one or more processors, cause the system to perform operations comprising:

accessing descriptive vectors that describe items, each descriptive vector comprising one or more values indicative of an extent to which one or more characteristics are present in a respective item of the items;

determining one or more vector distances between one or more pairs of the descriptive vectors;

generating a hierarchy of vector clusters by clustering the descriptive vectors into the vector clusters based on the determined one or more vector distances;

determining centroid vectors of the vector clusters in the hierarchy, each centroid vector corresponding to a respective vector cluster in the hierarchy;

summing one or more inter-cluster vector distances between one or more pairs of the centroid vectors;

summing for each of the vector clusters, one or more intra-cluster vector distances between one or more pairs of descriptive vectors;

determining a plurality of scores of the hierarchy by applying a plurality of weightings to the summed inter-cluster vector distances and the summed intra-cluster vector distances, wherein each score of the plurality of scores corresponds to a respective weighting of the plurality of weightings, and wherein a particular weighting of the plurality of weightings corresponds to an extreme score of the plurality of scores; and

selecting a subset of the vector clusters in the hierarchy based on the weighting that corresponds to the extreme score.

Assignments (8)
RELEASE (REEL 054066 / FRAME 0064) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063605/0001 →
RELEASE (REEL 053473 / FRAME 0001) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063603/0001 →
SECURITY INTEREST Recorded May 8, 2023
From: GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE, INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC
To: ARES CAPITAL CORPORATION
Reel/Frame 063574/0632 →
SECURITY INTEREST Recorded Apr 28, 2023
From: GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE, INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC
To: CITIBANK, N.A.
Reel/Frame 063561/0381 →
SECURITY AGREEMENT Recorded Jan 31, 2023
From: GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE, INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 063560/0547 →
CORRECTIVE ASSIGNMENT TO CORRECT THE PATENTS LISTED ON SCHEDULE 1 RECORDED ON 6-9-2020 PREVIOUSLY RECORDED ON REEL 053473 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE SUPPLEMENTAL IP SECURITY AGREEMENT. Recorded Oct 7, 2020
From: A.C. NIELSEN (ARGENTINA) S.A.; A.C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A
Reel/Frame 054066/0064 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Jun 9, 2020
From: A. C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NIELSEN UK FINANCE I, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A.
Reel/Frame 053473/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2019
From: VARTAKAVI, ANEESH; DIMARIA, PETER C.; CREMER, MARKUS K.; POPP, PHILLIP
To: GRACENOTE, INC.
Reel/Frame 048016/0732 →