IP Library Granted Patent US 10,891,513
Granted Patent B2
US 10,891,513 · App. 16/562,825 · Granted Jan 12, 2021

System and method for cascading image clustering using distribution over auto-generated labels

Inventors: Andrew J. Yeager (Mountain View, CA); Ji Fang (Mountain View, CA)
Assignee: Medallia, Inc.
G06K9/6222G06K9/627G06K9/6256G06K9/6277G06K9/726G06Q30/0282G06K9/6263G06K2209/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,891,513
App. No.
16/562,825
Granted
Jan 12, 2021
Kind
B2
Abstract

Embodiments of the present invention provide a system that can be used to classify a feedback image in a user review into a semantically meaningful class. During operation, the system analyzes the captions of feedback images in a set of user reviews and determines a set of training labels from the captions. The system then trains an image classifier with the set of training labels and the feedback images. Subsequently, the system generates a signature for a respective feedback image in a new set of user reviews using the image classifier. The signature indicates a likelihood of the image matching a respective label in the set of training labels. Based on the signature, the system can allocate the image to an image cluster.

Claims (40)

1. A computer-implemented method for facilitating cascading image clustering, the method comprising:

selecting a plurality of seed clusters from a dataset comprising a set of reviews, wherein the set of reviews comprises images, wherein selecting the plurality of seed clusters comprises:

forming an initial seed cluster comprising a randomly selected initial image and a set of neighbor images of the initial image;

forming, in sequence, a plurality of next seed clusters until a total number of seed clusters, including the plurality of next seed clusters and the initial seed cluster, exceeds a threshold, wherein each next seed cluster comprises a next image and a set of neighbor images of the next image, wherein the next image being used in the next cluster currently being formed has the largest average distance from images associated with the initial seed cluster and images associated with any next seed clusters that have already been formed, and wherein the initial seed cluster and the plurality of next seed clusters comprise the plurality of seed clusters;

merging the plurality of seed clusters based on cluster pairing criteria and merging criteria to generate at least one converged cluster; and

classifying each of the at least one converged cluster with a semantically meaningful category.

2. The method of claim 1 , wherein the merging comprises:

iteratively selecting pairs of seed clusters; and

merging the selected pairs of seed clusters when the selected pairs of clusters satisfy a merge condition.

3. The method of claim 2 , wherein iteratively selecting pairs of seed clusters comprises selecting two seed clusters determined to be most similar to each other.

4. The method of claim 3 , further comprises determining similarity between two seed clusters by determining an average distance between images respectively associated with the two seed clusters.

5. The method of claim 2 , wherein the merge condition requires that an average distance between images associated with the selected pair of seed clusters be below a threshold.

6. The method of claim 1 , wherein the merging comprises:

merging a first selected pair of seed clusters to a first merged cluster, the first select pair of seed clusters comprising a first cluster and a second cluster;

merging a second selected pair of seed clusters to a second merged cluster, the second select pair of seed clusters comprising a third cluster and a fourth cluster; and

merging the first merged cluster and the second merged cluster into a third merged cluster if a first distance average between the first cluster and the third cluster, a second distance average between the first cluster and the fourth cluster, a third distance average between the second cluster and the third cluster, and a fourth distance average between the second cluster and the fourth cluster are each below a threshold.

7. The method of claim 6 , wherein the third merged cluster is one of the at least one converged cluster.

8. The method of claim 1 , further comprising generating a binary tree based on the at least one converged cluster.

9. The method of claim 8 , wherein the binary tree comprises a root, wherein the root is labeled with the semantically meaningful category.

10. The method of claim 9 , wherein traversal down the binary tree represents finer levels of granularity.

11. The method of claim 8 , further comprising:

receiving a new image that is classified in the semantically meaningful category; and

adding the new image to the binary tree.

12. A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method, the method comprising:

selecting a plurality of seed clusters from a dataset comprising a set of reviews, wherein the set of reviews comprises images, wherein selecting the plurality of seed clusters comprises:

forming an initial seed cluster comprising a randomly selected initial image and a set of neighbor images of the initial image;

forming, in sequence, a plurality of next seed clusters until a total number of seed clusters, including the plurality of next seed clusters and the initial seed cluster, exceeds a threshold, wherein each next seed cluster comprises a next image and a set of neighbor images of the next image, wherein the next image being used in the next cluster currently being formed has the largest average distance from images associated with the initial seed cluster and images associated with any next seed clusters that have already been formed, and wherein the initial seed cluster and the plurality of next seed clusters comprise the plurality of seed clusters;

merging the plurality of seed clusters based on cluster pairing criteria and merging criteria to generate at least one converged cluster; and

classifying each of the at least one converged cluster with a semantically meaningful category.

13. The computer-readable storage medium of claim 12 , wherein the merging comprises:

iteratively selecting pairs of seed clusters; and

merging the selected pairs of seed clusters when the selected pairs of clusters satisfy a merge condition.

14. The computer-readable storage medium of claim 13 , wherein iteratively selecting pairs of seed clusters comprises selecting two seed clusters determined to be most similar to each other.

15. The computer-readable storage medium of claim 14 , further comprises determining similarity between two seed clusters by determining an average distance between images respectively associated with the two seed clusters.

16. The computer-readable storage medium of claim 13 , wherein the merge condition requires that an average distance between images associated with the selected pair of seed clusters be below a threshold.

17. The computer-readable storage medium of claim 12 , wherein the merging comprises:

merging a first selected pair of seed clusters to a first merged cluster, the first select pair of seed clusters comprising a first cluster and a second cluster;

merging a second selected pair of seed clusters to a second merged cluster, the second select pair of seed clusters comprising a third cluster and a fourth cluster; and

merging the first merged cluster and the second merged cluster into a third merged cluster if a first distance average between the first cluster and the third cluster, a second distance average between the first cluster and the fourth cluster, a third distance average between the second cluster and the third cluster, and a fourth distance average between the second cluster and the fourth cluster are each below a threshold.

18. The computer-readable storage medium of claim 12 , wherein the third merged cluster is one of the at least one converged cluster.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Apr 13, 2022
From: WELLS FARGO BANK NA
To: MEDALLION, INC
Reel/Frame 059581/0865 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE LIST OF PATENT PROPERTY NUMBER TO INCLUDE TWO PATENTS THAT WERE MISSING FROM THE ORIGINAL FILING PREVIOUSLY RECORDED AT REEL: 057968 FRAME: 0430. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 1, 2021
From: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS ADMINISTRATIVE AGENT
To: MEDALLIA, INC.
Reel/Frame 057982/0092 →
SECURITY INTEREST Recorded Oct 29, 2021
From: MEDALLIA, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 057964/0016 →
RELEASE OF SECURITY INTEREST Recorded Oct 29, 2021
From: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS ADMINISTRATIVE AGENT
To: MEDALLIA, INC.
Reel/Frame 057968/0430 →
SECURITY INTEREST Recorded Jul 28, 2021
From: MEDALLIA, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 057011/0012 →
Continuity (2)
Continuation 15669800 · Aug 4, 2017
Related Publication 20200210760A1 · Jul 2, 2020