IP Library Granted Patent US 9,355,330
Granted Patent B2
US 9,355,330 · App. 14/111,149 · Granted May 31, 2016

In-video product annotation with web information mining

Inventors: Tat Seng Chua (Singapore, SG); Guangda Li (Singapore, SG); Zheng Lu (Singapore, SG); Meng Wang (Singapore, SG)
Assignee: National University of Singapore
G06K9/4642G06F17/30799G06K9/00744
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,355,330
App. No.
14/111,149
Granted
May 31, 2016
Kind
B2
Abstract

A system provides product annotation in a video to one or more users. The system receives a video from a user, where the video includes multiple video frames. The system extracts multiple key frames from the video and generates a visual representation of the key frame. The system compares the visual representation of the key frame with a plurality of product visual signatures, where each visual signature identifies a product. Based on the comparison of the visual representation of the key frame and a product visual signature, the system determines whether the key frame contains the product identified by the visual signature of the product. To generate the plurality of product visual signatures, the system collects multiple training images comprising multiple of expert product images obtained from an expert product repository, each of which is associated with multiple product images obtained from multiple web resources.

Claims (54)

1. A computer method for providing product annotation in a video to one or more users, the method comprising:

generating a product visual signature for a product by at least:

collecting an unannotated expert product image of the product from an expert product repository,

searching for a plurality of unannotated product images from a plurality of web resources different from the expert product repository, the plurality of unannotated product images related to the unannotated expert product image,

selecting a subset of the plurality of unannotated product images by filtering the plurality of unannotated product images based on a similarity measure to the unannotated expert product image, and

generating the product visual signature from the unannotated expert product image and the subset of the plurality of unannotated product images;

receiving a video for product annotation, the video comprising a plurality of video frames;

extracting a plurality of key frames from the video frames; and

for each key frame:

generating a visual representation of the key framed;

comparing the visual representation with a plurality of product visual signatures including the product visual signature; and

determining, based on the comparison, that the key frame contains the product identified by the product visual signature.

2. The method of claim 1 , wherein extracting a plurality of key frames from the video comprises:

extracting each of the plurality of key frames at a fixed point of the video.

3. The method of claim 1 , wherein generating the visual signature of a key frame comprises:

extracting a plurality of visual features from the key frame;

grouping the plurality of visual features into a plurality of clusters; and

generating multi-dimensional bag visual words histogram as the visual signature of the key frame.

4. The method of claim 3 , wherein the plurality of visual features of a key frame are scale invariance feature transform (SIFT) descriptors of the key frame.

5. The method of claim 1 , wherein generating the subset of the plurality of unannotated product images represent a set of training images for generating the product visual signature.

6. The method of claim 1 , wherein generating the product visual signature further comprises:

applying a collective sparsification scheme to the, subset of the plurality of unannotated product images, wherein information unrelated to the product contained in a related product image is reduced in generating the product visual signature.

7. The method of claim 1 , wherein generating the product visual signature further comprises:

iteratively updating the product visual signature through a pre-determined number of iterations, wherein each of the iterations computes a respective similarity measure.

8. The method of claim 1 , further comprising: collecting a plurality of unannotated expert product images of the product at different views of the product, wherein the subset of the subset of the plurality of unannotated product images comprise unannotated product images corresponding to the plurality of unannotated expert product images.

9. The method of claim 1 , wherein determining that the key frame contains the product identified comprises:

estimating product relevance between the visual representation of the key frame with each product visual signature of the plurality of the product visual signatures; and

determining that the key frame contains the product identified by the product visual signature based on the estimated product relevance.

10. A non-transitory computer-readable storage medium storing executable computer program instructions for providing on-demand digital assets hosting services to one or more users, the computer program instructions when executed by a processor cause a system to perform operations comprising:

generating a product visual signature for a product by at least:

collecting an unannotated expert product image of the product from an expert product repository,

searching for a plurality of unannotated product images from a plurality of web resources different from the expert product repository, the plurality of unannotated product images related to the unannotated expert product image,

selecting a subset of the plurality of unannotated product images by filtering the plurality of unannotated product images based on a similarity measure to the unannotated expert product image, and

generating the product visual signature from the unannotated expert product image and the subset of the plurality of unannotated product images;

receiving a video from a user for product annotation, the video comprising a plurality of video frames;

extracting a plurality of key frames from the video; and

for each key frame:

extracting a plurality of visual features from the key frame;

grouping the plurality of visual features into a plurality of clusters; and

generating a multi-dimensional bag visual words histogram as a visual representation of the key frame;

comparing the visual representation with a plurality of product visual signatures comprising the product visual signature;

determining, based on the comparison, whether the key frame contains the product identified by the product visual signature.

11. The computer-readable storage medium of claim 10 , wherein the operations further comprise:

extracting each of the plurality of key frames at a fixed point of the video.

12. The computer-readable storage medium of claim 10 , wherein the plurality of visual features of a key frame are scale invariance feature transform (SIFT) descriptors of the key frame.

13. The computer-readable storage medium of claim 10 , wherein generating the subset of the plurality of unannotated product images represent a set of training images for generating the product visual signature.

14. The computer-readable storage medium of claim 10 , wherein generating the product visual signature further comprises:

applying a collective sparsification scheme to the, subset of the plurality of unannotated product images, wherein information unrelated to the product contained in a related product image is reduced in generating the product visual signature.

15. The computer-readable storage medium of claim 10 , wherein generating the product visual signature further comprises:

iteratively updating the product visual signature through a pre-determined number of iterations, wherein each of the iterations computes a respective similarity measure.

16. The computer-readable storage medium of claim 10 , wherein the operations further comprise: collecting a plurality of unannotated expert product images of the product at different views of the product, wherein the subset of the subset of the plurality of unannotated product images comprise unannotated product images corresponding to the plurality of unannotated expert product images.

17. The computer-readable storage medium of claim 10 , wherein determining whether the key frame contains the product comprises:

estimating product relevance between the visual representation of the key frame with each product visual signature of the plurality of the product visual signatures; and

determining that thee key frame contains the product identified by the product visual signature based on the estimated product relevance.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2019
From: NATIONAL UNIVERSITY OF SINGAPORE
To: VISENZE PTE LTD
Reel/Frame 051289/0104 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2014
From: CHUA, TAT SENG; LI, GUANGDA; LU, ZHENG; WANG, MENG
To: NATIONAL UNIVERSITY OF SINGAPORE
Reel/Frame 032133/0713 →
Continuity (2)
Provisional Application 61474328 · Apr 12, 2011
Related Publication 20140029801A1 · Jan 30, 2014