IP Library Granted Patent US 11,288,823
Granted Patent B2
US 11,288,823 · App. 16/841,681 · Granted Mar 29, 2022

Logo recognition in images and videos

Inventors: Jose Pio Pereira (Cupertino, CA); Kyle Brocklehurst (Mountain View, CA); Sunil Suresh Kulkarni (San Jose, CA); Peter Wendt (San Jose, CA)
Assignee: Gracenote, Inc.
G06T7/337G06K9/4671G06K9/6267G06T7/11G06T7/60G06K9/4642G06K2009/4666G06K2209/25G06T2207/20052
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,288,823
App. No.
16/841,681
Granted
Mar 29, 2022
Kind
B2
Abstract

Methods, apparatus, systems and articles of manufacture of logo recognition in images and videos are disclosed. An example method to detect a specific brand in images and video streams comprises accepting luminance images at a scale in an x direction Sx and a different scale in a y direction Sy in a neural network, and training the neural network with a set of training images for detected features associated with a specific brand.

Claims (78)

1. A method to detect a particular brand in images from a video stream, the method comprising:

segmenting and measuring a detected logo and its associated product to determine a brand;

identifying a physical location of the detected logo and product;

classifying the logo as being on a wearable product, located on a banner, or on a fixture; and

mapping the product and brand to a three-dimensional (3D) map of an event where the logo and product were detected.

2. The method of claim 1 , wherein:

the logo and product are located in a retail display;

identifying the physical location of the logo and product comprises identifying a physical location of the retail display; and

the 3D map of the event identifies the physical location of the retail display.

3. The method of claim 1 , wherein:

the logo and product are located for display in an outdoor setting;

identifying the physical location of the logo and product comprises identifying a physical location in the outdoor setting; and

the 3D map of the event identifies the physical location in the outdoor setting.

4. The method of claim 1 , further comprising detecting the logo in the images by:

applying a saliency analysis and segmentation of selected regions in a selected video frame of the video stream to determine segmented likely logo regions;

processing the segmented likely logo regions with feature matching using correlation to generate a first match, neural network classification using a convolutional neural network to generate a second match, and text recognition using character segmentation and string matching to generate a third match;

deciding a most likely logo match by combining results from the first match, the second match, and the third match; and

detecting the logo as the most likely logo match.

5. The method of claim 4 , wherein applying the saliency analysis comprises:

applying a discrete cosine transform (DCT) on the segmented likely logo regions of an image in the selected video frame to determine spectral saliency of each segmented likely logo region.

6. The method of claim 5 , wherein applying the saliency analysis further comprises:

measuring multi-scale similarity at two higher scales and a smaller scale of the spectral saliency of each likely logo region, wherein the multi-scale similarity measures include orientation gradient histograms, hue, saturation, value (HSV) histograms, and stroke width transform (SWT) statistics including total number of strokes, number of horizontal strokes, number of vertical strokes, stroke density, and number of loops.

7. The method of claim 4 , wherein applying the segmentation comprises:

applying a stroke width transform (SWT) analysis to the selected regions to generate SWT statistics;

applying a graph-based segmentation algorithm to establish word boxes around likely logo character strings; and

analyzing each of the word boxes to produce a set of character segmentations to delineate the characters in the likely logo character strings.

8. A computing system comprising:

one or more processors; and

a non-transitory, computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform a set of operations for detecting a particular brand in images from a video stream, the set of operations comprising:

segmenting and measuring a detected logo and its associated product to determine a brand;

identifying a physical location of the detected logo and product;

classifying the logo as being on a wearable product, located on a banner, or on a fixture; and

mapping the product and brand to a three-dimensional (3D) map of an event where the logo and product were detected.

9. The computing system of claim 8 , wherein:

the logo and product are located in a retail display;

identifying the physical location of the logo and product comprises identifying a physical location of the retail display; and

the 3D map of the event identifies the physical location of the retail display.

10. The computing system of claim 8 , wherein:

the logo and product are located for display in an outdoor setting;

identifying the physical location of the logo and product comprises identifying a physical location in the outdoor setting; and

the 3D map of the event identifies the physical location in the outdoor setting.

11. The computing system of claim 8 , the set of operations further comprising detecting the logo in the images by:

applying a saliency analysis and segmentation of selected regions in a selected video frame of the video stream to determine segmented likely logo regions;

processing the segmented likely logo regions with feature matching using correlation to generate a first match, neural network classification using a convolutional neural network to generate a second match, and text recognition using character segmentation and string matching to generate a third match;

deciding a most likely logo match by combining results from the first match, the second match, and the third match; and

detecting the logo as the most likely logo match.

12. The computing system of claim 11 , wherein applying the saliency analysis comprises:

applying a discrete cosine transform (DCT) on the segmented likely logo regions of an image in the selected video frame to determine spectral saliency of each segmented likely logo region.

13. The computing system of claim 12 , wherein applying the saliency analysis further comprises:

measuring multi-scale similarity at two higher scales and a smaller scale of the spectral saliency of each likely logo region, wherein the multi-scale similarity measures include orientation gradient histograms, hue, saturation, value (HSV) histograms, and stroke width transform (SWT) statistics including total number of strokes, number of horizontal strokes, number of vertical strokes, stroke density, and number of loops.

14. The computing system of claim 11 , wherein applying the segmentation comprises:

applying a stroke width transform (SWT) analysis to the selected regions to generate SWT statistics;

applying a graph-based segmentation algorithm to establish word boxes around likely logo character strings; and

analyzing each of the word boxes to produce a set of character segmentations to delineate the characters in the likely logo character strings.

15. A non-transitory, computer readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform a set of operations for detecting a particular brand in images from a video stream, the set of operations comprising:

segmenting and measuring a detected logo and its associated product to determine a brand;

identifying a physical location of the detected logo and product;

classifying the logo as being on a wearable product, located on a banner, or on a fixture; and

mapping the product and brand to a three-dimensional (3D) map of an event where the logo and product were detected.

16. The non-transitory, computer readable medium of claim 15 , wherein:

the logo and product are located in a retail display;

identifying the physical location of the logo and product comprises identifying a physical location of the retail display; and

the 3D map of the event identifies the physical location of the retail display.

17. The non-transitory, computer readable medium of claim 15 , wherein:

the logo and product are located for display in an outdoor setting;

identifying the physical location of the logo and product comprises identifying a physical location in the outdoor setting; and

the 3D map of the event identifies the physical location in the outdoor setting.

18. The non-transitory, computer readable medium of claim 15 , the set of operations further comprising detecting the logo in the images by:

applying a saliency analysis and segmentation of selected regions in a selected video frame of the video stream to determine segmented likely logo regions;

processing the segmented likely logo regions with feature matching using correlation to generate a first match, neural network classification using a convolutional neural network to generate a second match, and text recognition using character segmentation and string matching to generate a third match;

deciding a most likely logo match by combining results from the first match, the second match, and the third match; and

detecting the logo as the most likely logo match.

19. The non-transitory, computer readable medium of claim 18 , wherein applying the saliency analysis comprises:

applying a discrete cosine transform (DCT) on the segmented likely logo regions of an image in the selected video frame to determine spectral saliency of each segmented likely logo region.

20. The non-transitory, computer readable medium of claim 18 , wherein applying the segmentation comprises:

applying a stroke width transform (SWT) analysis to the selected regions to generate SWT statistics;

applying a graph-based segmentation algorithm to establish word boxes around likely logo character strings; and

analyzing each of the word boxes to produce a set of character segmentations to delineate the characters in the likely logo character strings.

Assignments (8)
RELEASE (REEL 054066 / FRAME 0064) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063605/0001 →
RELEASE (REEL 053473 / FRAME 0001) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063603/0001 →
SECURITY INTEREST Recorded May 8, 2023
From: GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE, INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC
To: ARES CAPITAL CORPORATION
Reel/Frame 063574/0632 →
SECURITY INTEREST Recorded Apr 28, 2023
From: GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE, INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC
To: CITIBANK, N.A.
Reel/Frame 063561/0381 →
SECURITY AGREEMENT Recorded Jan 31, 2023
From: GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE, INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 063560/0547 →
CORRECTIVE ASSIGNMENT TO CORRECT THE PATENTS LISTED ON SCHEDULE 1 RECORDED ON 6-9-2020 PREVIOUSLY RECORDED ON REEL 053473 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE SUPPLEMENTAL IP SECURITY AGREEMENT. Recorded Oct 7, 2020
From: A.C. NIELSEN (ARGENTINA) S.A.; A.C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A
Reel/Frame 054066/0064 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Jun 9, 2020
From: A. C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NIELSEN UK FINANCE I, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A.
Reel/Frame 053473/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2020
From: PEREIRA, JOSE PIO; BROCKLEHURST, KYLE; KULKARNI, SUNIL SURESH; WENDT, PETER
To: GRACENOTE, INC.
Reel/Frame 052344/0976 →
Continuity (4)
Continuation 16018011 · Jun 25, 2018
Division 15172826 · Jun 3, 2016
Provisional Application 62171820 · Jun 5, 2015
Related Publication 20200372662A1 · Nov 26, 2020
Cited By (1)
US 12,437,529