IP Library Granted Patent US 12,437,529
Granted Patent B2
US 12,437,529 · App. 18/507,560 · Granted Oct 7, 2025

Logo recognition in images and videos

Inventors: Jose Pio Pereira (Supertino, CA); Kyle Brocklehurst (Mountain View, CA); Sunil Suresh Kulkarni (San Jose, CA); Peter Wendt (San Jose, CA)
Assignee: Gracenote, Inc.
G06V10/82G06F18/24G06T7/11G06T7/337G06T7/60G06V10/462G06V10/764G06T2207/20052G06V10/50G06V2201/09
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,529
App. No.
18/507,560
Granted
Oct 7, 2025
Kind
B2
Abstract

Accurately detection of logos in media content on media presentation devices is addressed. Logos and products are detected in media content produced in retail deployments using a camera. Logo recognition uses saliency analysis, segmentation techniques, and stroke analysis to segment likely logo regions. Logo recognition may suitably employ feature extraction, signature representation, and logo matching. These three approaches make use of neural network based classification and optical character recognition (OCR). One method for OCR recognizes individual characters then performs string matching. Another OCR method uses segment level character recognition with N-gram matching. Synthetic image generation for training of a neural net classifier and utilizing transfer learning features of neural networks are employed to support fast addition of new logos for recognition.

Claims (45)

1. A method to map a product logo to three dimensional (3D) space, the method comprising:

detecting, by a computing system, based on image analysis of a video frame of the video stream, (i) a logo of a particular brand and (ii) an associated product; and

mapping, by the computing system, the detected logo and product to a 3D environment where the associated product was located.

2. The method of claim 1 , wherein the 3D environment is an indoor or outdoor event.

3. The method of claim 1 , wherein the 3D environment is an indoor or outdoor retail display.

4. The method of claim 1 , wherein detecting the logo is based on segmentation of one or more regions of the video frame of the video stream.

5. The method of claim 1 , wherein detecting the logo comprises:

applying a saliency analysis and segmentation of one or more regions in the video frame of the video stream to determine segmented likely-logo regions;

processing the segmented likely-logo regions using feature matching to generate a first match, using neural network classification to generate a second match, and using text recognition with string matching to generate a third match;

deciding a most likely logo match based on one or more of the first match, the second match, or the third match; and

detecting the logo as the most likely logo match.

6. The method of claim 5 , wherein applying the segmentation comprises:

applying a stroke width transform (SWT) analysis to the one or more regions to generate SWT statistics;

applying a graph-based segmentation algorithm to establish word boxes around likely logo character strings; and

analyzing each of the word boxes to produce a set of character segmentations to delineate characters in the likely logo character strings.

7. The method of claim 1 , further comprising segmenting, by the computing system, the detected logo and the associated product to determine the particular brand.

8. A computing system comprising:

one or more processors; and

non-transitory, computer-readable storage, storing instructions executable by the one or more processors to carry out operations for mapping a product logo to physical space, the operations including:

detecting based on image analysis of a video frame of the video stream, (i) a logo of a particular brand and (ii) an associated product, and

mapping the detected logo and product to a 3D environment where the associated product was located.

9. The computing system of claim 8 , wherein the 3D environment is an indoor or outdoor event.

10. The computing system of claim 8 , wherein the 3D environment is an indoor or outdoor retail display.

11. The computing system of claim 8 , wherein detecting the logo is based on segmentation of one or more regions of the video frame of the video stream.

12. The computing system of claim 8 , wherein detecting the logo comprises:

applying a saliency analysis and segmentation of one or more regions in the video frame of the video stream to determine segmented likely-logo regions;

processing the segmented likely-logo regions using feature matching to generate a first match, using neural network classification to generate a second match, and using text recognition with string matching to generate a third match;

deciding a most likely logo match based on one or more of the first match, the second match, or the third match; and

detecting the logo as the most likely logo match.

13. The computing system of claim 12 , wherein applying the segmentation comprises:

applying a stroke width transform (SWT) analysis to the one or more regions to generate SWT statistics;

applying a graph-based segmentation algorithm to establish word boxes around likely logo character strings; and

analyzing each of the word boxes to produce a set of character segmentations to delineate characters in the likely logo character strings.

14. The computing system of claim 8 , wherein the operations additionally include segmenting the detected logo and the associated product to determine the particular brand.

15. Non-transitory computer-readable storage storing having program code executable by one or more processors to cause the one or more processors to carry out operations for mapping a product logo to physical space, the operations comprising:

detecting based on image analysis of a video frame of the video stream, (i) a logo of a particular brand and (ii) an associated product; and

mapping the detected logo and product to a 3D environment where the associated product was located.

16. The non-transitory computer-readable storage of claim 15 , wherein the 3D environment is an indoor or outdoor event.

17. The non-transitory computer-readable storage of claim 15 , wherein the 3D environment is an indoor or outdoor retail display.

18. The non-transitory computer-readable storage of claim 15 , wherein detecting the logo is based on segmentation of one or more regions of the video frame of the video stream.

19. The non-transitory computer-readable storage of claim 15 , wherein detecting the logo comprises:

applying a saliency analysis and segmentation of one or more regions in the video frame of the video stream to determine segmented likely-logo regions;

processing the segmented likely-logo regions using feature matching to generate a first match, using neural network classification to generate a second match, and using text recognition with string matching to generate a third match;

deciding a most likely logo match based on one or more of the first match, the second match, or the third match; and

detecting the logo as the most likely logo match.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: PEREIRA, JOSE PIO; BROCKLEHURST, KYLE; KULKARNI, SUNIL SURESH; WENDT, PETER
To: GRACENOTE, INC.
Reel/Frame 065545/0953 →
Continuity (6)
Continuation 17672963 · Feb 16, 2022
Continuation 16841681 · Apr 7, 2020
Continuation 16018011 · Jun 25, 2018
Division 15172826 · Jun 3, 2016
Provisional Application 62171820 · Jun 5, 2015
Related Publication 20240096082A1 · Mar 21, 2024
References Cited (26)
US 5287272A · Rutenberg et al. · 1994 [cited by applicant]
US 6226041B1 · Florencio et al. · 2001 [cited by applicant]
US 6282317B1 · Luo et al. · 2001 [cited by applicant]
US 6714924B1 · McClanahan · 2004 [cited by applicant]
US 7593865B2 · Cirulli et al. · 2009 [cited by applicant]
US 8171030B2 · Pereira et al. · 2012 [cited by applicant]
US 8189945B2 · Stojancic et al. · 2012 [cited by applicant]
US 8195689B2 · Ramanathan et al. · 2012 [cited by applicant]
US 8229227B2 · Stojancic et al. · 2012 [cited by applicant]
US 8335786B2 · Pereira et al. · 2012 [cited by applicant]
US 8655878B1 · Kulkarni et al. · 2014 [cited by applicant]
US 8959108B2 · Pereira et al. · 2015 [cited by applicant]
US 9158995B2 · Rodriguez-Serrano et al. · 2015 [cited by applicant]
US 9628837B2 · Davidson et al. · 2017 [cited by applicant]
US 9666116B2 · Jung et al. · 2017 [cited by applicant]
US 11288823B2 · Pereira · 2022 [cited by examiner]
US 11861888B2 · Pereira · 2024 [cited by examiner]
US 20020097436A1 · Yokoyama et al. · 2002 [cited by applicant]
US 20030076448A1 · Pan et al. · 2003 [cited by applicant]
US 20100290701A1 · Puneet et al. · 2010 [cited by applicant]
US 20140079321A1 · Huynh-Thu et al. · 2014 [cited by applicant]
US 20150062197A1 · Jung et al. · 2015 [cited by applicant]
US 20160042253A1 · Sawhney et al. · 2016 [cited by applicant]
EP 2259207 · 2012 [cited by applicant]
Richard J.M. Den Hollander et al., “Logo Recognition in Video Stills by String Matching”, Delft University of Technology, Oct. 2003 (4 pages). [cited by applicant]
Goncalo Filipe Palaio Oliveira, “Sabado—Smart Brand Detection,” Sep. 2, 2015 (92 pages). [cited by applicant]