IP Library Granted Patent US 11,544,510
Granted Patent B2
US 11,544,510 · App. 16/509,062 · Granted Jan 3, 2023

System and method for multi-modal image classification

Inventors: Yogen Chaudhari (Reston, VA); Sean Pinkney (Reston, VA); Prashanth Venkatraman (Reston, VA); Ashwath Rajendran (Reston, VA); Jay Parikh (Reston, VA)
Assignee: Comscore, Inc.
G06K9/628G06F40/10G06K9/6256G06K9/6264G06K9/6269G06N3/08G06Q30/0201G06Q30/0241G06V30/153
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,544,510
App. No.
16/509,062
Granted
Jan 3, 2023
Kind
B2
Abstract

Systems and methods for classifying images (e.g., ads) are described. An image is accessed. Optical character recognition is performed on at least a first portion of the image. Image recognition is performed via a convolutional neural network on at least a second portion of the image. At least one class for the image is automatically identified, via a fully connected neural network, based on one or more predictions, each of the one or more predictions being based on both the optical character recognition and the image recognition. Finally, the at least one class identified for the image is output.

Claims (63)

1. A computer-implemented method, comprising:

preprocessing an image to identify a plurality of elements from the image;

performing optical character recognition on at least a first element of the plurality of elements from the image to identify text in the first element;

converting the text in the first element into a term feature vector;

performing image recognition, via a convolutional neural network, on at least a second element of the plurality of elements from the image to identify an image feature;

predicting, via a fully connected neural network and a concatenation of the term feature vector and the image feature, (i) that an advertisement is present in the image, (ii) a category of the advertisement, and (iii) a brand based on the category of the advertisement; and

outputting the category of the advertisement.

2. The method of claim 1 , wherein preprocessing the image comprises:

eroding or dilating thickness of the plurality of elements; or

adjusting contrast in and around the plurality of elements.

3. The method of claim 1 , further comprising:

cleaning, via a natural language processing, the text in the first element.

4. The method of claim 1 , further comprising:

resizing the image such that the text in the plurality of elements is recognizable by the convolutional neural network.

5. The method of claim 4 ,

wherein the convolutional neural network determines edges or boundaries of the plurality of elements within the resized image.

6. The method of claim 5 , wherein the image feature comprises a logo of the brand.

7. The method of claim 1 , wherein:

the fully connected neural network comprises a plurality of layers, each of the layers comprising a plurality of artificial neurons,

each of the neurons performs one or more calculations using one or more parameters, and

each of the neurons of a first layer of the plurality of layers is connected to each of the neurons of a second layer of the plurality of layers.

8. The method of claim 7 , wherein:

the convolutional neural network comprises an input layer, a plurality of hidden layers, and an output layer,

the plurality of hidden layers comprises a convolutional layer, a rectified linear unit layer, and a pooling layer,

neurons of each of the layers transforms a three-dimensional input volume to a three-dimensional output volume of neuron activations, and

the neurons of one of the layers do not connect to all of the neurons of another of the layers.

9. The method of claim 1 , wherein preprocessing the image comprises manipulating the image via a function obtained from an image library.

10. The method of claim 1 , further comprising retrieving the image from a remote server.

11. The method of claim 1 , wherein the category of the advertisement is at least one of travel, automotive, finance, consumer goods, retail, restaurants, sports, entertainment, telecommunications, healthcare, insurance, computer technology, education, or business services.

12. A system comprising one or more processors coupled to one or more storage devices storing instructions that, when executed, cause the one or more processors to perform the following operations:

preprocessing an image to identify a plurality of elements from the image;

performing optical character recognition on at least a first element of the plurality of elements from the image to identify text in the first element;

converting the text in the first element into a term feature vector;

performing image recognition via a convolutional neural network, on at least a second element of the plurality of elements from the image to identify an image feature;

predicting, via a fully connected neural network and a concatenation of the term feature vector and the image feature, (i) that an advertisement is present in the image, (ii) a category of the advertisement, and (iii) a brand based on the category of the advertisement; and

outputting the category of the advertisement.

13. The system of claim 12 , wherein:

the fully connected neural network comprises a plurality of layers, each of the layers comprising a plurality of artificial neurons,

each of the neurons performs one or more calculations using one or more parameters, and

each of the neurons of a first layer of the plurality of layers is connected to each of the neurons of a second layer of the plurality of layers.

14. The system of claim 13 , wherein:

the convolutional neural network comprises an input layer, a plurality of hidden layers, and an output layer,

the plurality of hidden layers comprises a convolutional layer, a rectified linear unit layer, and a pooling layer,

neurons of each of the layers transforms a three-dimensional input volume to a three-dimensional output volume of neuron activations, and

the neurons of one of the layers do not connect to all of the neurons of another of the layers.

15. The system of claim 12 , wherein the category of the advertisement is at least one of travel, automotive, finance, consumer goods, retail, restaurants, sports, entertainment, telecommunications, healthcare, insurance, computer technology, education, or business services.

16. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method, the method comprising:

preprocessing an image to identify a plurality of elements from the image;

performing optical character recognition on at least a first element of the plurality of elements from the image to identify text in the first element;

converting the text in the first element into a term feature vector;

performing image recognition, via a convolutional neural network, on at least a second element of the plurality of elements from the image to identify an image feature;

predicting, via a fully connected neural network and a concatenation of the term feature vector and the image feature, (i) that an advertisement is present in the image, (ii) a category of the advertisement, and (iii) a brand based on the category of the advertisement; and

outputting the category of the advertisement.

17. The medium of claim 16 , wherein:

the fully connected neural network comprises a plurality of layers, each of the layers comprising a plurality of artificial neurons,

each of the neurons performs one or more calculations using one or more parameters, and

each of the neurons of a first layer of the plurality of layers is connected to each of the neurons of a second layer of the plurality of layers.

18. The medium of claim 17 , wherein:

the convolutional neural network comprises an input layer, a plurality of hidden layers, and an output layer,

the plurality of hidden layers comprises a convolutional layer, a rectified linear unit layer, and a pooling layer,

neurons of each of the layers transforms a three-dimensional input volume to a three-dimensional output volume of neuron activations, and

the neurons of one of the layers do not connect to all of the neurons of another of the layers.

19. The medium of claim 16 , wherein the category of the advertisement is at least one of travel, automotive, finance, consumer goods, retail, restaurants, sports, entertainment, telecommunications, healthcare, insurance, computer technology, education, or business services.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Jun 2, 2026
From: BLUE TORCH FINANCE LLC
To: COMSCORE, INC.; PROXIMIC, LLC; RENTRAK, LLC (F/N/A RENTRAK CORPORATION)
Reel/Frame 075679/0830 →
RELEASE OF SECURITY INTEREST Recorded Jan 16, 2025
From: BANK OF AMERICA, N.A.
To: COMSCORE, INC.
Reel/Frame 069934/0573 →
SECURITY INTEREST Recorded Jan 3, 2025
From: COMSCORE, INC.; PROXIMIC, LLC; RENTRAK, LLC
To: BLUE TORCH FINANCE LLC
Reel/Frame 069818/0446 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded May 6, 2021
From: COMSCORE, INC.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 057279/0767 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2019
From: CHAUDHARI, YOGEN; PINKNEY, SEAN; VENKATRAMAN, PRASHANTH; RAJENDRAN, ASHWATH; PARIKH, JAY
To: COMSCORE, INC.
Reel/Frame 049729/0220 →