IP Library Granted Patent US 12,205,387
Granted Patent B2
US 12,205,387 · App. 18/429,066 · Granted Jan 21, 2025

System and method for using artificial intelligence (AI) to analyze social media content

Inventor: Allen O'Neill (Carlow, IE)
Assignee: Social Voice Ltd.
G06V30/10G06V10/25G06V20/49G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,387
App. No.
18/429,066
Granted
Jan 21, 2025
Kind
B2
Abstract

Systems and methods for determining attributes in media content. A computing device may be configured to obtain the media content based on a received media content identifier and segment the media content into scenes. The computing device may analyze viewer engagement metrics to identify a scene associated with a viewer engagement score that exceeds a threshold, select video frames from the identified scene, and identify primary objects in the series of images in the scene. The computing device may add a bounding box around the identified primary objects in one or more selected frames and perform text extraction within the bounding box. The computing device may determine object attributes of the identified primary objects, querying a database to identify topics of interest (ToIs) based on the extracted text and the determined object attributes, and performing a responsive action in response to identifying the one or more ToIs.

Claims (61)

1. A method of reducing a search space by determining attributes in media content, comprising:

receiving a media content identifier;

obtaining the media content based on the received media content identifier;

segmenting the media content into scenes based on visual and audio cues, wherein each of the scenes includes series of images;

analyzing viewer engagement metrics across the scenes to identify a scene associated with a viewer engagement score that exceeds a threshold;

selecting video frames from the identified scene;

identifying primary objects in each selected frame that occur across the timespan of the series of images in the scene;

adding a bounding box around the identified primary objects in one or more selected frame;

performing text extraction within the bounding box;

determining object attributes of the identified primary objects;

querying a database to identify topics of interest (ToIs) based on the extracted text and the determined object attributes; and

performing a responsive action in response to identifying the one or more ToIs.

2. The method of claim 1 , wherein identifying the primary objects in each selected frame that occur across the timespan of the series of images in the scene comprises:

using a convolutional neural network (CNN) model to identify the primary objects in each selected frame.

3. The method of claim 1 , wherein performing the text extraction within the bounding box comprises:

using optical character recognition (OCR) to perform text extraction within the bounding box.

4. The method of claim 1 , further comprising using a large generative artificial intelligence model (LXM) or another natural language processing (NLP) technique to covert non-Latin characters in the extracted text and complete incomplete data.

5. The method of claim 1 , further comprising using predictive text and word completion techniques for a partially obscured or incomplete text in the extracted text.

6. The method of claim 1 , further comprising:

retrieving external data from an external source for product identification,

wherein querying the database to identify the ToIs based on the extracted text and the determined object attributes comprises querying the database to identify the ToIs based on the extracted text, the determined object attributes, and the retrieved external data.

7. The method of claim 1 , further comprising:

generating signatures for ToIs with non-textual or obscure labels; and

using signatures to search and update a ToI knowledge repository.

8. A computing device, comprising:

a processor configured to:

receive a media content identifier;

obtain media content based on the received media content identifier;

segment the media content into scenes based on visual and audio cues, wherein each of the scenes includes series of images;

analyze viewer engagement metrics across the scenes to identify a scene associated with a viewer engagement score that exceeds a threshold;

select video frames from the identified scene;

identify primary objects in each selected frame that occur across the timespan of the series of images in the scene;

add a bounding box around the identified primary objects in one or more selected frame;

perform text extraction within the bounding box;

determine object attributes of the identified primary objects;

query a database to identify topics of interest (ToIs) based on the extracted text and the determined object attributes; and

perform a responsive action in response to identifying the one or more ToIs.

9. The computing device of claim 8 , wherein the processor is configured to identify the primary objects in each selected frame that occur across the timespan of the series of images in the scene by:

using a convolutional neural network (CNN) model to identify the primary objects in each selected frame.

10. The computing device of claim 8 , wherein the processor is configured to perform the text extraction within the bounding box by:

using optical character recognition (OCR) to perform text extraction within the bounding box.

11. The computing device of claim 8 , wherein the processor is further configured to use a large generative artificial intelligence model (LXM) or another natural language processing (NLP) technique to covert non-Latin characters in the extracted text and complete incomplete data.

12. The computing device of claim 8 , wherein the processor is further configured to use predictive text and word completion techniques for a partially obscured or incomplete text in the extracted text.

13. The computing device of claim 8 , wherein:

the processor is further configured to retrieve external data from an external source for product identification; and

the processor is configured to query the database to identify the ToIs based on the extracted text, the determined object attributes, and the retrieved external data.

14. The computing device of claim 8 , the processor is further configured to:

generate signatures for ToIs with non-textual or obscure labels; and

use signatures to search and update a ToI knowledge repository.

15. A non-transitory processor-readable medium having stored thereon processor-readable instructions configured to cause a processor in a computing device to perform operations for reducing a search space by determining attributes in media content, the operations comprising:

receiving a media content identifier;

obtaining the media content based on the received media content identifier;

segmenting the media content into scenes based on visual and audio cues, wherein each of the scenes includes series of images;

analyzing viewer engagement metrics across the scenes to identify a scene associated with a viewer engagement score that exceeds a threshold;

selecting video frames from the identified scene;

identifying primary objects in each selected frame that occur across the timespan of the series of images in the scene;

adding a bounding box around the identified primary objects in one or more selected frame;

performing text extraction within the bounding box;

determining object attributes of the identified primary objects;

querying a database to identify topics of interest (ToIs) based on the extracted text and the determined object attributes; and

performing a responsive action in response to identifying the one or more ToIs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2024
From: O'NEILL, ALLEN
To: SOCIAL VOICE LTD.
Reel/Frame 066719/0751 →
Continuity (2)
Provisional Application 63442168 · Jan 31, 2023
Related Publication 20240412542A1 · Dec 12, 2024
References Cited (32)
US 8379101B2 · Mathe · 2013 [cited by examiner]
US 8896721B2 · Mathe · 2014 [cited by examiner]
US 9934223B2 · Houh · 2018 [cited by examiner]
US 10417879B2 · Moussette · 2019 [cited by examiner]
US 11221671B2 · Stent · 2022 [cited by examiner]
US 11244167B2 · Zhao · 2022 [cited by examiner]
US 11341185B1 · Hamid · 2022 [cited by applicant]
US 11399264B2 · Zaltzman · 2022 [cited by examiner]
US 11416002B1 · Day · 2022 [cited by examiner]
US 11983913B2 · Eswara · 2024 [cited by examiner]
US 12014562B2 · Proschowsky · 2024 [cited by examiner]
US 20070118873A1 · Houh · 2007 [cited by examiner]
US 20150378998A1 · Houh et al. · 2015 [cited by applicant]
US 20160062464A1 · Moussette · 2016 [cited by examiner]
US 20160224869A1 · Clark-Polner · 2016 [cited by examiner]
US 20170213243A1 · Dollard · 2017 [cited by examiner]
US 20180286272A1 · McDermott · 2018 [cited by examiner]
US 20200249753A1 · Stent · 2020 [cited by examiner]
US 20210136537A1 · Zaltzman · 2021 [cited by examiner]
US 20210248376A1 · Zhao · 2021 [cited by examiner]
US 20220269882A1 · Proschowsky · 2022 [cited by examiner]
US 20220342930A1 · Chandrashekar et al. · 2022 [cited by applicant]
US 20220360847A1 · Chukoskie · 2022 [cited by examiner]
US 20220366665A1 · Eswara · 2022 [cited by examiner]
US 20240045558A1 · Sarkar · 2024 [cited by examiner]
WO WO2020034672A1 · 2020 [cited by examiner]
Li et al., “DIGMN: Dynamic Intent Guided Meta Network for Differentiated User Engagement Forecasting in Online Professional Social Platforms.” arXiv preprint arXiv:2210.12402 (2022). (Year: 2022). [cited by examiner]
Chang et al., Using machine learning to extract insights from consumer data. (2022). Encyclopaedia of Data Science and Machine Learning. 1-17. (Year: 2022). [cited by examiner]
Yang et al., Mining Chinese social media UGC: a big-data framework for analyzing Douban movie reviews. Journal of Big Data 3, 3 (2016). https://doi.org/10.1186/s40537-015-0037-9 (Year: 2016). [cited by examiner]
Hanafi et al., “Deep Learning for Recommender System Based on Application Domain Classification Perspective: a Review.” Journal of Theoretical & Applied Information Technology 96, No. 14 (2018). (Year: 2018). [cited by examiner]
Liu et al., “Crystalline: Lowering the Cost for Developers to Collect and Organize Information for Decision Making.” In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pp. 1-16. 2022. (Year… [cited by examiner]
European Patent Office, “International Search Report and Written Opinion”, issued in related International Patent Application No. PCT/IB2024/050897, mail date Apr. 8, 2024. (12 pages). [cited by applicant]