IP Library › Granted Patent US 12,389,076
Granted Patent B2
US 12,389,076 · App. 18/141,059 · Granted Aug 12, 2025

Systems and methods for providing supplemental content related to a queried object

Inventors: Mustafa Coskun (Kayseri, TR); Vehbi Cagri Gungor (Kayseri, TR); Dhananjay Lal (Englewood, CO); Reda Harb (Issaquah, WA)
Assignee: Adeia Guides Inc.
H04N21/4722H04N21/44008H04N21/4532H04N21/8133
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,389,076
App. No.
18/141,059
Granted
Aug 12, 2025
Kind
B2
Abstract

Systems and methods are described for generating for display a media asset, and receiving a query regarding an object depicted in the media asset at a first time point within a presentation duration of the media asset. The system and methods may, based on receiving the query, determine one or more second presentation points within the presentation duration of the media asset related to the object, identify the one or more second presentation points as supplemental content, and generate for display the supplemental content while the media asset is being generated for display.

Claims (75)

1. A computer-implemented system, comprising:

control circuitry configured to:

generate for display a media asset; and

input/output (I/O) circuitry configured to:

receive a query regarding an object depicted in the media asset at a first time point within a presentation duration of the media asset;

wherein the control circuitry is further configured to:

identify supplemental content related to the object by:

determining a presentation point within the presentation duration of the media asset related to the object;

determining, based on a user profile of a user associated with the query, whether the presentation point was previously consumed by the user of the user profile; and

based on determining that the presentation point was not previously consumed by the user of the user profile, refraining from using the presentation point as the supplemental content; and

identifying other content to be used as the supplemental content; and

generate for display the other content as the supplemental content while the media asset is being generated for display.

2. The system of claim 1 , wherein the control circuitry is configured to determine an identity of the object in a context of the media asset by:

identifying a plurality of portions of the media asset that are related to the object depicted at the first time point of the media asset and associated with the query; and

using one or more attributes of the plurality of portions of the media asset to determine the identity of the object in the context of the media asset.

3. The system of claim 2 , wherein the control circuitry is configured to determine the identity of the object in the context of the media asset by:

determining a type of the object depicted at the first time point of the media asset and associated with the query, wherein the plurality of portions of the media asset that are related to the object are identified based on depicting one or more objects of the same type as the object;

comparing the object associated with the query to the one or more objects depicted in the plurality of portions of the media asset;

determining, based on the comparing, one or more matching objects in the plurality of portions that match the object depicted at the first time point of the media asset and associated with the query; and

using the one or more matching objects to determine the identity of the object in the context of the media asset.

4. The system of claim 2 , wherein the control circuitry is further configured to:

train a machine learning model to receive as input an attribute related to a particular object depicted in the media asset and output an indication of an identity of the particular object in the context of the media asset;

input, to the trained machine learning model, a particular attribute related to the object and one or more attributes related to the plurality of portions of the media asset, wherein the one or more attributes are different than the particular attribute of the object; and

determine an output of the trained machine learning model indicating the identity of the object in the context of the media asset.

5. The system of claim 2 , wherein the control circuitry is further configured to:

generate a knowledge graph comprising a plurality of nodes, the plurality of nodes comprising a first node corresponding to a particular attribute related to the object and one or more other nodes corresponding to one or more attributes related to the plurality of portions of the media asset; and

use the knowledge graph to determine the identity of the object in the context of the media asset.

6. The system of claim 1 , wherein:

the media asset is an episodic media asset comprising a plurality of episodes of a series;

the first time point occurs during a first episode of the plurality of episodes; and

the presentation point occurs during a second episode of the plurality of episodes that is earlier in the series than the first episode or later in the series than the first episode.

7. The system of claim 1 , wherein:

the media asset comprises a plurality of related media assets;

the first time point occurs during a first related media asset of the plurality of related media assets; and

the presentation point occurs during a second related media asset corresponding to a prequel of, or a sequel to, the first related media asset.

8. The system of claim 1 , wherein the other content and the media asset are displayed simultaneously on a same device.

9. The system of claim 1 , wherein the other content and the media asset are displayed simultaneously on different devices.

10. The system of claim 1 , wherein the control circuitry is configured to identify the other content to be used as the supplemental content by:

obtaining the other content from an external source that is distinct from within the presentation duration of the media asset, based on determining that the presentation point was not previously consumed by the user of the user profile.

11. A computer-implemented method, comprising:

generating for display a media asset;

receiving a query regarding an object depicted in the media asset at a first time point within a presentation duration of the media asset;

identifying supplemental content related to the object by:

determining a presentation point within the presentation duration of the media asset related to the object;

determining, based on a user profile of a user associated with the query, whether the presentation point was previously consumed by the user of the user profile; and

based on determining that the presentation point was not previously consumed by the user of the user profile, refraining from using the presentation point as the supplemental content; and

identifying other content to be used as the supplemental content; and

generating for display the other content as the supplemental content while the media asset is being generated for display.

12. The method of claim 11 , further comprising determining an identity of the object in a context of the media asset by:

identifying a plurality of portions of the media asset that are related to the object depicted at the first time point of the media asset and associated with the query; and

using one or more attributes of the plurality of portions of the media asset to determine the identity of the object in the context of the media asset.

13. The method of claim 12 , wherein determining the identity of the object in the context of the media asset further comprises:

determining a type of the object depicted at the first time point of the media asset and associated with the query, wherein the plurality of portions of the media asset that are related to the object are identified based on depicting one or more objects of the same type as the object;

comparing the object associated with the query to the one or more objects depicted in the plurality of portions of the media asset;

determining, based on the comparing, one or more matching objects in the plurality of portions that match the object depicted at the first time point of the media asset and associated with the query; and

using the one or more matching objects to determine the identity of the object in the context of the media asset.

14. The method of claim 12 , further comprising:

training a machine learning model to receive as input an attribute related to a particular object depicted in the media asset and output an indication of an identity of the particular object in the context of the media asset;

inputting, to the trained machine learning model, a particular attribute related to the object and one or more attributes related to the plurality of portions of the media asset, wherein the one or more attributes are different than the particular attribute of the object; and

determining an output of the trained machine learning model indicating the identity of the object in the context of the media asset.

15. The method of claim 12 , further comprising:

generating a knowledge graph comprising a plurality of nodes, the plurality of nodes comprising a first node corresponding to a particular attribute related to the object and one or more other nodes corresponding to one or more attributes related to the plurality of portions of the media asset; and

using the knowledge graph to determine the identity of the object in the context of the media asset.

16. The method of claim 11 , wherein:

the media asset is an episodic media asset comprising a plurality of episodes of a series;

the first time point occurs during a first episode of the plurality of episodes; and

the presentation point occurs during a second episode of the plurality of episodes that is earlier in the series than the first episode or later in the series than the first episode.

17. The method of claim 11 , wherein:

the media asset comprises a plurality of related media assets;

the first time point occurs during a first related media asset of the plurality of related media assets; and

the presentation point occurs during a second related media asset corresponding to a prequel of, or a sequel to, the first related media asset.

18. The method of claim 11 , wherein identifying the other content to be used as the supplemental content comprises:

obtaining the other content from an external source that is distinct from within the presentation duration of the media asset, based on determining that the presentation point was not previously consumed by the user of the user profile.

19. The method of claim 11 , wherein the other content and the media asset are displayed simultaneously on a same device.

20. The method of claim 11 , wherein the other content and the media asset are displayed simultaneously on different devices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2023
From: COSKUN, MUSTAFA; GUNGOR, VEHBI CAGRI; LAL, DHANANJAY; HARB, REDA
To: ADEIA GUIDES INC.
Reel/Frame 064222/0666 →
Continuity (1)
Related Publication 20240364970A1 · Oct 31, 2024
References Cited (32)
US 10628501B2 · Santiago · 2020 [cited by applicant]
US 10681432B2 · Vehovsky et al. · 2020 [cited by applicant]
US 10869094B2 · Stathacopoulos · 2020 [cited by applicant]
US 11010436B1 · Peng et al. · 2021 [cited by applicant]
US 11170817B2 · Pham et al. · 2021 [cited by applicant]
US 20050137958A1 · Huber · 2005 [cited by examiner]
US 20140331264A1 · Schneiderman et al. · 2014 [cited by applicant]
US 20160381434A1 · Pulido · 2016 [cited by examiner]
US 20200128294A1 · Gupta et al. · 2020 [cited by applicant]
US 20210321166A1 · Jeong · 2021 [cited by examiner]
US 20210360331A1 · Craner · 2021 [cited by applicant]
“LucidVideo—Create & Share Short Videos—Fast & Free”, retrieved at https://web.archive.org/web/20220616125437/https:/lucidvideo.ai/, on Jul. 11, 2023. [cited by applicant]
Almog, Uri, “Object Detection With Deep Learning: RCNN, Anchors, Non-Maximum-Suppression”, https://medium.com/swlh/object-detection-with-deep-learning-renn-anchors-non-maximum-suppression-ce5a83c7c62b, Oct. 3, 2020. [cited by applicant]
Anwar, Taha, “Introduction to Video Classification and Human Activity Recognition”, https://learnopencv.com/introduction-to-video-classification-and-human-activity-recognition/, Mar. 8, 2021. [cited by applicant]
Bastani, Favyen, et al., “MIRIS: Fast Object Track Queries in Video”, In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, 2020, 1907-1921. [cited by applicant]
Coimbra De Andrade, Douglas, “Recognizing Speech Commands Using Recurrent Neural Networks with Attention”, https://towardsdatascience.com/recognizing-speech-commands-using-recurrent-neural-networks-with-attention-c2b2ba… [cited by applicant]
Golub, Gene H., et al., “Tikhonov regularization and total least squares”, Siam J. Matrix Anal. Appl., vol. 21, No. 1, 1999, 185-194. [cited by applicant]
Gunjal, Satish, “Multivariate Linear Regression From Scratch With Python”, retrieved at https://web.archive.org/web/20210616210641/https:/satishgunjal.com/multivariate_lr/#page-title, on Jul. 11, 2023. [cited by applicant]
He, Kaiming, et al., “Mask R-CNN”, In Proceedings of the IEEE International Conference on Computer Vision, 2017, 2961-2969. [cited by applicant]
Hossain, Md Zakir, et al., “A comprehensive survey of deep learning for image captioning”, ACM Computing Surveys (CsUR) 51, No. 6, 2018, 1-36. [cited by applicant]
Karpathy, Andrej, et al., “Deep visual-semantic alignments for generating image descriptions”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, 3128-3137. [cited by applicant]
Khattak, Asad, et al., “An efficient deep learning technique for facial emotion recognition”, Multimedia Tools and Applications 81, No. 2, 2022, 1649-1683. [cited by applicant]
Liu, Wei, et al., “SSD: Single Shot MultiBox Detector”, In European Conference on Computer Vision (ECCV) 2016, Springer, Part I, LNCS 9905, 2016, 21-37. [cited by applicant]
Lokoč, Jakub, et al., “A Framework for Effective Known-item Search in Video”, In Proceedings of the 27th ACM International Conference on Multimedia, 2019, 1777-1785. [cited by applicant]
Rao, Anyi, et al., “A Local-to-Global Approach to Multi-modal Movie Scene Segmentation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 10146-10155. [cited by applicant]
Redmon, Joseph, et al., “YOLOv3: An Incremental Improvement”, Technical Report, University of Washington, 2018, 1-6. [cited by applicant]
Rendle, Steffen, et al., “BPR: Bayesian Personalized Ranking from Implicit Feedback”, UAI 2009, https://arxiv.org/ftp/arxiv/papers/1205/1205.2618.pdf, 2009, 452-461. [cited by applicant]
Rui, Yong, et al., “Exploring Video Structure Beyond the Shots”, In Proceedings. IEEE International Conference on Multimedia Computing and Systems (Cat. No. 98TB100241), 1998, 237-240. [cited by applicant]
Shetty, Badreesh, “5 Classification Algorithms for Machine Learning”, https://builtin.com/data-science/supervised-machine-learning-classification, Apr. 23, 2023. [cited by applicant]
Theckedath, Dhananjay, et al., “Detecting Afect States Using VGG16, ResNet50 and SE-ResNet50 Networks”, SN Computer Science 1:79, 2020, 1-7. [cited by applicant]
Wang, Limin, et al., “Temporal Segment Networks: Towards Good Practices for Deep Action Recognition”, European Conference on Computer Vision (ECCV), 2016, 1-16. [cited by applicant]
Zhang, Muhan, et al., “Link Prediction Based on Graph Neural Networks”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), 2018, 1-11. [cited by applicant]