IP Library › Granted Patent US 12,651,457
Granted Patent B1
US 12,651,457 · App. 17/852,124 · Granted Jun 9, 2026

Automated content recognition machine learning model generation

Inventors: Mohamed Kamal Omar (Seattle, WA); Ashutosh Sanan (Seattle, WA); Xiaohang Sun (Bellevue, WA); Han-Kai Hsu (Seattle, WA); Wentao Zhu (Redmond, WA); Xiang Hao (Redmond, WA); Ahmed Aly Saad Ahmed (Bothell, WA)
Assignee: Amazon Technologies, Inc.
G06V20/41G06F16/738G06V10/774G06V10/776G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,457
App. No.
17/852,124
Granted
Jun 9, 2026
Kind
B1
Abstract

Systems and techniques for training a machine learning model to identify content labels within a video catalog using multimodal inputs are described. The techniques include receiving a multimodal input including the content. The techniques include determining a first selection of video data in data clusters based on the input and determining metadata indicating a correlation between the first selection and the attribute. Subsequently a second selection is selected based on the metadata and the first selection that may be used as a training dataset. A machine learning model is trained using the training data to determine instances of the attribute and build a content repository that summarizes the video data using the attribute labels.

Claims (78)

1 . A method, comprising:

receiving, from a user computing device comprising one or more processors and memory, user input data indicative of a new attribute of video content for adding labels to video data stored in a database of video data, wherein the video data stored in the database is not currently labelled for the new attribute, wherein the input data is further indicative of a type of machine learning model to be trained to label video data having the new attribute;

determining embedding values for a machine learning model of the type of machine learning model based on the user input data indicative of the new attribute;

determining a first subset of data having a similarity to the new attribute from a database of video data based on the user input data and the embedding values, wherein the video data of the database is labelled with previously determined labels according to one or more attributes different from the new attribute and is not currently labelled for the new attribute;

receiving indications of matching for the first subset of data, the indications of matching associated with whether an instance of the first subset of data is associated with the new attribute;

determining a training dataset, from the first subset of data, based on the indications of matching;

generating a trained machine learning model to label video data having the new attribute by training the machine learning model using the training dataset;

determining labels for a second subset of data using the trained machine learning model;

determining a precision of the trained machine learning model based on the labels for the second subset of data and label verification data that indicates whether the labels are accurately applied to the second subset of data;

performing at least one of:

further training the trained machine learning model using the label verification data in response to the precision being below a threshold; or

generating an alert that the trained machine learning model is trained in response to the precision being equal to or above the threshold;

using the trained machine learning model to identify video data in the database having the new attribute; and

storing indications of labels for the new attribute associated with the identified video data within the database such that labelling of the video data of the database is updated according to the new attribute and continues to provide the previously determined labels for the one or more attributes different from the new attribute.

2 . The method of claim 1 , further comprising:

searching a second database of existing machine learning models for the type of machine learning model, and wherein determining the embedding values comprises one of:

generating the embedding values in response to the database of existing machine learning models not including the type of machine learning model; or

selecting the embedding values of an existing machine learning model in response to the database including the type of machine learning model.

3 . The method of claim 1 , wherein:

the machine learning model comprises a convolutional neural network and a linear classifier, with the linear classifier configured to receive an output of the convolutional neural network;

the embedding values comprise the embedding values of the convolutional neural network; and

training the machine learning model comprises training the linear classifier using the training dataset.

4 . The method of claim 1 , wherein:

the labels comprise attributes of the video data; and

the user input data comprises two or more of text data, audio data, or video data.

5 . A method, comprising:

receiving a user query associated with a new attribute of video content for adding labels to video data having the new attribute stored in a database different from labels for any attributes currently associated with the video data stored in the database;

determining a first subset of the video data having a similarity to the new attribute based at least in part on the user query;

determining metadata indicating a correlation between the first subset of data and the new attribute;

determining a second subset of the video data based at least in part on the first subset of data and the metadata;

determining a training dataset based at least in part on the second subset of data;

training a machine learning model using the training dataset and the metadata;

determining, using the trained machine learning model, one or more instances of the new attribute within the video data stored in the database; and

storing indications of labels for the new attribute associated with the database identifying one or more locations of the determined one or more instances of the new attribute within the video data stored in the database.

6 . The method of claim 5 , wherein the new attribute is a first attribute and further comprising:

receiving a second user query for a second attribute;

determining a third subset of the video data based at least in part on the second user query;

determining second metadata indicating a second correlation between the third subset of data and the second attribute;

determining a second training dataset for the user query based at least in part on the third subset of data and the second metadata; and

training the machine learning model to determine instances of the second attribute in addition to the first attribute.

7 . The method of claim 5 , wherein determining the first subset of data comprises searching indexed data of the database of video data for keywords associated with the user query.

8 . The method of claim 5 , wherein determining the training dataset comprises iteratively:

receiving second metadata for a subset of data indicating a second correlation between the subset of data and the new attribute; and

determining a subsequent subset of data from the video data based at least in part on the subset of data and the second metadata.

9 . The method of claim 5 , further comprising determining a type of machine learning model by searching a database of existing machine learning models based at least in part on the user query or the new attribute.

10 . The method of claim 5 , further comprising determining an expected accuracy and precision of the machine learning model based at least in part on the training dataset prior to training the machine learning model.

11 . The method of claim 5 , further comprising:

determining one or more additional instances for one or more additional attributes within the video data using the machine learning model; and

determining a content summary for individual videos of the video data based at least in part on the one or more additional instances of the one or more additional attributes.

12 . The method of claim 11 , wherein the machine learning model comprises a feature extraction backbone and a linear classifier, and wherein training the machine learning model comprises training the linear classifier to identify instances of the one or more additional attributes.

13 . The method of claim 5 , wherein determining the metadata comprises presenting the first subset of data on a user interface and receiving indications with respect to positive instances and negative instances of the new attribute within the first subset of data.

14 . The method of claim 5 , wherein the user query comprises at least one of:

image data;

audio data; or

text data.

15 . A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving a user query associated with a new attribute of video content for adding labels to video data having the new attribute stored in a database that is different from any attributes currently associated with the video data stored in the database;

determining a first subset of the data having a similarity to the new attribute based at least in part on the user query;

determining metadata indicating a correlation between the first subset of data and the new attribute;

determining a second subset of the video data based at least in part on the first subset of data and the metadata;

determining a training dataset based at least in part on the second subset of data;

training a machine learning model using the training dataset and the metadata;

determining, using the trained machine learning model, one or more instances of the new attribute within the video data stored in the database of video data; and

storing indications of labels for the new attribute associated with the database identifying one or more locations of the determined one or more instances of the new attribute within the video data stored in the database.

16 . The non-transitory computer-readable medium of claim 15 , wherein determining the training dataset comprises iteratively:

receiving second metadata for a subset of data indicating a second correlation between the subset of data and the new attribute; and

determining a subsequent subset of data from the video data based at least in part on the subset of data and the second metadata.

17 . The non-transitory computer-readable medium of claim 15 , further comprising:

determining one or more additional instances for one or more additional attributes within the video data using the machine learning model; and

determining a content summary for individual videos of the video data based at least in part on the one or more additional instances of the one or more additional attributes.

18 . The non-transitory computer-readable medium of claim 17 , wherein the machine learning model comprises a feature extraction backbone and a linear classifier, and wherein training the machine learning model comprises training the linear classifier to identify instances of the one or more additional attributes.

19 . The non-transitory computer-readable medium of claim 15 , wherein determining the metadata comprises presenting the first subset of data on a user interface and receiving indications with respect to positive instances and negative instances of the new attribute within the first subset of data.

20 . The non-transitory computer-readable medium of claim 15 , wherein the new attribute is a first attribute and further comprising:

receiving a second user query for a second attribute;

determining a third subset of the video data based at least in part on the second user query;

determining second metadata indicating a second correlation between the third subset of data and the second attribute;

determining a second training dataset for the user query based at least in part on the third subset of data and the second metadata; and

training the machine learning model to determine instances of the second attribute in addition to the first attribute.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2022
From: OMAR, MOHAMED KAMAL; SANAN, ASHUTOSH; SUN, XIAOHANG; HSU, HAN-KAI; ZHU, WENTAO; HAO, XIANG; AHMED, AHMED ALY SAAD
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 060342/0041 →
References Cited (21)
US 9875445B2 · Amer · 2018 [cited by examiner]
US 11093798B2 · Torres · 2021 [cited by examiner]
US 11816618B1 · Cheek, Jr. · 2023 [cited by examiner]
US 20120243789A1 · Yang · 2012 [cited by examiner]
US 20190215551A1 · Modarresi · 2019 [cited by examiner]
US 20190362222A1 · Chen · 2019 [cited by examiner]
US 20200380298A1 · Aggarwal · 2020 [cited by examiner]
US 20210201209A1 · Sghiouer · 2021 [cited by examiner]
US 20230172167A1 · Eftelioglu · 2023 [cited by examiner]
Marques, O., Furht, B. Muse: A Content-Based Image Search and Retrieval System Using Relevance Feedback. Multimedia Tools and Applications 17, 21-50 (2002). (Year: 2002). [cited by examiner]
I. A. Kabary, I. Giangreco, H. Schuldt, F. Matulic and M. Norrie, “QUEST: Towards a Multi-modal CBIR Framework Combining Query-by-Example, Query-by-Sketch, and Text Search,” 2013 IEEE International Symposium on Multimed… [cited by examiner]
Tejaswi Nayak, U., Sujatha, C., Kamat, T.V., Desai, P. (2021). Video Retrieval Using Residual Networks. In: Thampi, S.M., Gelenbe , E., Atiquzzaman, M., Chaudhary, V., Li, KC. (eds) Advances in Computing and Network Com… [cited by examiner]
Phan, T.-C.; Phan, A.-C.; Cao, H.-P.; Trieu, T.-N. Content-Based Video Big Data Retrieval with Extensive Features and Deep Learning. Appl. Sci. 2022, 12, 6753. (Year: 2022). [cited by examiner]
Rao, Y., Liu, W., Fan, B. et al. A novel relevance feedback method for CBIR. World Wide Web 21, 1505-1522 (2018). (Year: 2018). [cited by examiner]
Ansari, Mohammed & Mohammed, Muzammil. (2015). Content based Video Retrieval Systems—Methods, Techniques, Trends and Challenges. International Journal of Computer Applications. 112. 975-8887 (Year: 2015). [cited by examiner]
Al-Mohamade, Abeer et al. “Multiple Query Content-Based Image Retrieval Using Relevance Feature Weight Learning.” Journal of imaging vol. 6,1 2. Jan. 17, 2020) (Year: 2020). [cited by examiner]
Marques, O., Furht, B. MUSE: A Content-Based Image Search and Retrieval System Using Relevance Feedback. Multimedia Tools and Applications 17, 21-50 (Year: 2002). [cited by examiner]
Stefan Siersdorfer, Jose San Pedro, and Mark Sanderson. 2009. Automatic video tagging using content redundancy. In Proceedings of the 32nd international ACM SIGIR conference on Research and development in information re… [cited by examiner]
Sayyed, Sufiyan & Patkar, Sayali & Patil, Atharva & Patil, Mahendra. (2020). Analyzing Video to Find Particular Object Timestamp. (Year: 2020). [cited by examiner]
Jie Hwan Lee et al. “Audio Query-Based Music Source Separation” Center for Super Intelligence, Seoul National University, Korea (Year: 2019). [cited by examiner]
Nina Shvetsova et al., Everything at Once—Multi-modal Fusion Transformer for Video Retrieval, Dec. 9, 2021, (Year: 2021). [cited by examiner]