IP Library Granted Patent US 12,205,361
Granted Patent B2
US 12,205,361 · App. 17/885,655 · Granted Jan 21, 2025

System and method for facilitating graphic-recognition training of a recognition model

Inventors: David Joshua Eigen (New York, NY); Matthew Zeiler (Fort Lee, NJ)
Assignee: Clarifai, Inc
G06V10/82G06F18/217G06F18/2414G06F18/28G06T5/00G06T7/337G06T11/60G06V30/1914G06V30/1916G06V30/19173G06V30/194G06T2207/20081G06T2207/20084G06V2201/09
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,361
App. No.
17/885,655
Granted
Jan 21, 2025
Kind
B2
Abstract

Methods and computer readable media for facilitating training of a recognition model. An embodiment includes generating media items based on information associated with a representation of a graphic, the information including content other than the graphic, content based on at least one transformation parameter set, and content comprising the graphic integrated with the other content, then using a recognition model to process the media items to generate predictions related to recognition of the graphic for the media items, the generated predictions including an indication of a predicted location of the graphic in a first media item. The process also includes presenting an indication of the predicted location on an area of the first media item via a user interface to a user, then obtaining a reference feedback set that includes reference indications related to recognition of the graphic for the media items and including user feedback concerning the indication of the predicted location of the graphic, and then updating the recognition model based on the reference feedback.

Claims (46)

1. A method for facilitating training of a recognition model comprising:

generating media items based on information associated with a representation of a graphic, the information comprising content other than the graphic, content based on at least one transformation parameter set, and content comprising the graphic integrated with other content;

generating, by using a recognition model to process the media items, predictions related to recognition of the graphic for the media items, the generated predictions comprising an indication of a predicted location of the graphic in a first media item;

presenting, via a user interface to a user, an indication of the predicted location on an area of the first media item;

obtaining a reference feedback set, the reference feedback set comprising reference indications related to recognition of the graphic for the media items and including user feedback via the user interface concerning the indication of the predicted location of the graphic on an area of the first media item; and

updating the recognition model based on the reference feedback set including the user feedback concerning the indication of the predicted location of the graphic.

2. The method of claim 1 , wherein generating the media items comprises using a plurality of transformation parameter sets, wherein each of the transformation parameter sets comprise different parameters from each other,

and wherein generating the media items comprises generating at least some media items by applying a different one of the transformation parameter sets to the representation of the graphic such that each media item of the at least some media items has a different transformed representation of the graphic from one another.

3. The method of claim 2 , wherein obtaining the transformation parameter sets comprises randomly generating at least some of the transformation parameter sets.

4. The method of claim 2 , wherein the transformation parameter sets comprise occlusion parameters, and

wherein generating the at least some media items comprises generating one or more of the at least some media items based on the occlusion parameters such that each media item has a transformed representation of the graphic in which at least of a portion of the representation of the graphic is at least one of missing or hidden.

5. The method of claim 2 , wherein the transformation parameter sets comprise at least one of blurring effect parameters, camera effect parameters, motion effect parameters, shadow effect parameters, pattern effect parameters, or texture effect parameters, sharpening parameters, softening parameters, brightness parameters, contrast parameters, or recoloring parameters.

6. The method of claim 2 , wherein the transformation parameter sets comprise compression parameters.

7. The method of claim 1 , wherein the recognition model comprises a neural network, and wherein each of the media items is an image.

8. The method of claim 1 , wherein generating the media items further comprises generating the first media item such that at least some content other than the graphic appears opaquely over at least a portion of the graphic on the first media item, and

wherein generating the predictions comprises generating the indication of the predicted location based on the first media item in which the at least some content opaquely appears over at least a portion of the graphic.

9. The method of claim 1 , wherein obtaining the reference indications comprises:

obtaining a reference indication of an object to be recognized via the user interface by the user of the indication of the predicted location, and

wherein updating the recognition model comprises updating the recognition model based on the reference indication.

10. The method of claim 1 , further comprising:

generating representations of the graphic such that each of the representations of the graphic has a different size from one another,

wherein generating the media items comprises generating at least some media items such that each media item of the at least some media items comprises a different representation of the graphic.

11. The method of claim 1 , wherein the recognition model:

determines similarities or differences between the generated predictions and their corresponding reference indications, and

updates neural network aspects of the recognition model based on the determined similarities or differences.

12. A non-transitory, computer-readable media storing instructions for facilitating training of a recognition model that, when executed by a one or more processors, cause operations comprising:

generating media items based on information associated with a representation of a graphic, the information comprising content other than the graphic, content based on at least one transformation parameter set, and content comprising the graphic integrated with other content;

generating, by using a recognition model to process the media items, predictions related to recognition of the graphic for the media items, the generated predictions comprising an indication of a predicted location of the graphic in a first media item;

presenting, via a user interface to a user, an indication of the predicted location on an area of the first media item;

obtaining a reference feedback set, the reference feedback set comprising reference indications related to recognition of the graphic for the media items and including user feedback via the user interface concerning the indication of the predicted location of the graphic on an area of the first media item; and

updating the recognition model based on the reference feedback set including the user feedback concerning the indication of the predicted location of the graphic.

13. The non-transitory, computer-readable media of claim 12 , wherein generating the media items comprises using a plurality of transformation parameter sets, wherein each of the transformation parameter sets comprise different parameters from each other,

and wherein generating the media items comprises generating at least some media items by applying a different one of the transformation parameter sets to the representation of the graphic such that each media item of the at least some media items has a different transformed representation of the graphic from one another.

14. The non-transitory, computer-readable media of claim 13 , wherein obtaining the transformation parameter sets comprises randomly generating at least some of the transformation parameter sets.

15. The non-transitory, computer-readable media of claim 13 , wherein the transformation parameter sets comprise occlusion parameters, and

wherein generating the at least some media items comprises generating one or more of the at least some media items based on the occlusion parameters such that each media item has a transformed representation of the graphic in which at least of a portion of the representation of the graphic is at least one of missing or hidden.

16. The non-transitory, computer-readable media of claim 13 , wherein the transformation parameter sets comprise at least one of blurring effect parameters, camera effect parameters, motion effect parameters, shadow effect parameters, pattern effect parameters, or texture effect parameters, sharpening parameters, softening parameters, brightness parameters, contrast parameters, or recoloring parameters.

17. The non-transitory, computer-readable media of claim 13 , wherein the transformation parameter sets comprise compression parameters.

18. The non-transitory, computer-readable media of claim 12 , wherein generating the media items further comprises generating the first media item such that at least some content other than the graphic appears opaquely over at least a portion of the graphic on the first media item, and

wherein generating the predictions comprises generating the indication of the predicted location based on the first media item in which the at least some content opaquely appears over at least a portion of the graphic.

19. The non-transitory, computer-readable media of claim 12 , wherein obtaining the reference indications comprises:

obtaining a reference indication of an object to be recognized via the user interface by the user of the indication of the predicted location, and

wherein updating the recognition model comprises updating the recognition model based on the reference indication.

20. The non-transitory, computer-readable media of claim 12 , further comprising:

generating representations of the graphic such that each of the representations of the graphic has a different size from one another,

wherein generating the media items comprises generating at least some media items such that each media item of the at least some media items comprises a different representation of the graphic.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2026
From: CLARIFAI, INC.
To: NEBIUS BV
Reel/Frame 075712/0109 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2026
From: EIGEN, DAVID JOSHUA; ZEILER, MATTHEW
To: CLARIFAI, INC.
Reel/Frame 074805/0665 →
Continuity (4)
Continuation 16998384 · Aug 20, 2020
Continuation 16214636 · Dec 10, 2018
Continuation 15475900 · Mar 31, 2017
Related Publication 20220383649A1 · Dec 1, 2022
References Cited (24)
US 10007863B1 · Pereira · 2018 [cited by applicant]
US 10163043B2 · Eigen · 2018 [cited by examiner]
US 10776675B2 · Eigen · 2020 [cited by examiner]
US 11417130B2 · Eigen · 2022 [cited by examiner]
US 20090297007A1 · Cosatto · 2009 [cited by applicant]
US 20160358024A1 · Krishnakumar et al. · 2016 [cited by applicant]
US 20170116498A1 · Raveane et al. · 2017 [cited by applicant]
US 20170140236A1 · Price · 2017 [cited by applicant]
US 20170304732A1 · Velic et al. · 2017 [cited by applicant]
US 20180082106A1 · Inaba · 2018 [cited by applicant]
US 20190156202A1 · Falk · 2019 [cited by applicant]
US 20210224990A1 · Varekamp · 2021 [cited by applicant]
US 20230237246A1 · Welinder · 2023 [cited by examiner]
IDS dated Aug. 20, 2022 which was filed in connection with U.S. Appl. No. 16/998,384. [cited by applicant]
892 Form dated Apr. 13, 2022 which was issued in connection with U.S. Appl. No. 16/998,384. [cited by applicant]
892 Form dated Jul. 5, 2022 which was issued in connection with U.S. Appl. No. 16/998,384. [cited by applicant]
Notice of Allowance issued on Jul. 5, 2022, in related U.S. Appl. No. 16/998,384, 9 pages. [cited by applicant]
Huang et al., “Vehicle Logo Recognition System Based on Convolutional Neural Networks with a Pretraining Strategy_” IEEE Transactions of Intelligent Transportation Systems, vol. 16, No. 4, Aug. 2015, pp. 1951-1960. [cited by applicant]
IDS dated Apr. 9, 2019 which was filed in connection with U.S. Appl. No. 16/214,636. [cited by applicant]
892 Form dated Jan. 27, 2020 which was issued in connection with U.S. Appl. No. 16/214,636. [cited by applicant]
Notice of Allowance issued on May 14, 2020, in related U.S. Appl. No. 16/214,636, 9 pages. [cited by applicant]
892 Form dated May 10, 2018, in related U.S. Appl. No. 15/475,900, 1 pages. [cited by applicant]
892 Form dated Sep. 4, 2018, in related U.S. Appl. No. 15/475,900, 1 pages. [cited by applicant]
Notice of Allowance issued on Sep. 4, 2018, in related U.S. Appl. No. 15/475,900, 8 pages. [cited by applicant]