IP Library Granted Patent US 12,646,318
Granted Patent B2
US 12,646,318 · App. 18/173,779 · Granted Jun 2, 2026

System and method for fast adaptive brands logos detection on video with open set approach

Inventors: Andrei Boiarov (Sofia, BG); Ilya Shimchik (Istanbul, TR); Nikita Firsakov (Istanbul, TR); Pavlo Bredikhin (Kharkov, UA); Sergey Ulasen (Singapore, SG); Serg Bell (Singapore, SG); Stanislav Protasov (Singapore, SG); Nikolay Dobrovolskiy (Sofia, BG)
Assignees: Constructor Technology AG; Constructor Education and Research Genossenschaft
G06V20/41G06V10/774G06V10/945G06V20/48G06V20/49G06V20/70G06V20/44G06V2201/09
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,318
App. No.
18/173,779
Granted
Jun 2, 2026
Kind
B2
Abstract

A system and method for performing brand detection in a video is disclosed herein. The method comprises receiving the video for performing the brand detection thereon; splitting the video for obtaining a plurality of video frames; performing an open set detection on each input video frame from the plurality of video frames, which comprises proposing one or more bounding boxes on the input video frames on regions of the video frame that potentially include brand media; cropping the one or more bounding boxes; providing the cropped bounding boxes to a classification module for obtaining embedding vectors corresponding to each of the cropped bounding boxes; and comparing the embedding vectors of the cropped bounding boxes with embedding vectors of one or more brand reference images provided by a user for computing instances of brand detection in each video frame of the plurality of video frames.

Claims (33)

1 . A method for performing brand detection in a video, the method comprising:

receiving the video with a video splitter;

splitting the video with the video splitter to obtain a plurality of video frames;

providing the plurality of video frames to a brand detector for performing open set detection on each input video frame from the plurality of video frames, wherein the open set detection comprises:

proposing, by a localization module one or more bounding boxes on the input video frames on regions of a video frame that potentially include brand media,

cropping by a cropping module the one or more bounding boxes from the input video frames to obtain cropped bounding boxes,

providing the cropped bounding boxes to a classification module to obtain embedding vectors corresponding to each of the cropped bounding boxes, and

comparing, with a comparator module, the embedding vectors of the cropped bounding boxes with embedding vectors of one or more brand reference images provided by a user for computing instances of brand detection in each video frame of the plurality of video frames;

segmenting the brand media within the cropped bounding boxes by determining an exact region in which a brand logo is occupied within the cropped bounding boxes by a semantic segmentation model, wherein the exact region is only an area of the brand logo within the cropped bounding boxes; and

computing, by a brand appearance computing unit, one or more parameters associated with per-brand, per-appearance statistics of the brand logo in the video using the comparing.

2 . The method of claim 1 , wherein the brand media and the one or more brand reference images include brand logos, brand taglines, and brand ambassador images.

3 . The method of claim 1 , further comprising training the classification module in an open set approach using self-supervised learning (Supervised Contrastive learning) and few-shot learning.

4 . The method of claim 1 , further comprising resolving a scene understanding task by the semantic segmentation model.

5 . The method of claim 1 , further comprising detecting whether the one or more brand reference images appear in the video at a crucial moment by a video action recognition module.

6 . The method of claim 5 , further comprising identifying whether the one or more brand reference images appear in an area of a screen where a user's attention is focused by the video action recognition module.

7 . The method of claim 1 , further comprising allowing the user to label new brand reference images in the video frames for retraining the classification module.

8 . The method of claim 1 , further comprising allowing the user to provide new brand reference images for retraining the classification module.

9 . A system for performing brand detection in a video, the system comprising:

a video splitter to receive the video for performing the brand detection thereon, the video splitter configured to split the video to obtain a plurality of video frames;

a brand detector for performing an open set detection on each input video frame from the plurality of video frames, wherein the brand detector comprises:

a localization module to propose one or more bounding boxes on the input video frames on regions of a video frame that potentially include a brand media,

a cropping module to crop the one or more bounding boxes from the input video frames to obtain cropped bounding boxes,

a classification module to receive the cropped bounding boxes and obtaining embedding vectors corresponding to each of the cropped bounding boxes, and

a comparator module to compare the embedding vectors of the cropped bounding boxes with embedding vectors of one or more brand reference images provided by a user for computing instances of brand detection in each video frame of the plurality of video frames;

a semantic segmentation model to determine an exact area occupied by a brand logo within the cropped bounding boxes, wherein the exact area is only an area of the brand logo within the cropped bounding boxes; and

a brand appearance computing unit to compute one or more parameters associated with per-brand, per-appearance statistics of the brand logo in the video based on the comparing by the comparator module.

10 . The system of claim 9 , wherein the brand media and the one or more brand reference images include brand logos, brand taglines, and brand ambassador images.

11 . The system of claim 9 , wherein the classification module is trained in an open set approach using self-supervised learning (Supervised Contrastive learning) and few-shot learning.

12 . The system of claim 11 , wherein the semantic segmentation model is configured to perform a scene understanding task.

13 . The system of claim 9 , further comprising a video action recognition module to detect whether the one or more brand reference images appear in the video at a crucial moment.

14 . The system of claim 13 , wherein the video action recognition module is further configured to identify whether the one or more brand reference images appear in an area of a screen where a user's attention is focused.

15 . The system of claim 9 , further comprising a user interface to allow the user to label new brand reference images in the video frames for retraining the classification module.

16 . The system of claim 15 , wherein the user interface is further configured to allow the user to provide the new brand reference images for retraining the classification module.

Assignments (3)
CHANGE OF NAME Recorded Apr 21, 2026
From: AUTONOMOUS AG
To: CONSTRUCTOR TECHNOLOGY AG
Reel/Frame 075453/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2026
From: FIRSAKOV, NIKITA
To: AUTONOMOUS AG
Reel/Frame 074430/0224 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2026
From: BOIAROV, ANDREI; SHIMCHIK, ILYA; BREDIKHIN, PAVLO; ULASEN, SERGEY; BELL, SERG; PROTASOV, STANISLAV; DOBROVOLSKIY, NIKOLAY
To: CONSTRUCTOR TECHNOLOGY AG
Reel/Frame 074430/0149 →
Continuity (1)
Related Publication 20240290093A1 · Aug 29, 2024
References Cited (20)
US 20160227277A1 · Schlesinger · 2016 [cited by examiner]
US 20200372662A1 · Pereira · 2020 [cited by examiner]
US 20220129911A1 · Gabale · 2022 [cited by examiner]
US 20230353839A1 · Chun · 2023 [cited by examiner]
US 20250029170A1 · Chachek · 2025 [cited by examiner]
CN 110837615A · 2020 [cited by applicant]
IN 201941004531A · 2020 [cited by applicant]
WO WO2021194090A1 · 2021 [cited by applicant]
Fehérvári, I., & Appalaraju, S. (Jan. 2019). Scalable logo recognition using proxies. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV) (pp. 715-725). IEEE. (Year: 2019). [cited by examiner]
Su, J. C., Maji, S., & Hariharan, B. (Aug. 2020). When does self-supervision improve few-shot learning?. In European conference on computer vision (pp. 645-666). Cham: Springer International Publishing. (Year: 2020). [cited by examiner]
Ermakov, M., & Makarov, I. (Nov. 2022). Few-shot logo recognition in the wild. In 2022 IEEE 22nd International Symposium on Computational Intelligence and Informatics and 8th IEEE (CINTI-MACRo) (pp. 000393-000398). IEEE… [cited by examiner]
Tüzkö, A., Herrmann, C., Manger, D., & Beyerer, J. (2017). Open set logo detection and retrieval. arXiv preprint arXiv:1710.10891. (Year: 2017). [cited by examiner]
Li, C., Fehervari, I., Zhao, X., Macedo, I., & Appalaraju, S. (Jan. 2022). SeeTek: Very Large-Scale Open-set Logo Recognition with Text-Aware Metric Learning. In 2022 IEEE/CVF Winter Conference on Applications of Comput… [cited by examiner]
Dong Wu, et al. YOLOP:YouOnlyLookOnceforPanopticDrivingPerception, Mar. 26, 2022, 2108.11250.pdf. [cited by applicant]
Chenge Li, et.al. See Tek: Very Large-Scale Open-set Logo Recognition with Text-Aware Metric Learning, 2022, Li SeeTek.pdf. [cited by applicant]
Olaf Ronneberger, et al. U-Net: Convolutional Networks for Biomedical Image Segmentation, May 18, 2015, 1505.04597.pdf. [cited by applicant]
Christoph Feichtenhofer, et al. SlowFast Networks for Video Recognition, Oct Oct. 29, 2019, 1812.03982.pdf. [cited by applicant]
Haoqi Fan, et al. Multiscale Vision Transformers, 2021, Fan_Multiscale_Vision_Transformers_ICCV_2021_paper.pdf. [cited by applicant]
Jake Snell, et al. Prototypical Networks for Few-shot Learning, Jun. 19, 2017, 1703.05175.pdf. [cited by applicant]
Prannay Khosla, et al. Supervised Contrastive Learning, Mar. 10, 2021, 2004.11362.pdf. [cited by applicant]