IP Library › Granted Patent US 12,621,544
Granted Patent B2
US 12,621,544 · App. 18/883,720 · Granted May 5, 2026

Dynamically altering preprocessing of streaming video data by a large language model dependent upon review by the large language model

Inventor: Amol Ajgaonkar (Chandler, AZ)
Assignee: Insight Direct USA, Inc.
H04N21/84G06V10/25G06V10/761G11B27/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,621,544
App. No.
18/883,720
Granted
May 5, 2026
Kind
B2
Abstract

A method of preprocessing incoming video data can include preprocessing incoming video data, by a computer processor, according to preprocessing parameters, wherein the preprocessing includes formatting the incoming video data to create first video data of a first region of interest. The method can also include providing the first video data to a first transformer model along with a first prompt requesting the first transformer model to describe the first video data, describing the first video data by the first transformer model to generate at least one description of the first video data, providing the at least one description to a second transformer model with a second prompt requesting the second transformer model to review the at least one description and determine altered preprocessing parameters for preprocessing the incoming video data, and generating, by the second transformer model, altered preprocessing parameters based upon the at least one description.

Claims (52)

1 . A method of preprocessing incoming video data having at least one region of interest, the method comprising:

accessing the incoming video;

preprocessing the incoming video data, by a computer processor, according to preprocessing parameters, wherein the preprocessing includes formatting the incoming video data to create first video data of a first region of interest;

providing the first video data to a first transformer model along with a first prompt requesting the first transformer model to describe the first video data;

describing the first video data by the first transformer model to generate at least one description of the first video data;

providing the at least one description to a second transformer model with a second prompt requesting the second transformer model to review the at least one description and determine altered preprocessing parameters for preprocessing the incoming video data;

generating, by the second transformer model, altered preprocessing parameters based upon the at least one description;

providing the altered preprocessing parameters to the computer processor; and

preprocessing the incoming video data according to the altered preprocessing parameters to create second video data.

2 . The method of claim 1 , wherein the altered preprocessing parameters change at least one of the following video edits of the incoming video data as compared to the unaltered preprocessing parameters: crop, grayscale, contrast, brightness, color threshold, resize, blur, hue saturation value, sharpen, erosion, dilation, Laplacian image processing, Sobel image processing, pyramid up, and pyramid down.

3 . The method of claim 1 , further comprising:

formatting, by the second transformer model, the altered preprocessing parameters so as to be accepted and applied by the computer processor.

4 . The method of claim 1 , further comprising:

providing the altered preprocessing parameters to a formatting module; and

formatting the altered preprocessing parameters, by the formatting module, into a format that is acceptable by the computer processor,

wherein the formatting module provides the altered preprocessing parameters in an acceptable format to the computer processor.

5 . The method of claim 1 , wherein the first transformer model and the second transformer model each include a large language model.

6 . The method of claim 1 , wherein the first transformer model is different from the second transformer model.

7 . The method of claim 1 , wherein the first transformer model and the second transformer model are the same transformer model.

8 . The method of claim 1 , further comprising:

publishing the first video data to an endpoint,

wherein accessing the first video data includes subscribing to the endpoint.

9 . The method of claim 8 , wherein the endpoint is hosted by a gateway.

10 . The method of claim 1 , wherein the incoming video data is received from a camera.

11 . A method of preprocessing incoming video data having at least one region of interest, the method comprising:

preprocessing the incoming video data, by a computer processor, according to preprocessing parameters, wherein the preprocessing includes formatting the incoming video data to create first video data of a first region of interest;

accessing the first video data by an AI model;

processing the first video data by the AI model to determine a first output that is indicative of a first inference dependent upon the first video data;

providing the first video data, the first output, and a first prompt to a first transformer model with the first prompt requesting the first transformer model to describe the first video data;

generating, by the first transformer model, at least one description of the first video data from the first video data;

providing the at least one description and a second prompt to a second transformer model with the second prompt requesting the second transformer model to determine altered preprocessing parameters that alter the incoming video data to create second video data; and

determining, by the second transformer model, the altered preprocessing parameters dependent upon the at least one description with the altered preprocessing parameters altering the incoming video data to create the second video data.

12 . The method of claim 11 , further comprising:

changing the preprocessing parameters in a configuration file to be the altered preprocessing parameters;

accessing the altered preprocessing parameters by the computer processor; and

preprocessing the incoming video data according to the altered preprocessing parameters to create the second video data.

13 . The method of claim 12 , further comprising:

providing the second video data and a third prompt to the first transformer model with the third prompt requesting the first transformer model to describe the second video data; and

describing, by the first transformer model, the second video data to create at least one description of the second video data.

14 . The method of claim 13 , further comprising:

compiling the at least one description of the first video data and the at least one description of the second video data into an overall video data description.

15 . The method of claim 11 , further comprising:

providing the altered preprocessing parameters to a formatting module;

formatting the altered preprocessing parameters, by the formatting module, into a format that is acceptable by the computer processor; and

providing the altered preprocessing parameters in an acceptable format to the computer processor.

16 . The method of claim 11 , wherein the first transformer model and the second transformer model each include a large language model.

17 . The method of claim 11 , wherein the first transformer model and the second transformer model are the same transformer model.

18 . The method of claim 11 , further comprising:

generating multiple vector embeddings corresponding to multiple descriptions of the at least one description of the first video data.

19 . The method of claim 18 , wherein the at least one description of the first video data includes multiple descriptions with each description being generated by the first transformer model and corresponding to one frame of multiple frames that form the first video data.

20 . The method of claim 19 , further comprising:

searching the multiple descriptions to find at least one relevant frame of the first video data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2024
From: AJGAONKAR, AMOL
To: INSIGHT DIRECT USA, INC.
Reel/Frame 068573/0461 →
Continuity (2)
Provisional Application 63553204 · Feb 14, 2024
Related Publication 20250259653A1 · Aug 14, 2025
References Cited (23)
US 9635307B1 · Mysore Vijaya Kumar et al. · 2017 [cited by applicant]
US 11910073B1 · Sharma et al. · 2024 [cited by applicant]
US 11954151B1 · Jain et al. · 2024 [cited by applicant]
US 11995412B1 · Mishra · 2024 [cited by applicant]
US 12056918B1 · Goyal et al. · 2024 [cited by applicant]
US 12283291B1 · Sarfati et al. · 2025 [cited by applicant]
US 12301960B1 · Ben-cohen et al. · 2025 [cited by applicant]
US 20110063500A1 · Loher · 2011 [cited by examiner]
US 20200320116A1 · Wu · 2020 [cited by applicant]
US 20210326393A1 · Aggarwal et al. · 2021 [cited by applicant]
US 20220076707A1 · Walker et al. · 2022 [cited by applicant]
US 20220108208A1 · Li et al. · 2022 [cited by applicant]
US 20230019360A1 · Whatmough et al. · 2023 [cited by applicant]
US 20240114146A1 · Kawai · 2024 [cited by examiner]
US 20240177443A1 · Ajgaonkar · 2024 [cited by examiner]
US 20240205520A1 · Carbajo et al. · 2024 [cited by applicant]
US 20240232937A1 · D'Auria · 2024 [cited by examiner]
US 20240395042A1 · Boiarov et al. · 2024 [cited by applicant]
US 20240406521A1 · Ramesh et al. · 2024 [cited by applicant]
US 20250124689A1 · Williams et al. · 2025 [cited by applicant]
US 20250182483A1 · Zhao et al. · 2025 [cited by applicant]
US 20250190503A1 · Lee et al. · 2025 [cited by applicant]
US 20250238968A1 · Ibrahim et al. · 2025 [cited by applicant]