IP Library Granted Patent US 12,633,112
Granted Patent B2
US 12,633,112 · App. 18/750,644 · Granted May 19, 2026

Methods and apparatus for identifying video-derived data

Inventors: Rishabh Goyal (San Mateo, CA); Song Cao (Foster City, CA)
Assignee: Verkada Inc.
G06V10/945G06F16/71G06F16/75G06V10/82G06V20/41G06V20/52H04N5/2628H04N7/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,112
App. No.
18/750,644
Granted
May 19, 2026
Kind
B2
Abstract

A method for generating and selecting images of objects based on video data and text data includes receiving, at a processor of a video camera system, a video stream including a series of video frames depicting at least one object. A set of at least one classification for the object is generated. Additionally, an image that depicts the object and that includes a cropped portion of a video frame from the series of video frames is generated. A set of at least one index key is generated based on the set of at least one classification, and the image is stored based on the set of at least one index key. The processor receives a signal representing a text input from a user, and the processor performs at least one of (1) retrieval of the image or (2) generation of an alert.

Claims (65)

1 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:

receive a plurality of temporally arranged images, each image from the plurality of temporally arranged images including a depiction of an object;

generate a set of at least one identification for the object based on at least one image from the plurality of temporally arranged images;

generate a set of at least one cropped image, each cropped image from the set of at least one cropped image including a region of an image from the plurality of temporally arranged images and from a plurality of regions of that image, the image being different from each remaining image from the plurality of temporally arranged images, the region being smaller than an entirety of the image and depicting the object;

cause storage, in a memory operably coupled to the processor, of at least one cropped image from the set of at least one cropped image based on an index that is (1) configured for use with a search operation and (2) associated with the set of at least one identification;

receive, from a compute device of a user, a signal that encodes at least one identification from the set of at least one identification;

retrieve, using the search operation and the index, the at least one cropped image from the memory based on the signal; and

cause transmission of the at least one cropped image to a remote compute device for display.

2 . The non-transitory, processor-readable medium of claim 1 , wherein the plurality of temporally arranged images includes images generated by at least one sensor of a video camera.

3 . The non-transitory, processor-readable medium of claim 1 , further storing instructions to cause the processor to cause transmission of an alert to the compute device of the user in response to the signal.

4 . The non-transitory, processor-readable medium of claim 1 , wherein:

the set of at least one identification is associated with first embedded data;

the signal is associated with second embedded data; and

the search operation retrieves the at least one cropped image based on a comparison between the first embedded data and the second embedded data.

5 . The non-transitory, processor-readable medium of claim 1 , wherein the set of at least one identification is generated based on a motion associated with the object within at least two images from the plurality of temporally arranged images.

6 . The non-transitory, processor-readable medium of claim 1 , wherein:

the set of at least one identification is generated using a machine learning model and a plurality of images from the plurality of temporally arranged images; and

the set of at least one identification indicates an action.

7 . The non-transitory, processor-readable medium of claim 1 , further storing instructions to cause the processor to:

cause metadata associated with the plurality of temporally arranged images to be stored at the memory based on the index, the instructions to retrieve the at least one cropped image including instructions to retrieve the at least one cropped image from the memory based on the metadata.

8 . The non-transitory, processor-readable medium of claim 7 , wherein:

the metadata indicates a first time period; and

the signal encodes a second time period, the instructions to retrieve the at least one cropped image including instructions to retrieve the at least one cropped image from the memory based on the second time period including the first time period.

9 . The non-transitory, processor-readable medium of claim 1 , wherein:

the signal is a first signal;

the non-transitory, processor-readable medium further stores instructions to cause the processor to receive a second signal associated with a time period; and

the instructions to retrieve the at least one cropped image include instructions to retrieve the at least one cropped image from the memory based on the first signal and the second signal.

10 . The non-transitory, processor-readable medium of claim 1 , wherein the plurality of temporally arranged images is a first plurality of temporally arranged images, and the non-transitory, processor-readable medium further stores instructions to cause the processor to:

generate a second plurality of temporally arranged images based on the first plurality of temporally arranged images, the second plurality of temporally arranged images (1) having a resolution lower than a resolution of the first plurality of temporally arranged images and (2) used to generate the set of at least one identification.

11 . The non-transitory, processor-readable medium of claim 10 , wherein the instructions to generate the set of at least one identification include instructions to provide the second plurality of temporally arranged images as input to a machine learning model.

12 . An apparatus, comprising:

a processor; and

a memory operably coupled to the processor, the memory storing instructions to cause the processor to:

receive a plurality of temporally arranged images, each image from the plurality of temporally arranged images including a depiction of an object,

generate a plurality of temporally arranged compressed images based on the plurality of temporally arranged images,

generate an identification for the object based on at least one compressed image from the plurality of temporally arranged compressed images,

generate a set of at least one cropped image based on the identification, each cropped image from the set of at least one cropped image including a region of an image from the plurality of temporally arranged images, the region of the image being from a plurality of regions of that image, the image being different from each remaining image from the plurality of temporally arranged images, the region of the image being smaller than an entirety of the image and depicting the object, and

cause at least one cropped image from the set of at least one cropped image to be indexed in a searchable data structure based on the identification.

13 . The apparatus of claim 12 , wherein the memory further stores instructions to cause the processor to:

use the identification as a key to perform a search the searchable data structure for the at least one cropped image; and

retrieve the at least one cropped image from the searchable data structure in response to identifying, based on the search, the at least one cropped image indexed in the searchable data structure.

14 . The apparatus of claim 12 , further comprising a video camera operably coupled to the processor, the video camera configured to generate the plurality of temporally arranged images.

15 . The apparatus of claim 12 , wherein the memory further stores instructions to cause the processor to detect a motion of the object by providing the plurality of temporally arranged compressed images as input to a Kalman filter, the identification for the object being generated based on the motion.

16 . The apparatus of claim 12 , wherein:

the instructions to generate the identification include instructions to provide at least two temporally arranged compressed images from the plurality of temporally arranged compressed images as input to a machine learning model to generate the identification; and

the identification indicates an activity performed by the object.

17 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:

receive a plurality of temporally arranged images, each image from the plurality of temporally arranged images including a depiction of an object;

generate a plurality of temporally arranged compressed images based on the plurality of temporally arranged images;

generate detection data based on the plurality of temporally arranged compressed images;

generate a plurality of closeup images based on the plurality of temporally arranged images, each closeup image from the plurality of closeup images depicting an enlarged view of the object;

cause the plurality of closeup images and the detection data to be stored in association with each other at a database, the plurality of closeup images being indexed at the database based on the detection data;

retrieve, from the database and based on a text prompt that indicates a portion of the detection data, at least one closeup image that is from the plurality of closeup images and that is associated with the portion of the detection data; and

cause transmission of at least one signal to cause display of the at least one closeup image via a user interface of a remote compute device.

18 . The non-transitory, processor-readable medium of claim 17 , wherein:

the detection data includes a feature vector associated with at least one feature of the object.

19 . The non-transitory, processor-readable medium of claim 17 , further storing instructions to cause the processor to generate embedded data by providing the text prompt as input to a machine learning model, the instructions to retrieve the at least one closeup image including instructions to retrieve the at least one closeup image based on the embedded data.

20 . The non-transitory, processor-readable medium of claim 17 , wherein:

the text prompt is a first text prompt;

the at least one closeup image is a subset of selected closeup images from the plurality of closeup images;

the at least one signal is at least one first signal; and

the non-transitory, processor-readable medium further stores instructions to cause the processor to:

receive a second text prompt that indicates a subset of the portion of the detection data,

select at least one specific closeup image that is from the subset of selected closeup images and that is associated with the subset of the portion of the detection data, and

cause transmission of at least one second signal to cause display of the at least one specific closeup image via the user interface of the remote compute device.

Assignments (2)
SECURITY INTEREST Recorded Oct 1, 2024
From: VERKADA INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 068758/0910 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2024
From: GOYAL, RISHABH; CAO, SONG
To: VERKADA INC.
Reel/Frame 067854/0698 →
Continuity (2)
Division 18477280 · Sep 28, 2023
Related Publication 20250111665A1 · Apr 3, 2025
References Cited (23)
US 10386999B2 · Burns et al. · 2019 [cited by applicant]
US 12056918B1 · Goyal · 2024 [cited by examiner]
US 20040006628A1 · Shepard et al. · 2004 [cited by applicant]
US 20120062732A1 · Marman · 2012 [cited by examiner]
US 20150070587A1 · Emeott et al. · 2015 [cited by applicant]
US 20160314354A1 · Teuton et al. · 2016 [cited by applicant]
US 20170249386A1 · Ostrovsky-Berman et al. · 2017 [cited by applicant]
US 20170257612A1 · Emeott et al. · 2017 [cited by applicant]
US 20180365498A1 · Teuton et al. · 2018 [cited by applicant]
US 20190236371A1 · Boonmee et al. · 2019 [cited by applicant]
US 20210150235A1 · Meskimen et al. · 2021 [cited by applicant]
US 20210304736A1 · Kothari et al. · 2021 [cited by applicant]
US 20230019211A1 · Wang et al. · 2023 [cited by applicant]
US 20230351102A1 · Tran · 2023 [cited by applicant]
CN 116415034A · 2023 [cited by applicant]
WO WO2019100350A1 · 2019 [cited by applicant]
WO WO2021092632A2 · 2021 [cited by examiner]
WO WO2021092631A9 · 2021 [cited by applicant]
International Search Report and Written Opinion for PCT Application No. PCT/US2024/048116, by Verkada Inc., mailed Dec. 16, 2024; 15 pages. [cited by applicant]
Co-pending U.S. Appl. No. 18/477,280, inventors Goyal; Rishabh et al., filed Sep. 28, 2023. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/477,280 dated Apr. 15, 2024, 13 pages. [cited by applicant]
OPENAI: CLIP: Connecting text and images, retrieved from internet at https://openai.com/research/clip, on Dec. 28, 2023, 23 pages. [cited by applicant]
OPENAI: DALL-E 2, retrieved from internet at https://openai.com/dall-e-2, on Dec. 28, 2023, 17 pages. [cited by applicant]