IP Library Granted Patent US 12,596,747
Granted Patent B2
US 12,596,747 · App. 18/814,322 · Granted Apr 7, 2026

Natural language search over security videos

Inventors: Amit Rozner (Yehud, IL); Yohay Falik (Petah Tiqwa, IL); Ran Levy (Rishon Lezion, IL)
Assignee: TYCO FIRE & SECURITY GMBH
G06F16/73
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,747
App. No.
18/814,322
Granted
Apr 7, 2026
Kind
B2
Abstract

A system may be configured to provide natural language search over security videos. In some aspects, the system may generate a first representation of sampled video information in a multidimensional format via a first machine learning model, and receive a request including a natural language input. Further, the system may generate a second representation of the natural language input in the multidimensional format via a second machine learning model that is a different from the first machine learning model, and determine that the first representation has a predefined relationship with the second representation. In addition, the system may present the second representation as a response to the request based on the first representation having the predefined relationship with the second representation.

Claims (36)

1 . A method, comprising:

generating a first representation of sampled video information in a multidimensional format via a first machine learning model;

receiving a request including a natural language input;

generating a second representation of the natural language input in the multidimensional format via a second machine learning model that is a different from the first machine learning model;

determining that the first representation has a predefined relationship with the second representation based on a learned proximity metric between the first representation and the second representation in the multidimensional format; and

presenting the second representation as a response to the request based on the first representation having the predefined relationship with the second representation.

2 . The method of claim 1 , wherein receiving the request includes receiving a search query for a plurality of video frames corresponding to an event defined by the natural language input.

3 . The method of claim 1 , wherein receiving the request includes receiving the request for an alert identifying an occurrence of an event corresponding to the natural language input.

4 . The method of claim 1 , further comprising:

receiving video capture information from a video capture device; and

sampling the video capture information to generate the sampled video information.

5 . The method of claim 1 , wherein at least one of the first machine learning model and the second machine learning model is a transformer model.

6 . The method of claim 1 , wherein the first machine learning model is a convolutional neural network.

7 . The method of claim 1 , further comprising jointly training the first machine learning model and the second machine learning model using a common process.

8 . A system comprising:

at least one memory storing instructions thereon; and

at least one processor coupled to the at least one memory and configured by the instructions to:

generate a first representation of sampled video information in a multidimensional format via a first machine learning model;

receive a request including a natural language input;

generate a second representation of the natural language input in the multidimensional format via a second machine learning model that is a different from the first machine learning model;

determine that the first representation has a predefined relationship with the second representation based on a learned proximity metric between the first representation and the second representation in the multidimensional format; and

present the second representation as a response to the request based on the first representation having the predefined relationship with the second representation.

9 . The system of claim 8 , wherein the at least one processor is further configured by the instructions to receive a search query for a plurality of video frames corresponding an event defined by the natural language input.

10 . The system of claim 8 , wherein the at least one processor is further configured by the instructions to receive the request for an alert identifying an occurrence of an event corresponding to the natural language input.

11 . The system of claim 8 , wherein the at least one processor is further configured by the instructions to receive video capture information from a video capture device, and sample the video capture information to generate the sampled video information.

12 . The system of claim 8 , wherein at least one of the first machine learning model and the second machine learning model is a transformer model.

13 . The system of claim 8 , wherein the first machine learning model is a convolutional neural network.

14 . A non-transitory computer-readable device having instructions thereon that, when executed by at least one computing device, causes the at least one computing device to perform operations comprising:

generating a first representation of sampled video information in a multidimensional format via a first machine learning model;

receiving a request including a natural language input;

generating a second representation of the natural language input in the multidimensional format via a second machine learning model that is a different from the first machine learning model;

determining that the first representation has a predefined relationship with the second representation based on a learned proximity metric between the first representation and the second representation in the multidimensional format; and

presenting the second representation as a response to the request based on the first representation having the predefined relationship with the second representation.

15 . The non-transitory computer-readable device of claim 14 , wherein receiving the request includes receiving a search query for a plurality of video frames corresponding to an event defined by the natural language input.

16 . The non-transitory computer-readable device of claim 14 , wherein receiving the request includes receiving the request for an alert identifying an occurrence of an event corresponding to the natural language input.

17 . The method of claim 7 , wherein the common process comprises coordinated training to align the first representation and the second representation.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 4, 2024
From: ROZNER, AMIT; FALIK, YOHAY; LEVY, RAN
To: JOHNSON CONTROLS TYCO IP HOLDINGS LLP
Reel/Frame 068488/0953 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 4, 2024
From: JOHNSON CONTROLS TYCO IP HOLDINGS LLP
To: TYCO FIRE & SECURITY GMBH
Reel/Frame 068846/0843 →
Continuity (2)
Provisional Application 63578869 · Aug 25, 2023
Related Publication 20250068674A1 · Feb 27, 2025
References Cited (9)
US 10768769B2 · Choe · 2020 [cited by examiner]
US 10896342B2 · Gavrilyuk et al. · 2021 [cited by applicant]
US 20190147284A1 · Gavrilyuk · 2019 [cited by examiner]
US 20210264227A1 · Ma · 2021 [cited by examiner]
US 20230086735A1 · Yan · 2023 [cited by examiner]
US 20240355100A1 · Katole · 2024 [cited by examiner]
Tellex et al., “Towards Surveillance Video Search by Natural Language Query”, Copyright 2009 ACM 978-1-60558-480-5/09/07 (Year: 2009). [cited by examiner]
Erozel et al., “Natural language querying for video databases”, Information Sciences 178 (2008) 2534-2552 (Year: 2008). [cited by examiner]
Hakeem et al., “Semantic Video Search using Natural Language Queries”, Copyright 2009 ACM 978-1-60558-608-3/09/10 (Year: 2009). [cited by examiner]