IP Library Granted Patent US 12688680
Granted Patent B2
US 12688680 · App. 18/519,276 · Granted Jul 21, 2026

Image analytics using embeddings

Inventors: Thomas Cilloni (Oxford, MS); Charles Fleming (Oxford, MS); Ramana Rao V. R. Kompella (Foster City, CA); Gaowen Liu (Austin, TX)
Assignee: Cisco Technology, Inc.
G06V10/774G06V10/945G06V20/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12688680
App. No.
18/519,276
Granted
Jul 21, 2026
Kind
B2
Abstract

In one implementation, a device receives, via a user interface, one or more parameters regarding formation of an embedding of an image. The device forms, in accordance with the one or more parameters, an embedding of the image by inputting it to a machine learning-based encoder model that was trained to maximize a measure of data utility of its output embeddings. The device provides the embedding for use to train an analytics model. The device causes the analytics model to be used to make inferences about embeddings derived from images captured by one or more cameras.

Claims (37)

1 . A method comprising:

receiving, at a device and via a user interface, one or more parameters regarding formation of an embedding of an image, wherein the one or more parameters specify a number of bits for each feature of the image represented in the embedding;

forming, by the device and in accordance with the one or more parameters, an embedding of the image by inputting it to a machine learning-based encoder model that was trained to maximize a measure of data utility of its output embeddings;

providing, by the device, the embedding for use to train an analytics model; and

causing, by the device, the analytics model to be used to make inferences about embeddings derived from images captured by one or more cameras.

2 . The method as in claim 1 , wherein the one or more parameters specify a file type for the embedding of the image.

3 . The method as in claim 1 , wherein the one or more parameters specify a location of the image.

4 . The method as in claim 1 , wherein the machine learning-based encoder model maximizes the measure of data utility based on a type of the analytics model indicated by the one or more parameters.

5 . The method as in claim 1 , wherein the one or more parameters specify a maximum size of the embedding and a maximum number of features of the image represented in the embedding.

6 . The method as in claim 1 , further comprising:

using embeddings formed by the machine learning-based encoder model as input to the analytics model.

7 . The method as in claim 1 , wherein the inferences by the analytics model comprise identification of a particular type of object depicted in the images from which the embeddings were derived.

8 . The method as in claim 7 , wherein the images are from a video feed from one or more cameras.

9 . The method as in claim 1 , wherein the device is an edge device in a network.

10 . An apparatus, comprising:

a network interface to communicate with a computer network;

a processor coupled to the network interface and configured to execute one or more processes; and

a memory configured to store a process that is executed by the processor, the process when executed configured to:

receive, via a user interface, one or more parameters regarding formation of an embedding of an image, wherein the one or more parameters specify a number of bits for each feature of the image represented in the embedding;

form, in accordance with the one or more parameters, an embedding of the image by inputting it to a machine learning-based encoder model that was trained to maximize a measure of data utility of its output embeddings;

provide the embedding for use to train an analytics model; and

cause the analytics model to be used to make inferences about embeddings derived from images captured by one or more cameras.

11 . The apparatus as in claim 10 , wherein the one or more parameters specify a file type for the embedding of the image.

12 . The apparatus as in claim 10 , wherein the one or more parameters specify a location of the image.

13 . The apparatus as in claim 10 , wherein the machine learning-based encoder model maximizes the measure of data utility based on a type of the analytics model indicated by the one or more parameters.

14 . The apparatus as in claim 10 , wherein the one or more parameters specify a maximum size of the embedding and a maximum number of features of the image represented in the embedding.

15 . The apparatus as in claim 10 , wherein the process when executed is further configured to:

use embeddings formed by the machine learning-based encoder model as input to the analytics model.

16 . The apparatus as in claim 10 , wherein the inferences by the analytics model comprise identification of a particular type of object depicted in the images from which the embeddings were derived.

17 . The apparatus as in claim 16 , wherein the images are from a video feed from one or more cameras.

18 . A tangible, non-transitory, computer-readable medium storing program instructions that cause a device to execute a process comprising:

receiving, at a device and via a user interface, one or more parameters regarding formation of an embedding of an image, wherein the one or more parameters specify a number of bits for each feature of the image represented in the embedding;

forming, by the device and in accordance with the one or more parameters, an embedding of the image by inputting it to a machine learning-based encoder model that was trained to maximize a measure of data utility of its output embeddings;

providing, by the device, the embedding for use to train an analytics model; and

causing, by the device, the analytics model to be used to make inferences about embeddings derived from images captured by one or more cameras.

19 . The method as in claim 1 , wherein the analytics model is trained to perform inference directly on the embedding without reconstructing the image from the embedding.

20 . The method as in claim 1 , wherein the machine learning-based encoder model was trained to maximize the measure of data utility by optimizing performance of a reference analytics model on embeddings generated by the machine learning-based encoder model during training.