IP Library › Granted Patent US 12,738,060
Granted Patent B1
US 12,738,060 · App. 18/759,294 · Granted Sep 15, 2026

Visual event processing using language models

Inventors: Kent V Lam (Diamond Bar, CA); Amey Laxman Gawde (Foothill Ranch, CA); Kevin Roderick Sellon (Santa Ana, CA); Pijung Roy Chen (Diamond Bar, CA); Dongyang Huang (Bellevue, WA); Shivakumar Venkatakrishna (Kirkland, WA)
Assignee: Amazon Technologies, Inc.
G06V20/47G06F16/7328G06V20/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,060
App. No.
18/759,294
Granted
Sep 15, 2026
Kind
B1
Abstract

Systems and methods for visual event processing using multi-modal language models include receiving, at a first device, first user input data representing a first user command to store occurrence of a type of visual, auditory or other event over time, the event associated with a physical activity involving movement, sound, or other type of detectable physical stimuli. An activity indicator describing the activity may be stored in the first device. The first device's functionality can be extended to process data from the sensors of a second device. As such, the activity indicator stored in the first device can be used to process a first query requesting that a model identify the activity from the data. The model may determine that the data depicts the activity and a first output (natural language, notification, smart home device response, etc.) may be generated indicating occurrence of the event.

Claims (114)

1 . A device, comprising:

one or more processors; and

non-transitory computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving first user input data representing a first user command to store occurrence of a visual event, the visual event associated with an activity physically performed by a user in association with an object other than the device;

generating, utilizing the first user input data, an activity indicator describing the visual event from the first user command;

storing the activity indicator in an activity catalog disposed on the device as one of multiple activity indicators;

receiving image data representing images in a field of view of a camera to the device, wherein the image data is received in response to motion being detected;

generating, utilizing the activity indicator stored in the activity catalog, a first query requesting that a language model (LM) identify the activity in association with the object and the user from the image data, wherein the LM is disposed on the device and is configured to determine results utilizing the image data as the input;

determining, at the device utilizing the LM and the first query, that the image data depicts the activity being performed in association with the object and the user;

storing first data indicating occurrence of the visual event in response to the activity being performed in association with the object and the user;

receiving second user input data representing a natural language request associated with stored visual events;

determining, utilizing the second user input data, that the natural language request is associated with the visual event; and

generating, utilizing the LM, a first natural language response to the second user input data, wherein generating the first natural language response is in response to the visual event being one of the visual events from the first data.

2 . The device of claim 1 , the operations further comprising:

determining, utilizing the LM, a description of an image data frame received at the device;

storing, at the device, data representing the description of the image data frame;

receiving third user input data requesting determination of whether a particular visual event has occurred, the particular visual event differing from the multiple activity indicators;

generating a second query requesting that the LM determine whether the particular visual event occurred;

determining, utilizing the LM and in response to the second query, that the description is associated with the particular visual event; and

generating a second natural language response to the third user input data, wherein generating the second natural language response is in response to determining that the description is associated with the particular visual event.

3 . The device of claim 1 , the operations further comprising:

determining account data associated with the device;

determining a routine associated with the account data, the routine indicating an action to be performed upon occurrence of an additional activity;

generating one of the multiple activity indicators utilizing the additional activity from the routine; and

wherein the LM selects the activity being performed from the multiple activity indicators.

4 . The device of claim 1 , the operations further comprising:

receiving third user input data requesting a summary of occurrences of the visual event over a period of time;

determining, utilizing the LM, additional occurrences of the activity from additional image data received over the period of time;

storing visual event identifiers of the activity on the device over the period of time; and

generating, utilizing the LM and the visual event identifiers, the summary of the occurrences of the visual event over the period of time.

5 . A method, comprising:

receiving first user input data representing a first user command to store an activity indictor representing an occurrence of a visual event, the visual event associated with an activity in an environment;

storing an activity indicator in a device;

generating, utilizing the activity indicator stored in the device, a first query requesting that a language model (LM) identify the activity from image data, wherein the LM is disposed on the device and is configured to determine results utilizing the image data as the input;

determining, at the device utilizing the LM and the first query, that the image data depicts the activity;

storing first data indicating occurrence of the visual event based at least in part on the LM determining that the image data depicts the activity;

receiving second user input data representing a natural language request associated with stored visual events;

determining, utilizing the LM and based at least in part on the second user input data, that the natural language request is associated with the visual event; and

generating a first natural language response to the second user input data indicating that the visual event has been detected, wherein generating the first natural language response is based at least in part on the first data indicating occurrence of the visual event.

6 . The method of claim 5 , further comprising:

determining, utilizing the LM, a description of an image data frame received at the device;

receiving third user input data requesting determination of whether an additional visual event has occurred, the additional visual event differing from the activity indicator;

determining, utilizing the LM and based at least in part on the third user input data, that the description is associated with the additional visual event; and

generating a second natural language response indicating occurrence of the additional visual event based at least in part on determining that the description is associated with the additional visual event.

7 . The method of claim 5 , further comprising:

determining a routine associated with the device, the routine indicating an action to be performed upon occurrence of an additional activity; and

generating an additional activity indicator based at least in part on the routine.

8 . The method of claim 5 , further comprising:

receiving third user input data requesting a summary of occurrences of the visual event over a period of time;

determining, utilizing the LM, additional occurrences of the activity from additional image data received over the period of time;

storing visual event identifiers of the activity on the device over the period of time; and

generating, utilizing the LM and the visual event identifiers, the summary of the occurrences of the visual event over the period of time.

9 . The method of claim 5 , further comprising:

determining a keyframe representing the image data;

inputting the keyframe into the LM;

receiving, from the LM, a description of the keyframe;

storing the description of the keyframe on the device;

querying at least one of the LM or another LM disposed on the device to determine, from the description of the keyframe, whether the image data depicts the activity; and

wherein determining that the image data depicts the activity comprises determining that the image data depicts the activity based at least in part on the description of the keyframe.

10 . The method of claim 5 , further comprising:

determining to store the first data including the image data and an indication that the activity was detected for a first period of time based at least in part on the visual event being detected;

receiving additional image data; and

determining to store the additional image data for a second period of time based at least in part on the additional image data being unassociated with one or more activity indicators stored on the device, the second period of time being less than the first period of time.

11 . The method of claim 5 , further comprising:

determining that the first user input data indicates an object associated with the activity and a user associated with the activity;

determining, utilizing the LM, that the activity determined from the image data is performed in association with the object and the user; and

wherein generating the first natural language response comprises generating the first natural language response based at least in part on the activity being performed in association with the object and the user.

12 . The method of claim 5 , further comprising:

determining, from the image data, a subset of frames of the image data to be analyzed by the LM at the device, the subset of frames determined based at least in part on pixel differences detected as between consecutive frames of the image data;

inputting the subset of frames into the LM along with the first query, wherein the first query prompts the LM to determine whether the activity is determined from individual frames of the subset of frames; and

receiving, from the LM, an identification of which of the individual frames are associated with the activity.

13 . A system, comprising:

one or more processors; and

non-transitory computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving first user input data representing a first user command to store an activity indictor representing an occurrence of a visual event, the visual event associated with an activity in an environment;

storing an activity indicator in a device;

generating, utilizing the activity indicator stored in the device, a first query requesting that a language model (LM) identify the activity from image data, wherein the LM is disposed on the device and is configured to determine results utilizing the image data as the input;

determining, at the device utilizing the LM and the first query, that the image data depicts the activity;

storing first data indicating occurrence of the visual event based at least in part on the LM determining that the image data depicts the activity;

receiving second user input data representing a natural language request associated with stored visual events;

determining, utilizing the LM and based at least in part on the second user input data, that the natural language request is associated with the visual event; and

generating a first natural language response to the second user input data indicating that the visual event has been detected, wherein generating the first natural language response is based at least in part on the first data indicating occurrence of the visual event.

14 . The system of claim 13 , the operations further comprising:

determining, utilizing the LM, a description of an image data frame received at the device;

receiving third user input data requesting determination of whether an additional visual event has occurred, the additional visual event differing from the activity indicator;

determining, utilizing the LM and based at least in part on the third user input data, that the description is associated with the additional visual event; and

generating a second natural language response indicating occurrence of the additional visual event based at least in part on determining that the description is associated with the additional visual event.

15 . The system of claim 13 , the operations further comprising:

determining a routine associated with the device, the routine indicating an action to be performed upon occurrence of an additional activity; and

generating an additional activity indicator based at least in part on the routine.

16 . The system of claim 13 , the operations further comprising:

receiving third user input data requesting a summary of occurrences of the visual event over a period of time;

determining, utilizing the LM, additional occurrences of the activity from additional image data received over the period of time;

storing visual event identifiers of the activity on the device over the period of time; and

generating, utilizing the LM and the visual event identifiers, the summary of the occurrences of the visual event over the period of time.

17 . The system of claim 13 , the operations further comprising:

determining a keyframe representing the image data;

inputting the keyframe into the LM;

receiving, from the LM, a description of the keyframe;

storing the description of the keyframe on the device;

querying at least one of the LM or another LM disposed on the device to determine, from the description of the keyframe, whether the image data depicts the activity; and

wherein determining that the image data depicts the activity comprises determining that the image data depicts the activity based at least in part on the description of the keyframe.

18 . The system of claim 13 , the operations further comprising:

determining to store the first data including the image data and an indication that the activity was detected for a first period of time based at least in part on the visual event being detected;

receiving additional image data; and

determining to store the additional image data for a second period of time based at least in part on the additional image data being unassociated with one or more activity indicators stored on the device, the second period of time being less than the first period of time.

19 . The system of claim 13 , the operations further comprising:

determining that the first user input data indicates an object associated with the activity and a user associated with the activity;

determining, utilizing the LM, that the activity determined from the image data is performed in association with the object and the user; and

wherein generating the first natural language response comprises generating the first natural language response based at least in part on the activity being performed in association with the object and the user.

20 . The system of claim 13 , the operations further comprising:

determining, from the image data, a subset of frames of the image data to be analyzed by the LM at the device, the subset of frames determined based at least in part on pixel differences detected as between consecutive frames of the image data;

inputting the subset of frames into the LM along with the first query, wherein the first query prompts the LM to determine whether the activity is determined from individual frames of the subset of frames; and

receiving, from the LM, an identification of which of the individual frames are associated with the activity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2024
From: LAM, KENT V; GAWDE, AMEY LAXMAN; SELLON, KEVIN RODERICK; CHEN, PIJUNG ROY; HUANG, DONGYANG; VENKATAKRISHNA, SHIVAKUMAR
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069212/0792 →
References Cited (6)
US 11393253B1 · Maron · 2022 [cited by examiner]
US 20200310749A1 · Miller · 2020 [cited by examiner]
US 20230334857A1 · Goldstein · 2023 [cited by examiner]
US 20250111674A1 · Li · 2025 [cited by examiner]
CN 117237488A · 2023 [cited by examiner]
Ju, Chen, et al. “Prompting visual-language models for efficient video understanding.” European conference on computer vision. Cham: Springer Nature Switzerland, 2022.https://link.springer.com/chapter/10.1007/978-3-031-… [cited by examiner]