IP Library Granted Patent US 12,705,110
Granted Patent B2
US 12,705,110 · App. 18/203,539 · Granted Aug 11, 2026

Curation of custom workflows using multiple cameras, with AI trained from the workflow

Inventors: David D. Lee (Palo Alto, CA); Andrew Augustine Wajs (Haarlem, NL)
Assignee: Scenera, Inc.
G06F9/5072G06F9/542G06F18/251G06F18/285G06N20/00G06V10/454G06V10/764G06V10/82G06V10/94G06V20/52H04N23/80H04N23/90G06V40/161G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,110
App. No.
18/203,539
Filed
May 30, 2023
Granted
Aug 11, 2026
Kind
B2
Art Unit
2661
USPC
382/155
Abstract

A multi-layer technology stack includes a sensor layer including image sensors, a device layer, and a cloud layer, with interfaces between the layers. A method to curate different custom workflows for multiple applications include the following. Requirements for custom sets of data packages for the applications is received. The custom set of data packages include sensor data packages (e.g., SceneData) and contextual metadata packages that contextualize the sensor data packages (e.g., SceneMarks). Based on the received requirements and capabilities of components in the technology stack, the custom workflow for that application is deployed. This includes a selection, configuration and linking of components from the technology stack. The custom workflow is implemented in the components of the technology stack by transmitting workflow control packages directly and/or indirectly via the interfaces to the different layers.

Claims (39)

1 . A workflow method for a plurality of non-human technological entities to develop a situational awareness of a space, the entities including cameras that view the monitored space, wherein the workflow method comprises:

the cameras capturing images of the monitored space;

artificial intelligence and/or machine learning (AI/ML) entities detecting events from the captured images;

generating SceneMarks with attributes that are descriptive of the detected events;

a first AI/ML entity performing analysis of the SceneMarks; and adding information to the attributes of at least one of the SceneMarks based on said analysis, and/or detecting an event based on said analysis and generating a new SceneMark for said detected event;

wherein the workflow method contextualizes the captured images from the SceneMarks to provide awareness of situations in the monitored space; and the first AI/ML entity is trained based on Scenemarks generated by execution of the workflow method.

2 . The workflow method of claim 1 wherein the first AI/ML entity is trained to predict a generation of SceneMarks.

3 . The workflow method of claim 2 wherein the first AI/ML entity is trained to predict the generation of SceneMarks, based on received SceneMarks.

4 . The workflow method of claim 1 wherein the first AI/ML entity is trained to predict a generation of attributes for SceneMarks.

5 . The workflow method of claim 1 further comprising:

grouping SceneMarks into Scenes, wherein the first AI/ML entity is trained to group SceneMarks into Scenes.

6 . The workflow method of claim 1 further comprising:

automatically triggering action based on the SceneMarks, wherein the first AI/ML entity is trained to learn which actions are triggered by which sequences of SceneMarks.

7 . The workflow method of claim 1 wherein the first AI/ML entity is trained to determine when a person or object identified in two different SceneMarks are the same person or object.

8 . The workflow method of claim 7 wherein the first AI/ML entity determines that the person or object in the two different SceneMarks are the same person or object, based on attributes of the two different SceneMarks.

9 . The workflow method of claim 7 wherein the first AI/ML entity determines that the person or object in the two different SceneMarks are the same person or object, based on timestamps of the two different SceneMarks and a known proximity of cameras capturing the images that generated the two different SceneMarks.

10 . The workflow method of claim 1 wherein the first AI/ML entity is a generative pre-trained transformer AI entity.

11 . The workflow method of claim 1 wherein the first AI/ML entity is trained based on labeled SceneMarks.

12 . The workflow method of claim 11 wherein the labeled SceneMarks include labels that group SceneMarks into Scenes.

13 . The workflow method of claim 11 wherein the labeled SceneMarks include labels that are descriptive of the situation in the monitored space.

14 . The workflow method of claim 1 wherein the first AI/ML entity is trained on SceneMark tokens.

15 . The workflow method of claim 1 wherein the situation occurs across multiple cameras viewing the space, and providing awareness of the situation is based on SceneMarks generated from images captured by the multiple cameras and is also based on a known proximity of the cameras.

16 . The workflow method of claim 1 further comprising:

dynamically adjusting the entities in the workflow and/or the analysis performed by the entities, based on the SceneMarks.

17 . The workflow method of claim 1 wherein the entities include other non-camera sensors, and the workflow method further comprises:

the non-camera sensors capture sensor data for the monitored space; and

the AI/ML entities detecting events from the captured images and the captured sensor data.

18 . The workflow method of claim 1 further comprising:

receiving descriptions of capabilities of different entities; and

transmitting workflow control packages to the entities to implement the workflow method on the entities.

19 . A system comprising:

a plurality of non-human technological entities, the entities including cameras that view a space and further including artificial intelligence and/or machine learning (AI/ML) entities; and

a service that configures the entities to implement a workflow to develop a situational awareness of the space, wherein the workflow includes:

capturing images of the space;

detecting events from the captured images;

generating SceneMarks with attributes that are descriptive of the detected events;

at least one AI/ML entity performing analysis of the SceneMarks; and adding information to the attributes of at least one of the SceneMarks based on said analysis, and/or detecting an event based on said analysis and generating a new SceneMark for said detected event;

wherein the workflow contextualizes the captured images from the SceneMarks to provide awareness of situations in the space; and a first of the AI/ML entities is trained based on Scenemarks generated by execution of the workflow.

20 . The system of claim 19 wherein the first AI/ML entity is trained to determine when a person or object identified in two different SceneMarks are the same person or object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2023
From: LEE, DAVID D.; WAJS, ANDREW AUGUSTINE
To: SCENERA, INC.
Reel/Frame 064989/0210 →
Continuity (6)
Continuation In Part 17341794 · Jun 8, 2021
Continuation 17084417 · Oct 29, 2020
Provisional Application 63020521 · May 5, 2020
Provisional Application 62928199 · Oct 30, 2019
Provisional Application 62928165 · Oct 30, 2019
Related Publication 20230325255A1 · Oct 12, 2023
References Cited (89)
US 6654047B2 · Lizaka · 2003 [cited by examiner]
US 10055853B1 · Fisher · 2018 [cited by examiner]
US 10412291B2 · Lee et al. · 2019 [cited by applicant]
US 10509459B2 · Lee et al. · 2019 [cited by applicant]
US 10693843B2 · Lee et al. · 2020 [cited by applicant]
US 11068742B2 · Lee · 2021 [cited by examiner]
US 11068991B2 · Przechocki · 2021 [cited by examiner]
US 11188758B2 · Lee · 2021 [cited by examiner]
US 11328513B1 · Osherovich · 2022 [cited by examiner]
US 11600070B2 · Elgamal et al. · 2023 [cited by applicant]
US 11663049B2 · Lee · 2023 [cited by examiner]
US 11961303B1 · Osherovich · 2024 [cited by examiner]
US 12272137B2 · Xu · 2025 [cited by examiner]
US 12400442B2 · Bates · 2025 [cited by examiner]
US 20050120330A1 · Ghai et al. · 2005 [cited by applicant]
US 20080028363A1 · Mathew · 2008 [cited by applicant]
US 20100185973A1 · Ali et al. · 2010 [cited by applicant]
US 20100207762A1 · Lee et al. · 2010 [cited by applicant]
US 20110173328A1 · Park et al. · 2011 [cited by applicant]
US 20120026335A1 · Brown · 2012 [cited by examiner]
US 20130055201A1 · No et al. · 2013 [cited by applicant]
US 20130124807A1 · Nielsen et al. · 2013 [cited by applicant]
US 20140201707A1 · Schroeder · 2014 [cited by applicant]
US 20140350997A1 · Holm et al. · 2014 [cited by applicant]
US 20140362223A1 · LaCroix · 2014 [cited by examiner]
US 20150012396A1 · Puerini · 2015 [cited by examiner]
US 20150127485A1 · Kakizawa · 2015 [cited by examiner]
US 20160004390A1 · Laska et al. · 2016 [cited by applicant]
US 20160034809A1 · Trenholm et al. · 2016 [cited by applicant]
US 20170006135A1 · Siebel et al. · 2017 [cited by applicant]
US 20170063886A1 · Muddu · 2017 [cited by examiner]
US 20170316586A1 · Ricci · 2017 [cited by applicant]
US 20170336858A1 · Lee et al. · 2017 [cited by applicant]
US 20170337425A1 · Lee · 2017 [cited by examiner]
US 20170339329A1 · Lee et al. · 2017 [cited by applicant]
US 20180018508A1 · Tusch · 2018 [cited by examiner]
US 20180068172A1 · Despiegel · 2018 [cited by examiner]
US 20180069838A1 · Lee et al. · 2018 [cited by applicant]
US 20180096209A1 · Matsuda · 2018 [cited by examiner]
US 20180348092A1 · Suresh et al. · 2018 [cited by applicant]
US 20190043201A1 · Strong · 2019 [cited by examiner]
US 20190095716A1 · Shrestha et al. · 2019 [cited by applicant]
US 20190147251A1 · Numata · 2019 [cited by examiner]
US 20190207866A1 · Pathak et al. · 2019 [cited by applicant]
US 20190258864A1 · Lee et al. · 2019 [cited by applicant]
US 20190354773A1 · Leizerovich · 2019 [cited by examiner]
US 20190354774A1 · Leizerovich · 2019 [cited by examiner]
US 20190354775A1 · Leizerovich · 2019 [cited by examiner]
US 20190354776A1 · Ribeiro · 2019 [cited by examiner]
US 20190356885A1 · Ribeiro · 2019 [cited by examiner]
US 20200053116A1 · Soroush et al. · 2020 [cited by applicant]
US 20200117897A1 · Froloff · 2020 [cited by examiner]
US 20200293803A1 · Wajs · 2020 [cited by examiner]
US 20200342324A1 · Sivaraman et al. · 2020 [cited by applicant]
US 20200372381A1 · Trim · 2020 [cited by examiner]
US 20210042638A1 · Novotny · 2021 [cited by applicant]
US 20210133456A1 · Lee · 2021 [cited by examiner]
US 20210133492A1 · Lee · 2021 [cited by examiner]
US 20210295094A1 · Lee · 2021 [cited by examiner]
US 20210306560A1 · Lee · 2021 [cited by examiner]
US 20210397848A1 · Lee · 2021 [cited by examiner]
US 20220114477A1 · Kou · 2022 [cited by examiner]
US 20220122015A1 · Bandopadhyay · 2022 [cited by examiner]
US 20220277193A1 · Wekel et al. · 2022 [cited by applicant]
US 20230267775A1 · Dhoot · 2023 [cited by examiner]
US 20230305903A1 · Lee · 2023 [cited by examiner]
US 20230325255A1 · Lee · 2023 [cited by examiner]
US 20250071027A1 · Bakshi · 2025 [cited by examiner]
US 20250209328A1 · Zhuang · 2025 [cited by examiner]
US 20250308219A1 · Smith et al. · 2025 [cited by applicant]
US 20260044711A1 · Hoshina · 2026 [cited by examiner]
Girdhar, R. et al., “Video Action Transformer Network,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 244-253. [cited by applicant]
Horev, R., “Bert Explained: State of the art language model for NLP,” Nov. 10, 2018, 8 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://towardsdatascience.com/bert-explained-state-o… [cited by applicant]
Kong, Y. et al., “Human Action Recognition and Prediction: A Survey,” arXiv.1806.11230, Jul. 2, 2018, pp. 1-20. [cited by applicant]
Loginova, K., “Attention in NLP,” Jun. 22, 2018, 16 pages, [Online] [Retrieved on Jan. 20, 2021] Retrieved from the Internet <URL: https://medium.com/@edloginova/aftention-in-nlp-734c6fa9d983>. [cited by applicant]
Olah, C. et al., “Attention and Augmented Recurrent Neural Networks,” Sep. 8, 2016, 19 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://distill.pub/2016/augmented-rnns/>. [cited by applicant]
Olah, C., “Understanding LSTM Networks,” Aug. 27, 2015, 8 pages, [Online] [Retrieved on Jan. 20, 2021] Retrieved from the Internet <URL: https://colah.github.io/posts/2015-08-Understanding-LSTMs/>. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/US20/58193, dated Feb. 2, 2021, 14 pages. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/US20/58198, dated Feb. 2, 2021, 12 pages. [cited by applicant]
Piergiovanni, A., et al., “Tiny Video Networks,” arXiv:1910.06961, Oct. 15, 2019, pp. 1-10. [cited by applicant]
Rivera-Soto, R.A. et al., “Sequence to Sequence Models for Generating Video Captions,” Stanford University, Jul. 2, 2017, pp. 1-7. [cited by applicant]
Rosset, C., “Turing-NLG: A 17-billion-parameter language model by Microsoft,” Feb. 13, 2020, 11 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://www.microsoft.com/en-us/research/blo… [cited by applicant]
Security World Market, “Androvideo's AI camera for security & safety,” Nov. 14, 2019, 6 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://www.securityworldmarket.com/int/News/Product… [cited by applicant]
Sharma, A.K., “Predicting Human Behaviour Activity using Deep Leaming (LSTM),” May 26, 2018, 12 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://medium.com/@chataks93/predicting-hum… [cited by applicant]
Sun, C. et al., “VideoBERT: A Joint Model for Video and Language Representation Learning,” arXiv:1904.01766, Sep. 11, 2019, pp. 1-13. [cited by applicant]
United States Office Action, U.S. Appl. No. 17/084,429, dated Feb. 4, 2021, 16 pages. [cited by applicant]
Vincent, J., “This Japanese AI security camera shows the future of surveillance will be automated,” Jun. 26, 2018, four pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://www.theverge… [cited by applicant]
United States Office Action, U.S. Appl. No. 17/345,648, filed Jan. 28, 2026, 16 pages. [cited by applicant]
United States Office Action, U.S. Appl. No. 18/203,535, dated Oct. 1, 2025, 37 pages. [cited by applicant]