IP Library Granted Patent US 12,705,109
Granted Patent B2
US 12,705,109 · App. 18/203,535 · Granted Aug 11, 2026

Curation of custom workflows using multiple cameras, with AI to provide awareness of situations

Inventors: David D. Lee (Palo Alto, CA); Andrew Augustine Wajs (Haarlem, NL)
Assignee: Scenera, Inc.
G06F9/5072G06F9/542G06F18/251G06F18/285G06N20/00G06V10/454G06V10/764G06V10/82G06V10/94G06V20/52H04N23/80H04N23/90G06V40/161G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,109
App. No.
18/203,535
Filed
May 30, 2023
Granted
Aug 11, 2026
Kind
B2
Art Unit
2661
USPC
382/155
Abstract

A multi-layer technology stack includes a sensor layer including image sensors, a device layer, and a cloud layer, with interfaces between the layers. A method to curate different custom workflows for multiple applications include the following. Requirements for custom sets of data packages for the applications is received. The custom set of data packages include sensor data packages (e.g., SceneData) and contextual metadata packages that contextualize the sensor data packages (e.g., SceneMarks). Based on the received requirements and capabilities of components in the technology stack, the custom workflow for that application is deployed. This includes a selection, configuration and linking of components from the technology stack. The custom workflow is implemented in the components of the technology stack by transmitting workflow control packages directly and/or indirectly via the interfaces to the different layers.

Claims (45)

1 . A method for enabling application-configured awareness of spaces, the method comprising:

receiving, via an application programming interface (API), requests from applications for different monitorings of spaces; and

configuring a plurality of non-human technological entities to implement workflows for the requested monitorings of spaces, the entities including cameras that view the monitored spaces, wherein the workflows include:

the cameras capturing images of the monitored spaces;

artificial intelligence and/or machine learning (AI/ML) entities detecting events from the captured images;

generating SceneMarks with attributes that are descriptive of the detected events;

transmitting the SceneMarks between the AI/ML entities;

at least one AI/ML entity performing analysis of SceneMarks received via the transmission between the AI/ML entities; and adding information to the attributes of at least one of the received SceneMarks based on said analysis, and/or detecting an event based on said analysis and generating a new SceneMark for said detected event; and

the workflows contextualize the captured images from the SceneMarks to provide awareness of situations in the monitored spaces.

2 . The method of claim 1 wherein the at least one AI/ML entity performs analysis of received SceneMarks to detect an anomaly in the situation in the monitored space.

3 . The method of claim 2 wherein the at least one AI/ML entity detects the anomaly based on a sequence of received SceneMarks.

4 . The method of claim 2 wherein the anomaly is one of: an unexpected occupancy of the space, an unusual movement of a person through the space, an unusual interaction between people in the space, an unexpected object in the space, or an unexpected condition for the space.

5 . The method of claim 2 wherein the at least one AI/ML entity detects the anomaly based on comparing received SceneMarks with SceneMarks produced by a normal situation in the monitored space.

6 . The method of claim 2 wherein the at least one AI/ML entity detects the anomaly based on comparing received SceneMarks with SceneMarks predicted for a normal situation in the monitored space.

7 . The method of claim 1 wherein the at least one AI/ML entity determines that a person or object identified in two different SceneMarks are the same person or object.

8 . The method of claim 7 wherein the at least one AI/ML entity determines that the person or object in the two different SceneMarks are the same person or object, based on attributes of the two different SceneMarks.

9 . The method of claim 7 wherein the at least one AI/ML entity determines that the person or object in the two different SceneMarks are the same person or object, based on timestamps of the two different SceneMarks and a known proximity of cameras capturing the images that generated the two different SceneMarks.

10 . The method of claim 1 wherein the workflow further includes:

automatically triggering an action based on a sequence of SceneMarks that are indicative of a predefined situation in the monitored space.

11 . The method of claim 1 wherein the workflow further includes:

classifying received SceneMarks into different predefined categories; and

automatically triggering different actions based on the category.

12 . The method of claim 1 wherein the workflow further includes:

generating a text description of the situation in the monitored space, based on received SceneMarks.

13 . The method of claim 12 wherein a generative pre-trained transformer entity generates the text description.

14 . The method of claim 12 wherein the workflow further includes:

returning the text description of the situation to the requesting application.

15 . The method of claim 12 wherein generating the text description comprises:

generating labels based on the received SceneMarks; and

matching the generated labels against predefined labels that describe different situations.

16 . The method of claim 1 wherein the workflow further includes:

accessing SceneMarks stored in a SceneMark database, wherein the workflows contextualize the captured images from SceneMarks including the stored SceneMarks.

17 . The method of claim 16 wherein a generative pre-trained transformer entity formulates a query to access the SceneMarks stored in the SceneMark database.

18 . The method of claim 16 wherein a generative pre-trained transformer entity formulates a natural language response based on SceneMarks returned from the SceneMark database in response to a query.

19 . The method of claim 1 wherein the SceneMarks include links to the captured images, and the SceneMarks are transmitted between entities but the captured images are not transmitted between entities.

20 . A system comprising:

a plurality of applications that make requests for different monitorings of spaces;

a plurality of non-human technological entities, the entities including cameras that view the monitored spaces and further including artificial intelligence and/or machine learning (AI/ML) entities;

a service that receives the requests and configures the entities to implement workflows for the requested monitorings of spaces, wherein the workflows include:

the cameras capturing images of the monitored spaces;

the AI/ML entities detecting events from the captured images;

generating SceneMarks with attributes that are descriptive of the detected events;

transmitting the SceneMarks between the AI/ML entities;

at least one AI/ML entity performing analysis of SceneMarks received via the transmission between the AI/ML entities; and adding information to the attributes of at least one of the received SceneMarks based on said analysis, and/or detecting an event based on said analysis and generating a new SceneMark for said detected event; and

the workflows contextualize the captured images from the SceneMarks to provide awareness of situations in the monitored spaces.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2023
From: LEE, DAVID D.; WAJS, ANDREW AUGUSTINE
To: SCENERA, INC.
Reel/Frame 064989/0236 →
Continuity (6)
Continuation In Part 17341794 · Jun 8, 2021
Continuation 17084417 · Oct 29, 2020
Provisional Application 63020521 · May 5, 2020
Provisional Application 62928199 · Oct 30, 2019
Provisional Application 62928165 · Oct 30, 2019
Related Publication 20230305903A1 · Sep 28, 2023
References Cited (88)
US 6654047B2 · Iizaka · 2003 [cited by examiner]
US 10055853B1 · Fisher · 2018 [cited by examiner]
US 10412291B2 · Lee et al. · 2019 [cited by applicant]
US 10509459B2 · Lee et al. · 2019 [cited by applicant]
US 10693843B2 · Lee et al. · 2020 [cited by applicant]
US 11068742B2 · Lee · 2021 [cited by examiner]
US 11068991B2 · Przechocki · 2021 [cited by examiner]
US 11188758B2 · Lee et al. · 2021 [cited by applicant]
US 11328513B1 · Osherovich · 2022 [cited by examiner]
US 11600070B2 · Elgamal et al. · 2023 [cited by applicant]
US 11663049B2 · Lee et al. · 2023 [cited by applicant]
US 11961303B1 · Osherovich et al. · 2024 [cited by applicant]
US 12272137B2 · Xu · 2025 [cited by examiner]
US 12400442B2 · Bates · 2025 [cited by examiner]
US 20050120330A1 · Ghai et al. · 2005 [cited by applicant]
US 20080028363A1 · Mathew · 2008 [cited by applicant]
US 20100185973A1 · Ali et al. · 2010 [cited by applicant]
US 20100207762A1 · Lee et al. · 2010 [cited by applicant]
US 20110173328A1 · Park et al. · 2011 [cited by applicant]
US 20120026335A1 · Brown · 2012 [cited by examiner]
US 20130055201A1 · No et al. · 2013 [cited by applicant]
US 20130124807A1 · Nielsen et al. · 2013 [cited by applicant]
US 20140201707A1 · Schroeder · 2014 [cited by applicant]
US 20140350997A1 · Holm et al. · 2014 [cited by applicant]
US 20140362223A1 · LaCroix · 2014 [cited by examiner]
US 20150012396A1 · Puerini · 2015 [cited by examiner]
US 20150127485A1 · Kakizawa · 2015 [cited by examiner]
US 20160004390A1 · Laska et al. · 2016 [cited by applicant]
US 20160034809A1 · Trenholm et al. · 2016 [cited by applicant]
US 20170006135A1 · Siebel et al. · 2017 [cited by applicant]
US 20170063886A1 · Muddu · 2017 [cited by examiner]
US 20170316586A1 · Ricci · 2017 [cited by applicant]
US 20170336858A1 · Lee et al. · 2017 [cited by applicant]
US 20170337425A1 · Lee · 2017 [cited by examiner]
US 20170339329A1 · Lee · 2017 [cited by examiner]
US 20180018508A1 · Tusch · 2018 [cited by applicant]
US 20180068172A1 · Despiegel · 2018 [cited by examiner]
US 20180069838A1 · Lee · 2018 [cited by examiner]
US 20180096209A1 · Matsuda · 2018 [cited by examiner]
US 20180348092A1 · Suresh et al. · 2018 [cited by applicant]
US 20190043201A1 · Strong et al. · 2019 [cited by applicant]
US 20190095716A1 · Shrestha · 2019 [cited by examiner]
US 20190147251A1 · Numata · 2019 [cited by examiner]
US 20190207866A1 · Pathak et al. · 2019 [cited by applicant]
US 20190258864A1 · Lee et al. · 2019 [cited by applicant]
US 20190354773A1 · Leizerovich · 2019 [cited by examiner]
US 20190354774A1 · Leizerovich · 2019 [cited by examiner]
US 20190354775A1 · Leizerovich · 2019 [cited by examiner]
US 20190354776A1 · Ribeiro · 2019 [cited by examiner]
US 20190356885A1 · Ribeiro · 2019 [cited by examiner]
US 20200053116A1 · Soroush et al. · 2020 [cited by applicant]
US 20200117897A1 · Froloff · 2020 [cited by examiner]
US 20200293803A1 · Wajs et al. · 2020 [cited by applicant]
US 20200342324A1 · Sivaraman et al. · 2020 [cited by applicant]
US 20200372381A1 · Trim · 2020 [cited by examiner]
US 20210042638A1 · Novotny · 2021 [cited by applicant]
US 20210133456A1 · Lee · 2021 [cited by examiner]
US 20210133492A1 · Lee et al. · 2021 [cited by applicant]
US 20210295094A1 · Lee et al. · 2021 [cited by applicant]
US 20210306560A1 · Lee et al. · 2021 [cited by applicant]
US 20210397848A1 · Lee et al. · 2021 [cited by applicant]
US 20220114477A1 · Kou · 2022 [cited by examiner]
US 20220122015A1 · Bandopadhyay · 2022 [cited by examiner]
US 20220277193A1 · Wekel et al. · 2022 [cited by applicant]
US 20230267775A1 · Dhoot · 2023 [cited by examiner]
US 20230305903A1 · Lee · 2023 [cited by examiner]
US 20230325255A1 · Lee · 2023 [cited by examiner]
US 20250209328A1 · Zhuang · 2025 [cited by examiner]
US 20250308219A1 · Smith et al. · 2025 [cited by applicant]
US 20260044711A1 · Hoshina · 2026 [cited by examiner]
Girdhar, R. et al., “Video Action Transformer Network,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 244-253. [cited by applicant]
Horev, R., “Bert Explained: State of the art language model for NLP,” Nov. 10, 2018, 8 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://towardsdatascience.com/bert-explained-state-o… [cited by applicant]
Kong, Y. et al., “Human Action Recognition and Prediction: A Survey,” arXiv.1806.11230, Jul. 2, 2018, pp. 1-20. [cited by applicant]
Loginova, K., “Attention in NLP,” Jun. 22, 2018, 16 pages, [Online] [Retrieved on Jan. 20, 2021] Retrieved from the Internet <URL: https://medium.com/@edloginova/aftention-in-nlp-734c6fa9d983>. [cited by applicant]
Olah, C. et al., “Attention and Augmented Recurrent Neural Networks,” Sep. 8, 2016, 19 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://distill.pub/2016/augmented-rnns/>. [cited by applicant]
Olah, C., “Understanding LSTM Networks,” Aug. 27, 2015, 8 pages, [Online] [Retrieved on Jan. 20, 2021] Retrieved from the Internet <URL: https://colah.github.io/posts/2015-08-Understanding-LSTMs/>. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/US20/58193, dated Feb. 2, 2021, 14 pages. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/US20/58198, dated Feb. 2, 2021, 12 pages. [cited by applicant]
Piergiovanni, A., et al., “Tiny Video Networks,” arXiv:1910.06961, Oct. 15, 2019, pp. 1-10. [cited by applicant]
Rivera-Soto, R.A. et al., “Sequence to Sequence Models for Generating Video Captions,” Stanford University, Jul. 2, 2017, pp. 1-7. [cited by applicant]
Rosset, C., “Turing-NLG: A 17-billion-parameter language model by Microsoft,” Feb. 13, 2020, 11 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://www.microsoft.com/en-us/research/blo… [cited by applicant]
Security World Market, “Androvideo's AI camera for security & safety,” Nov. 14, 2019, 6 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://www.securityworldmarket.com/int/News/Product… [cited by applicant]
Sharma, A.K., “Predicting Human Behaviour Activity using Deep Leaming (LSTM),” May 26, 2018, 12 pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://medium.com/@chataks93/predicting-hum… [cited by applicant]
Sun, C. et al., “VideoBERT: A Joint Model for Video and Language Representation Learning,” arXiv:1904.01766, Sep. 11, 2019, pp. 1-13. [cited by applicant]
United States Office Action, U.S. Appl. No. 17/084,429, dated Feb. 4, 2021, 16 pages. [cited by applicant]
Vincent, J., “This Japanese AI security camera shows the future of surveillance will be automated,” Jun. 26, 2018, four pages, [Online] [Retrieved on Jan. 21, 2021] Retrieved from the Internet <URL: https://www.theverge… [cited by applicant]
United States Office Action, U.S. Appl. No. 17/345,648, dated Jan. 28, 2026, 16 pages. [cited by applicant]
United States Office Action, U.S. Appl. No. 18/203,539, dated Oct. 1, 2025, 32 pages. [cited by applicant]