IP Library › Granted Patent US 12,525,345
Granted Patent B2
US 12,525,345 · App. 18/254,801 · Granted Jan 13, 2026

Systems and methods for detection of subject activity by processing video and other signals using artificial intelligence

Inventors: George Takla (San Francisco, CA); Jayadev Hondadkatte Shivayogi (San Francisco, CA); Shaung Liu (San Francisco, CA); Bethel Orozco (San Francisco, CA)
Assignee: Dignity Health
G16H40/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,525,345
App. No.
18/254,801
Granted
Jan 13, 2026
Kind
B2
Abstract

Various embodiments of a system and associated method for detection of subject activity by processing video and other signals using artificial intelligence are shown herein. In particular, a subject monitoring system is shown that monitors subjects using a video-capable camera or other suitable video capture device to identify a subject's status of an individual by monitoring subject actions in real-time. The system further monitors other persons in the room with a subject to identify their actions and identities to ensure safety of the subject and facility while preventing confusion of the system as multiple individuals step in and out of frame over the course of the collected video feed.

Claims (54)

1 . A system, comprising:

a processor in communication with a memory, the memory including instructions, which, when executed, cause the processor to:

receive a real-time video feed including a plurality of frames that capture a body;

extract a feature set for the body from the real-time video feed including a plurality of features descriptive of the body captured within an incoming frame of the plurality of frames of the video feed;

recognize, by a neural network of the processor, an action captured over the plurality of frames, the neural network configured to interpret one or more spatial characteristics of the body between each feature set of a plurality of feature sets, each feature set corresponding with a respective frame of the plurality of frames of the video feed; and

generate an alert based on the recognized action.

2 . The system of claim 1 , wherein the memory includes instructions, which, when executed, further cause the processor to:

extract skeleton joint data for the body within a current frame of a plurality of frames of the real-time video feed; and

extract the feature set for the body based on a geometry of the skeleton joint data.

3 . The system of claim 1 , wherein the memory includes instructions, which, when executed, further cause the processor to:

extract an ocular distance for the body from a first frame of the real-time video feed; and

identify the body within a second frame of the real-time video feed based on the extracted ocular distance for the body.

4 . The system of claim 2 , wherein the memory includes instructions, which, when executed, further cause the processor to:

apply one or more additional recognition tasks for the body over the plurality of frames using the extracted skeleton joint data; and

combine results of the one or more additional recognition tasks with action recognition data indicative of the recognized action, the results of the one or more additional recognition tasks being correctly attributed to a corresponding body captured within the real-time video feed based on one or more facial characteristics of the body.

5 . The system of claim 1 , wherein the memory includes instructions, which, when executed, further cause the processor to:

recognize the action being performed by the body using a current feature set and a plurality of prior feature sets from one or more previous frames of the real-time video feed with consideration for a variation in elapsed time between frames of the real-time video feed.

6 . The system of claim 5 , wherein the neural network is a recurrent neural network including a plurality of long short term memory (LSTM) units, each LSTM unit configured to store a current memory of a feature set of the plurality of feature sets and directly adjust the current memory stored therein with respect to an elapsed time relative to other feature sets of the plurality of feature sets associated with the real-time video feed.

7 . The system of claim 1 , wherein the memory includes instructions, which, when executed, further cause the processor to:

extract one or more facial characteristics of the body for correct attribution of the body across multiple frames of the real-time video feed.

8 . The system of claim 7 , wherein the one or more extracted facial characteristics include an ocular distance that defines a rectangle area around a face of the body captured within the real-time video feed to match the face from a first frame of the plurality of frames to a second frame of the plurality of frames of the real-time video feed.

9 . The system of claim 1 , further comprising:

a display in communication with the processor configured to display the alert based on the recognized action.

10 . The system of claim 1 , further comprising:

a mitigation module in association with the processor and configured to broadcast the alert to one or more additional devices.

11 . A method, comprising:

receiving, at a processor, a real-time video feed including a plurality of frames that capture a body;

extracting a feature set for the body from the real-time video feed including a plurality of features descriptive of the body captured within an incoming frame of the video feed;

recognizing, by a neural network of the processor, an action captured over the plurality of frames, the neural network configured to interpret one or more spatial characteristics of the body between each feature set of a plurality of feature sets; and

generating an alert based on the recognized action.

12 . The method of claim 11 , further comprising:

adjusting a memory stored within a long short term memory unit of the neural network with respect to an elapsed time relative to other feature sets of the plurality of feature sets associated with the real-time video feed.

13 . The method of claim 11 , further comprising:

broadcasting the alert to one or more additional devices in communication with the processor.

14 . The method of claim 11 , further comprising:

applying one or more additional recognition tasks for the body over the plurality of frames using extracted skeleton joint data; and

combining results of the one or more additional recognition tasks with action recognition data indicative of the recognized action, the results of the one or more additional recognition tasks being correctly attributed to a corresponding body captured within the real-time video feed based on one or more facial characteristics of the body.

15 . The method of claim 14 , further comprising:

generating the alert based on the results of the one or more additional recognition tasks.

16 . The method of claim 11 , further comprising:

extracting skeleton joint data for the body within a current frame of the plurality of frames of the real-time video feed; and

extracting the feature set for the body based on the skeleton joint data.

17 . A system, comprising:

a processor in communication with a memory, the memory including instructions, which, when executed, cause the processor to:

receive a real-time video feed including a plurality of frames that capture a body;

extract a feature set for the body from the real-time video feed including a plurality of features descriptive of the body captured within an incoming frame of the video feed;

recognize, by a neural network of the processor, an action captured over the plurality of frames, the neural network configured to interpret one or more spatial characteristics of the body between each feature set of a plurality of feature sets captured across the plurality of frames with respect to an elapsed time between each frame of the plurality of frames; and

generate an alert based on the recognized action.

18 . The system of claim 17 , wherein the neural network is a recurrent neural network including a plurality of long short term memory (LSTM) units, wherein each LSTM unit stores a current memory of a feature set of the plurality of feature sets.

19 . The system of claim 18 , wherein each LSTM unit of the plurality of LSTM units is configured to directly adjust the current memory stored therein with respect to an elapsed time relative to other feature sets of the plurality of feature sets associated with the real-time video feed.

20 . The system of claim 18 , wherein a feature set of the plurality of feature sets corresponds to a respective frame of the plurality of frames of the real-time video feed.

21 . The system of claim 18 , wherein the memory includes instructions, which, when executed, further cause the processor to:

extract skeleton joint data for the body within a current frame of the plurality of frames of the real-time video feed; and

extract the feature set for the body based on the skeleton joint data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: TAKLA, GEORGE; SHIVAYOGI, JAYADEV HONDADKATTE; LIU, SHAUNG; OROZCO, BETHEL
To: DIGNITY HEALTH
Reel/Frame 065150/0232 →
Continuity (2)
Provisional Application 63121489 · Dec 4, 2020
Related Publication 20240029877A1 · Jan 25, 2024
References Cited (17)
US 7541935B2 · Dring et al. · 2009 [cited by applicant]
US 8532737B2 · Cervantes · 2013 [cited by applicant]
US 11100633B2 · Ngo Dinh · 2021 [cited by examiner]
US 20180129873A1 · Alghazzawi et al. · 2018 [cited by applicant]
US 20210267491A1 · Guibene · 2021 [cited by examiner]
US 20240021019A1 · Yu · 2024 [cited by examiner]
European Patent Office, Extended European Patent Office, Application No. 21901610.2,Oct. 1, 2024, 10 pages. [cited by applicant]
Ahmedt-Aristizabal, D. et al., “Deep Motion Analysis for Epileptic Seizure Classification,” International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Jul. 18, 2018, pp. 3578-3581. [cited by applicant]
Amiribesheli, M. et al., “A review of smart homes in healthcare,” Journal of Ambient Intelligence and Hunanized Computing, vol. 6, No. 4, Mar. 14, 2015, pp. 495-517. [cited by applicant]
Islam, M. M. et al., “Deep Learning Based Systems Developed for Fall Detection: A Review,” IEEE Access, vol. 8, Sep. 4, 2020, 21 pages. [cited by applicant]
Popescu, A. et al., “Fusion Mechanisms for Human Activity Recognition Using Automated Machine Learning,” IEEE Access, vol. 8, Jul. 30. 2020, pp. 143996-144014. [cited by applicant]
Rodriguez, P. et al., “Deep Pain: Exploiting Long Short-Term Memory Networks for Facial Expression Classification,” IEEE Transactions on Cybernetics, vol. 52, No. 5, May 2022, pp. 3314-3324. [cited by applicant]
Yao, L. et al., “A fall detection method based on a joint motion map using double convolutional neural networks,” Multimedia Tools and Applications, vol. 81, No. 4, Jun. 22, 2020, pp. 4551-4568. [cited by applicant]
Patent Cooperation Treaty, International Search Report and Written Opinion, International Application No. PCT/US2021/062024, date of mailing Mar. 15, 2022, 14 pages. [cited by applicant]
Yu, M. et al., “Computer Vision Based Fall Detection by a Convolutional Neural Network,” Proceedings of the 19th ACM International Conference on Multimodal Interaction (ICMI '17), Nov. 13-17, 2017, pp. 416-420. [cited by applicant]
Gorodnichy, D. et al., “Target-based evaluation of face recognition technology for video surveillance applications,” 2014 IEEE Symposium on Computational Intelligence in Biometrics and Identity Management (CIBIM) Dec. 9… [cited by applicant]
Augustyniak, P. et al., “Graph-based representation of behavior in detection and prediction of daily living activities,” Computers in Biology and Medicine, vol. 95, Apr. 1, 2018, pp. 261-270. [cited by applicant]