IP Library › Granted Patent US 11,937,929
Granted Patent B2
US 11,937,929 · App. 17/397,675 · Granted Mar 26, 2024

Systems and methods for using mobile and wearable video capture and feedback plat-forms for therapy of mental disorders

Inventors: Catalin Voss (Stanford, CA); Nicholas Joseph Haber (Palo Alto, CA); Dennis Paul Wall (Palo Alto, CA); Aaron Scott Kline (Saratoga, CA); Terry Allen Winograd (Stanford, CA)
Assignee: The Board of Trustees of the Leland Stanford Junior University
A61B5/165A61B5/0002A61B5/0036A61B5/0205A61B5/1176A61B5/4836A61B5/6803A61B5/681G06V10/255G06V10/40G06V10/764G06V10/945G16H20/70G16H30/40G16H40/63G16H50/20G16H50/30A61B5/1114A61B5/1126A61B5/1128A61B5/7405A61B5/742A61B5/7455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,937,929
App. No.
17/397,675
Granted
Mar 26, 2024
Kind
B2
Abstract

Behavioral and mental health therapy systems in accordance with several embodiments of the invention include a wearable camera and/or a variety of sensors (accelerometer, microphone, among various other) connected to a computing system including a display, audio output, holographic output, and/or vibrotactile output to automatically recognize social cues from images captured by at least one camera and provide this information to the wearer via one or more outputs such as (but not limited to) displaying an image, displaying a holographic overlay, generating an audible signal, and/or generating a vibration.

Claims (30)

1. An image processing system, comprising:

at least one camera for capturing images of a surrounding environment; and

at least one processor and memory containing software;

wherein the software directs the at least one processor to:

obtain data comprising a sequence of images captured by the at least one camera;

detect a face for at least one person within a plurality of images in the sequence of images, wherein the at least one person is talking in at least one image of the plurality of images;

detect at least one emotional cue in the face based upon the plurality of images using a classifier that is trained using statistically representative social expression data that comprises image data of expressive talking sequences;

identify at least one emotion based on the at least one emotional cue; and

display at least one emotion indicator label in real time to provide therapeutic feedback to a user.

2. The image processing system of claim 1 , wherein the system comprises a wearable video capture system comprising at least one outward facing camera.

3. The image processing system of claim 2 , wherein the wearable video capture system is selected from the group consisting of a virtual reality headset, a mixed-reality headset, an augmented reality headset, and glasses comprising a heads-up display.

4. The image processing system of claim 2 , wherein the wearable video capture system communicates with at least one mobile device, wherein the at least one processor is executing on the at least one mobile device.

5. The image processing system of claim 1 , wherein the software directs the at least one processor to obtain supplementary data comprising data captured from at least one sensor selected from the group consisting of a microphone, an accelerometer, a gyroscope, an eye tracking sensor, a head-tracking sensor, a body temperature sensor, a heart rate sensor, a blood pressure sensor, and a skin conductivity sensor.

6. The image processing system of claim 1 , wherein displaying at least one emotion indicator label in real time to provide therapeutic feedback further comprises performing at least one of displaying a label within a heads-up display, generating an audible signal, generating a vibration, displaying a holographic overlay, and displaying an image.

7. The image processing system of claim 1 , wherein the software directs the at least one processor to process image data at a higher resolution within a region of interest related to a detected face within an image.

8. The image processing system of claim 7 , wherein the region of interest is a bounding region around the detected face, wherein processing the image data further comprises using a moving average filter to smoothen the bounding region of interest.

9. The image processing system of claim 1 , wherein the at least one emotional cue comprises information selected from the group consisting of facial expressions, facial muscle movements, body language, gestures, body pose, eye contact events, head pose, features of a conversation, fidgeting, and anxiety information.

10. The image processing system of claim 1 , wherein the classifier provides event-based social cues.

11. The image processing system of claim 1 , wherein the software directs the at least one processor to:

prompt a user to label data for a target individual with at least one emotional cue label; and

store the user-labeled data for the target individual in memory.

12. The image processing system of claim 1 , wherein the software directs the at least one processor to store social interaction data and provide a user interface for review of the social interaction data.

13. The image processing system of claim 1 , wherein the classifier is a regression machine that provides continuous social cues.

14. The image processing system of claim 1 , wherein the classifier is trained as visual time-dependent classifiers using video data of standard facial expressions and with expressive talking sequences.

15. The image processing system of claim 1 , wherein the software directs the at least one processor to detect gaze events using at least one inward-facing eye tracking data in conjunction with outward-facing video data.

16. The image processing system of claim 1 , wherein the software directs the at least one processor to provide a review of activities recorded and provide user behavioral data generated as a reaction to the recorded activities.

17. The image processing system of claim 1 , wherein the system comprises a smart phone, desktop computer, laptop computer, or tablet computer.

18. The image processing system of claim 1 , wherein the software further directs the at least one processor to perform neutral feature estimation and subtraction on the face that is detected in each of the plurality of images.

19. The image processing system of claim 1 , wherein the user has autism.

20. The image processing system of claim 1 , wherein said software further directs the at least one processor to weight data related to the at least one emotional cue of the at least one person.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2023
From: HABER, NICHOLAS JOSEPH; WALL, DENNIS PAUL; KLINE, AARON SCOTT; WINOGRAD, TERRY ALLEN
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
Reel/Frame 063938/0645 →
Continuity (4)
Continuation 17066979 · Oct 9, 2020
Continuation 15589877 · May 8, 2017
Provisional Application 62333108 · May 6, 2016
Related Publication 20220202330A1 · Jun 30, 2022