IP Library Granted Patent US 11,632,258
Granted Patent B1
US 11,632,258 · App. 17/226,315 · Granted Apr 18, 2023

Recognizing and mitigating displays of unacceptable and unhealthy behavior by participants of online video meetings

Inventor: Phil Libin (San Francisco, CA)
Assignee: All Turtles Corporation
H04L12/1831G06F1/163G06F3/165G06F18/214G06N20/00G06V20/46G06V40/107G06V40/172G06V40/28G08B3/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,632,258
App. No.
17/226,315
Granted
Apr 18, 2023
Kind
B1
Abstract

Handling unacceptable behavior by a participant in a video conference includes detecting the unacceptable behavior by the participant by applying machine learning to data about the participant received from one or more capturing devices and by using a predetermined list of bad habits, determining recognition accuracy for the unacceptable behavior, and providing a response to the unacceptable behavior that varies according to the recognition accuracy. The machine learning may include an initial training phase that, prior to deployment, is used to obtain a general recognition capability for each item on the predetermined list of bad habits. The one or more capturing devices may include a laptop with a camera and a microphone, a mobile device, autonomous cameras, add-on cameras, headsets, regular speakers, smart watches, wristbands, smart rings, and wearable sensors, smart eyewear, heads-up displays, headbands, and/or smart footwear.

Claims (36)

1. A method of handling unacceptable behavior by a participant in a video conference, comprising:

detecting the unacceptable behavior by the participant by applying machine learning to data about the participant received from one or more capturing devices and by using a predetermined list of bad habits;

determining recognition accuracy for the unacceptable behavior; and

providing a response to the unacceptable behavior that varies according to the recognition accuracy, wherein the recognition accuracy is estimated to be lower at an early stage of the system than at later stages of the system and wherein the response to the unacceptable behavior is an alert to the user for each detected episode in response to the recognition accuracy being estimated to be lower and wherein the participant is asked to provide confirmation for each detected episode.

2. A method, according to claim 1 , wherein the machine learning includes an initial training phase that, prior to deployment, is used to obtain a general recognition capability for each item on the predetermined list of bad habits.

3. A method, according to claim 1 , wherein the one or more capturing devices include a laptop with a camera and a microphone, a mobile device, autonomous cameras, add-on cameras, headsets, regular speakers, smart watches, wristbands, smart rings, and wearable sensors, smart eyewear, heads-up displays, headbands, and smart footwear.

4. A method, according to claim 1 , wherein the data about the participant includes at least one of: visual data, sound data, motion, proximity and chemical sensor data, heart rate, breathing rate, and blood pressure.

5. A method, according to claim 1 , wherein the list of bad habits includes nail biting, yelling, making unacceptable gestures, digging one's nose, yawning, blowing one's nose, combing one's hair, slouching, and looking away from a screen being used for the video conference.

6. A method, according to claim 5 , wherein technologies used to detect bad habits include facial recognition, sound recognition, gesture recognition, and hand movement recognition.

7. A method, according to claim 6 , wherein yawning is detected using a combination of the facial recognition technology, the sound recognition technology, and the gesture recognition technology.

8. A method, according to claim 6 , wherein face touching is detected using the facial recognition technology and hand movement recognition technology.

9. A method, according to claim 6 , wherein digging one's nose is detected using the facial recognition technology and hand movement recognition technology.

10. A method, according to claim 1 , wherein the confirmation provided by the participant is used to improve the recognition accuracy of the machine learning.

11. A method, according to claim 1 , wherein the confirmation provided by the participant is used to improve recognition speed of the machine learning.

12. A method, according to claim 1 , wherein the confirmation provided by the participant is used to provide early recognition of displays of the bad habits detected by the machine learning.

13. A method, according to claim 12 , wherein the early recognition is used to predict when a participant will exhibit unacceptable behavior.

14. A method, according to claim 13 , wherein the response is provided when the participant is predicted to exhibit unacceptable behavior.

15. A method, according to claim 13 , wherein the response includes at least one of: cutting off audio input of the user or cutting off video input of the user.

16. A method, according to claim 12 , wherein the early recognition is based on at least one of: observed behavioral norms of the participant and planned activities of the participant.

17. A method, according to claim 1 , wherein the participant is provided with only a warning if the participant is not providing audio input and/or video input to the video conference.

18. A method of preventing a user from touching a face of the user, comprising:

obtaining video frames of the user including the face of the user;

applying facial recognition technology to the video frames to detect locations of particular portions of the face;

detecting a position, shape, and trajectory of a moving hand of the user in the video frames;

predicting a final position and shape of the hand based on the position, shape, and trajectory of the hand and on the locations of the specific portions of the face; and

providing an alarm to the user in response to predicting a final position of the hand will be touching the face of the user, wherein the alarm varies according to a predicted final shape of the hand and according to predicting that a final position of the hand will be touching a specific one of the particular portions of the face and wherein the predicted final shape of the hand is one of: an open palm and open fingers and wherein a sound for the alarm provided in response to the predicted final shape being an open palm is less severe than a sound for the alarm provided in response to the predicted final shape being open fingers and the predicted final position being one of: a mouth, nose, eyes, and ears of the user.

19. A method, according to claim 18 , wherein the user is a participant in a video conference.

20. A method, according to claim 18 , wherein predicting a final position of the hand includes determining if the hand crosses an alert zone that is proximal to the face.

21. A method of preventing a user from touching a face of the user, comprising:

obtaining video frames of the user including the face of the user;

applying facial recognition technology to the video frames to detect locations of particular portions of the face;

detecting a position, shape, and trajectory of a moving hand of the user in the video frames;

predicting a final position and shape of the hand based on the position, shape, and trajectory of the hand and on the locations of the specific portions of the face; and

providing an alarm to the user in response to predicting a final position of the hand will be touching the face of the user, wherein the alarm varies according to a predicted final shape of the hand and according to predicting that a final position of the hand will be touching a specific one of the particular portions of the face and wherein the predicted final shape of the hand is one of: an open palm and open fingers and wherein a sound for the alarm becomes more severe as the predicted final position and the predicted final shape changes to the predicted final shape being open fingers and the predicted final position being one of: a mouth, nose, eyes, and ears of the user.

22. A method, according to claim 21 , wherein the user is a participant in a video conference.

23. A method, according to claim 21 , wherein predicting a final position of the hand includes determining if the hand crosses an alert zone that is proximal to the face.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2023
From: ALL TURTLES CORPORATION
To: MMHMM INC.
Reel/Frame 063385/0180 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2021
From: LIBIN, PHIL
To: ALL TURTLES CORPORATION
Reel/Frame 055937/0988 →
Continuity (1)
Provisional Application 63008769 · Apr 12, 2020
Cited By (5)
US 12,265,648 US 12,282,633 US 12,309,213 US 12,626,581 US 12,681,566