IP Library Granted Patent US 11,652,956
Granted Patent B2
US 11,652,956 · App. 17/172,552 · Granted May 16, 2023

Emotion recognition in video conferencing

Inventors: Victor Shaburov (Castro Valley, CA); Yurii Monastyrshyn (Santa Monica, CA)
Assignee: SNAP INC.
G06V40/176G06Q30/0281G06T7/337G06T7/344G06V10/7553G06V20/64G06V40/165G06V40/167G06V40/171G10L25/63H04N7/147H04N7/15G06T2207/10016G06T2207/30201G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,652,956
App. No.
17/172,552
Granted
May 16, 2023
Kind
B2
Abstract

Methods and systems for videoconferencing include recognition of emotions related to one videoconference participant such as a customer. This ultimately enables another videoconference participant, such as a service provider or supervisor, to handle angry, annoyed, or distressed customers. One example method includes the steps of receiving a video that includes a sequence of images, detecting at least one object of interest (e.g., a face), locating feature reference points of the at least one object of interest, aligning a virtual face mesh to the at least one object of interest based on the feature reference points, finding over the sequence of images at least one deformation of the virtual face mesh that reflect face mimics, determining that the at least one deformation refers to a facial emotion selected from a plurality of reference facial emotions, and generating a communication bearing data associated with the facial emotion.

Claims (60)

1. A method comprising:

receiving, by one or more processors, information about a conference between a plurality of participants including a first user and a second user;

determining an emotion of the first user based on a combination of one or more facial features of the first user and a speech emotion of the first user; and

generating a communication bearing data associated with the combination of the one or more facial features of the first user and the speech emotion of the first user.

2. The method of claim 1 , further comprising:

determining that the emotion of the first user is a negative emotion.

3. The method of claim 2 , further comprising:

in response to determining that the emotion is the negative emotion, transmitting the communication to a non-participant of the conference, the communication comprising data associated with the negative emotion.

4. The method of claim 1 , wherein the received information about the conference comprises a video including a sequence of images corresponding to a videoconference, further comprising:

detecting at least one object of interest in one or more of the images;

locating feature reference points of the at least one object of interest; and

determining that at least one deformation between two or more of the feature reference points refers to a facial emotion selected from a plurality of reference facial emotions.

5. The method of claim 4 , wherein determining that the at least one deformation refers to a facial emotion further comprises:

comparing the deformation between two or more feature reference points to reference facial parameters of the plurality of facial emotions; and

selecting the facial emotion based on the comparison of the deformation between two or more feature reference points to the reference facial parameters.

6. The method of claim 1 , wherein the first user is a service provider, further comprising:

establishing a videoconference between the service provider and the second user; and

transmitting the communication over a communications network to a third party.

7. The method of claim 1 , further comprising:

detecting one or more gestures; and

determining the one or more gestures are associated with a negative emotion.

8. The method of claim 1 , wherein the information includes an audio stream corresponding to the conference.

9. The method of claim 8 , further comprising:

extracting at least one voice feature from the audio stream;

comparing the extracted at least one voice feature to a plurality of reference voice features; and

detecting that the emotion of the first user is a negative emotion based on the comparison.

10. A system comprising:

one or more processors; and

a non-transitory processor-readable medium coupled to the one or more processors, the non-transitory processor-readable medium comprising processor-executable instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving information about a conference between a plurality of participants including a first user and a second user;

determining an emotion of the first user based on a combination of one or more facial features of the first user and a speech emotion of the first user; and

generating a communication bearing data associated with the combination of the one or more facial features of the first user and the speech emotion of the first user.

11. The system of claim 10 , wherein the operations further comprise:

determining that the emotion of the first user is a negative emotion.

12. The system of claim 11 , wherein the operations further comprise:

in response to determining that the emotion is the negative emotion, transmitting the communication to a non-participant of the conference, the communication comprising data associated with the negative emotion.

13. The system of claim 10 , wherein the received information about the conference comprises a video including a sequence of images corresponding to a videoconference, further comprising operations for:

detecting at least one object of interest in one or more of the images;

locating feature reference points of the at least one object of interest; and

determining that at least one deformation between two or more of the feature reference points refers to a facial emotion selected from a plurality of reference facial emotions.

14. The system of claim 13 , wherein determining that the at least one deformation refers to a facial emotion further comprises:

comparing the deformation between two or more feature reference points to reference facial parameters of the plurality of facial emotions; and

selecting the facial emotion based on the comparison of the deformation between two or more feature reference points to the reference facial parameters.

15. The system of claim 10 , wherein the first user is a service provider, further comprising operations for:

establishing a videoconference between the service provider and the second user; and

transmitting the communication over a communications network to a third party.

16. The system of claim 10 , further comprising operations for:

detecting one or more gestures; and

determining the one or more gestures are associated with a negative emotion.

17. The system of claim 10 , wherein the information includes an audio stream corresponding to the conference.

18. The system of claim 17 , further comprising operations for:

extracting at least one voice feature from the audio stream;

comparing the extracted at least one voice feature to a plurality of reference voice features; and

detecting that the emotion of the first user is a negative emotion based on the comparison.

19. A non-transitory processor-readable medium comprising processor-executable instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving information about a conference between a plurality of participants including a first user and a second user;

determining an emotion of the first user based on a combination of one or more facial features of the first user and a speech emotion of the first user; and

generating a communication bearing data associated with the combination of the one or more facial features of the first user and the speech emotion of the first user.

20. The non-transitory processor-readable medium of claim 19 , wherein the operations further comprise:

determining that the emotion of the first user is a negative emotion.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2023
From: SHABUROV, VICTOR; MONASTYRSHYN, YURII
To: LOOKSERY, INC.
Reel/Frame 063246/0076 →
MERGER Recorded Apr 6, 2023
From: LOOKSERY, INC.
To: AVATAR ACQUISITION CORP
Reel/Frame 063246/0297 →
MERGER Recorded Apr 6, 2023
From: AVATAR ACQUISITION CORP
To: AVATAR MERGER SUB II, LLC
Reel/Frame 063246/0380 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2023
From: AVATAR MERGER SUB II, LLC
To: SNAP INC.
Reel/Frame 063246/0493 →
Continuity (5)
Continuation 16260813 · Jan 29, 2019
Continuation 15816776 · Nov 17, 2017
Continuation 15430133 · Feb 10, 2017
Continuation 14661539 · Mar 18, 2015
Related Publication 20210192193A1 · Jun 24, 2021