IP Library Granted Patent US 10,949,655
Granted Patent B2
US 10,949,655 · App. 16/260,813 · Granted Mar 16, 2021

Emotion recognition in video conferencing

Inventors: Victor Shaburov (Castro Valley, CA); Yurii Monastyrshyn (Santa Monica, CA)
Assignee: Snap Inc.
G06K9/00315G06K9/00201G06K9/00248G06K9/00261G06K9/00281G06K9/6209G06Q30/0281G06T7/337G06T7/344G10L25/63H04N7/147H04N7/15G06T2207/10016G06T2207/30201G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,949,655
App. No.
16/260,813
Granted
Mar 16, 2021
Kind
B2
Abstract

Methods and systems for videoconferencing include recognition of emotions related to one videoconference participant such as a customer. This ultimately enables another videoconference participant, such as a service provider or supervisor, to handle angry annoyed, or distressed customers. One example method includes the steps of receiving a video that includes a sequence of images, detecting at least one object of interest (e.g., a face), locating feature reference points of the at least one object of interest, aligning a virtual face mesh to the at least one object of interest based on the feature reference points, finding over the sequence of images at least one deformation of the virtual face mesh that reflect face mimics, determining that the at least one deformation refers to a facial emotion selected from a plurality of reference facial emotions, and generating a communication bearing data associated with the facial emotion.

Claims (66)

1. A computer-implemented method comprising:

receiving, by one or more processors, information about a conference between a plurality of participants including a first user and a second user;

in an automated operation, based on the received information, determining an emotion of the first user;

determining that the emotion of the first user is a negative emotion; and

in response to determining that the emotion is the negative emotion, generating a communication for transmission to a non-participant of the conference, the communication comprising data associated with the negative emotion.

2. The method of claim 1 , wherein the received information about the conference comprises a video including a sequence of images corresponding to a videoconference, further comprising:

detecting at least one object of interest in one or more of the images;

locating feature reference points of the at least one object of interest; and

determining that at least one deformation between two or more of the feature reference points refers to a facial emotion selected from a plurality of reference facial emotions, wherein the negative emotion is determined from the selected facial emotion.

3. The method of claim 2 , further comprising:

aligning a virtual face mesh to the at least one object of interest based at least in part on the feature reference points; and

finding at least one facial deformation of the virtual face mesh associated with at least one facial mimic.

4. The method of claim 2 , wherein determining that the at least one deformation refers to a facial emotion further comprises:

comparing the deformation between two or more feature reference points to reference facial parameters of the plurality of facial emotions; and

selecting the facial emotion based on the comparison of the deformation between two or more feature reference points to the reference facial parameters.

5. The method of claim 1 , wherein the first user is a service provider and the non-participant is a third party, further comprising:

establishing a videoconference between the service provider and the second user associated with the negative emotion; and

transmitting the communication over a communications network to the third party.

6. The method of claim 1 , wherein the first user is a service provider and the non-participant is a third party, further comprising:

establishing a videoconference between the service provider, the second user, and the third party, the videoconference including the third party being established responsive to determining the negative emotion.

7. The method of claim 1 , further comprising:

detecting one or more gestures; and

determining the one or more gestures are associated with the negative emotion.

8. The method of claim 1 , wherein the information includes an audio stream corresponding to the conference.

9. The method of claim 8 further comprising:

extracting at least one voice feature from the audio stream;

comparing the extracted at least one voice feature to a plurality of reference voice features; and

detecting that the emotion of the first user is the negative emotion based on the comparison.

10. The method of claim 1 , wherein detecting that the emotion of the first user is the negative emotion based on a combination of detecting a facial emotion and a speech emotion of the first user.

11. A system comprising:

one or more processors; and

a non-transitory processor-readable medium coupled to the one or more processors, the non-transitory processor-readable medium comprising processor-executable instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving information about a conference between a plurality of participants including a first user and a second user;

in an automated operation, based on the received information, determining an emotion of the first user;

determining that the emotion of the first user is a negative emotion; and

in response to determining that the emotion is the negative emotion, generating a communication for transmission to a non-participant of the conference, the communication comprising data associated with the negative emotion.

12. The system of claim 11 , wherein the received information about the conference comprises a video including a sequence of images corresponding to a videoconference, and wherein the operations further comprise:

detecting at least one object of interest in one or more of the images;

locating feature reference points of the at least one object of interest; and

determining that at least one deformation between two or more of the feature reference points refers to a facial emotion selected from a plurality of reference facial emotions, wherein the negative emotion is determined from the selected facial emotion.

13. The system of claim 12 , wherein the operations further comprise:

aligning a virtual face mesh to the at least one object of interest based at least in part on the feature reference points; and

finding at least one facial deformation of the virtual face mesh associated with at least one facial mimic.

14. The system of claim 12 , wherein the operations further comprise:

comparing the deformation between two or more feature reference points to reference facial parameters of the plurality of facial emotions; and

selecting the facial emotion based on the comparison of the deformation between two or more feature reference points to the reference facial parameters.

15. The system of claim 11 , wherein the first user is a service provider and the non-participant is a third party, and wherein the operations further comprise:

establishing a videoconference between the service provider and the second user associated with the negative emotion; and

transmitting the communication over communications network to the third party.

16. The system of claim 11 , wherein the first user is a service provider and the non-participant is a third party, and wherein the operations further comprise:

establishing a videoconference between the service provider, the second user, and the third party, the videoconference including the third party being established responsive to determining the negative emotion.

17. The system of claim 11 , wherein the operations further comprise:

detecting one or more gestures; and

determining the one or more gestures are associated with the negative emotion.

18. A non-transitory processor-readable medium comprising processor-executable instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving information about a conference between a plurality of participants including a first user and a second user;

in an automated operation, based on the received information, determining an emotion of the first user;

determining that the emotion of the first user is a negative emotion; and

in response to determining that the emotion is the negative emotion, generating a communication for transmission to a non-participant of the conference, the communication comprising data associated with the negative emotion.

19. The non-transitory processor-readable medium of claim 18 , wherein the received information about the conference comprises a video including a sequence of images corresponding to a videoconference, and wherein the operations further comprise:

detecting at least one object of interest in one or more of the images;

locating feature reference points of the at least one object of interest; and

determining that at least one deformation between two or more of the feature reference points refers to a facial emotion selected from a plurality of reference facial emotions, wherein the negative emotion is determined from the selected facial emotion.

20. The non-transitory processor-readable medium of claim 19 , wherein the operations further comprise:

aligning a virtual face mesh to the at least one object of interest based at least in part on the feature reference points; and

finding at least one facial deformation of the virtual face mesh associated with at least one facial mimic.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2021
From: SHABUROV, VICTOR; MONASTYRSHYN, YURII
To: LOOKSERY INC.
Reel/Frame 055216/0486 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2021
From: LOOKSERY, INC.
To: AVATAR ACQUISITION CORP.
Reel/Frame 055216/0616 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2021
From: AVATAR ACQUISITION CORP.
To: AVATAR MERGER SUB II, LLC
Reel/Frame 055216/0699 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2021
From: AVATAR MERGER SUB II, LLC
To: SNAP INC.
Reel/Frame 055216/0728 →
Continuity (4)
Continuation 15816776 · Nov 17, 2017
Continuation 15430133 · Feb 10, 2017
Continuation 14661539 · Mar 18, 2015
Related Publication 20190156112A1 · May 23, 2019
Cited By (3)
US 12,216,519 US 12,645,280 US 12,711,451