IP Library Granted Patent US 9,852,328
Granted Patent B2
US 9,852,328 · App. 15/430,133 · Granted Dec 26, 2017

Emotion recognition in video conferencing

Inventors: Victor Shaburov (Castro Valley, CA); Yurii Monastyrshyn (Odessa, UA)
Assignee: SNAP INC.
G06K9/00315G06K9/00201G06K9/00248G06K9/00261G06K9/00281G06K9/6209G06Q30/0281G06T7/337G06T7/344G10L25/63H04N7/147H04N7/15G06T2207/10016G06T2207/30201G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,852,328
App. No.
15/430,133
Granted
Dec 26, 2017
Kind
B2
Abstract

Methods and systems for videoconferencing include recognition of emotions related to one videoconference participant such as a customer. This ultimately enables another videoconference participant, such as a service provider or supervisor, to handle angry, annoyed, or distressed customers. One example method includes the steps of receiving a video that includes a sequence of images, detecting at least one object of interest (e.g., a face), locating feature reference points of the at least one object of interest, aligning a virtual face mesh to the at least one object of interest based on the feature reference points, finding over the sequence of images at least one deformation of the virtual face mesh that reflect face mimics, determining that the at least one deformation refers to a facial emotion selected from a plurality of reference facial emotions, and generating a communication bearing data associated with the facial emotion.

Claims (70)

1. A computer-implemented method for video conferencing, the method comprising:

receiving a video including a sequence of images and an audio stream;

detecting at least one object of interest in one or more of the images;

locating feature reference points of the at least one object of interest;

aligning a virtual face mesh to the at least one object of interest in one or more of the images based at least in part on the feature reference points;

finding over the sequence of images at least one deformation of the virtual face mesh, wherein the at least one deformation is associated with at least one face mimic;

determining that the at least one deformation refers to a facial emotion selected from a plurality of reference facial emotions;

recognizing a speech emotion in the audio stream of the at least one object of interest; and

generating a communication bearing data associated with one or more of the facial emotion and the speech emotion.

2. The computer-implemented method of claim 1 , wherein recognizing the speech emotion comprises:

extracting at least one voice feature from the audio stream;

comparing the extracted at least one voice feature to a plurality of reference voice features; and

selecting the speech emotion based on the comparison of the extracted at least one voice feature to the plurality of reference voice features.

3. The computer-implemented method of claim 1 , wherein the object of interest is a first user and the video stream comprises speech from the first user and a second user, and the method further comprises:

recognizing a speech emotion of the second user; and

recognizing the speech emotion of the first user based on the speech of the first user and the speech emotion of the second user.

4. The computer-implemented method of claim 1 , wherein recognizing the speech emotion comprises recognizing a speech in the audio stream.

5. The computer-implemented method of claim 1 , wherein the communication bearing data associated with the facial emotion further includes data associated with the speech emotion.

6. The computer-implemented method of claim 1 further comprising:

combining the facial emotion and the speech emotion to generate an emotional status of an individual associated with the at least one object of interest.

7. The computer-implemented method of claim 1 , wherein recognizing the speech emotion further comprises:

identifying one or more keywords within the audio stream of the at least one object of interest;

determining at least one keyword of the one or more keywords is associated with a negative emotion; and

selecting a negative emotion as the speech emotion of the at least one object of interest.

8. A system, comprising:

one or more processors; and

a non-transitory processor-readable medium coupled to the one or more processors, the non-transitory processor-readable medium comprising processor-executable instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving a video including a sequence of images and an audio stream;

detecting at least one object of interest in one or more of the images;

locating feature reference points of the at least one object of interest;

aligning a virtual face mesh to the at least one object of interest in one or more of the images based at least in part on the feature reference points;

finding over the sequence of images at least one deformation of the virtual face mesh, wherein the at least one deformation is associated with at least one face mimic;

determining that the at least one deformation refers to a facial emotion selected from a plurality of reference facial emotions;

recognizing a speech emotion in the audio stream of the at least one object of interest; and

generating a communication bearing data associated with one or more of the facial emotion and the speech emotion.

9. The system of claim 8 , wherein recognizing the speech emotion comprises:

extracting at least one voice feature from the audio stream;

comparing the extracted at least one voice feature to a plurality of reference voice features; and

selecting the speech emotion based on the comparison of the extracted at least one voice feature to the plurality of reference voice features.

10. The system of claim 8 , wherein the object of interest is a first user and the video stream comprises speech from the first user and a second user, and the operations further comprise:

recognizing a speech emotion of the second user; and

recognizing the speech emotion of the first user based on the speech of the first user and the speech emotion of the second user.

11. The system of claim 8 , wherein recognizing the speech emotion comprises recognizing a speech in the audio stream.

12. The system of claim 8 , wherein the communication bearing data associated with the facial emotion further includes data associated with the speech emotion.

13. The system of claim 8 , wherein the operations further comprise:

combining the facial emotion and the speech emotion to generate an emotional status of an individual associated with the at least one object of interest.

14. The system of claim 8 , wherein recognizing the speech emotion further comprises:

identifying one or more keywords within the audio stream of the at least one object of interest;

determining at least one keyword of the one or more keywords is associated with a negative emotion; and

selecting a negative emotion as the speech emotion of the at least one object of interest.

15. A non-transitory processor-readable medium comprising processor-executable instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving a video including a sequence of images and an audio stream;

detecting at least one object of interest in one or more of the images;

locating feature reference points of the at least one object of interest;

aligning a virtual face mesh to the at least one object of interest in one or more of the images based at least in part on the feature reference points;

finding over the sequence of images at least one deformation of the virtual face mesh, wherein the at least one deformation is associated with at least one face mimic;

determining that the at least one deformation refers to a facial emotion selected from a plurality of reference facial emotions;

recognizing a speech emotion in the audio stream of the at least one object of interest; and

generating a communication bearing data associated with one or more of the facial emotion and the speech emotion.

16. The non-transitory processor-readable medium of claim 15 , wherein recognizing the speech emotion comprises:

extracting at least one voice feature from the audio stream;

comparing the extracted at least one voice feature to a plurality of reference voice features; and

selecting the speech emotion based on the comparison of the extracted at least one voice feature to the plurality of reference voice features.

17. The non-transitory processor-readable medium of claim 15 , wherein the object of interest is a first user and the video stream comprises speech from the first user and a second user, and the operations further comprise:

recognizing a speech emotion of the second user; and

recognizing the speech emotion of the first user based on the speech of the first user and the speech emotion of the second user.

18. The non-transitory processor-readable medium of claim 15 , wherein recognizing the speech emotion comprises recognizing a speech in the audio stream.

19. The non-transitory processor-readable medium of claim 15 , wherein the communication bearing data associated with the facial emotion further includes data associated with the speech emotion.

20. The non-transitory processor-readable medium of claim 15 , wherein the operations further comprise:

combining the facial emotion and the speech emotion to generate an emotional status of an individual associated with the at least one object of interest.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2017
From: AVATAR MERGER SUB II, LLC
To: SNAP INC.
Reel/Frame 044164/0377 →
Continuity (2)
Continuation 14661539 · Mar 18, 2015
Related Publication 20170154211A1 · Jun 1, 2017