IP Library Granted Patent US 10,204,625
Granted Patent B2
US 10,204,625 · App. 15/861,855 · Granted Feb 12, 2019

Audio analysis learning using video data

Inventors: Taniya Mishra (New York, NY); Rana el Kaliouby (Milton, MA)
Assignee: Affectiva, Inc.
G10L15/22B60R16/0373B60W50/10G06K9/00288G10L15/02G10L15/1807G10L15/25G10L21/055G10L25/51G10L25/63G10L25/90B60W2540/02B60W2540/04G10L21/0356G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,204,625
App. No.
15/861,855
Granted
Feb 12, 2019
Kind
B2
Abstract

Audio analysis learning is performed using video data. Video data is obtained, on a first computing device, wherein the video data includes images of one or more people. Audio data is obtained, on a second computing device, which corresponds to the video data. A face is identified within the video data. A first voice, from the audio data, is associated with the face within the video data. The face within the video data is analyzed for cognitive content. Audio features are extracted corresponding to the cognitive content of the video data. The audio data is segmented to correspond to an analyzed cognitive state. An audio classifier is learned, on a third computing device, based on the analyzing of the face within the video data. Further audio data is analyzed using the audio classifier.

Claims (54)

1. A computer-implemented method for audio analysis comprising:

obtaining video data, on a first computing device, wherein the video data includes images of one or more people;

obtaining audio data, on a second computing device, corresponding to the video data;

identifying a face within the video data;

associating a first voice, from the audio data, with the face within the video data;

analyzing the face within the video data for cognitive content;

learning an audio classifier, on a third computing device, based on the analyzing of the face within the video data;

identifying and separating a second voice from the obtained audio data corresponding to the video data but not associated with the face associated with a first voice, wherein the second voice is included in the learning and wherein the second voice corresponds to a second person; and

analyzing further audio data using the audio classifier.

2. The method of claim 1 further comprising extracting audio features corresponding to the cognitive content of the video data.

3. The method of claim 1 further comprising segmenting the audio data to correspond to an analyzed cognitive state.

4. The method of claim 3 wherein the segmenting the audio data is for a human sensorially detectable unit of time.

5. The method of claim 4 wherein the segmenting the audio data includes noticeable differences in intensity, duration, or pitch.

6. The method of claim 3 wherein the segmenting the audio data is for less than thirty seconds.

7. The method of claim 1 further comprising synchronizing the audio data with the video data.

8. The method of claim 1 further comprising analyzing a first voice for features.

9. The method of claim 8 wherein the analyzing the first voice for features includes evaluation of timbre.

10. The method of claim 8 wherein the analyzing the first voice for features includes evaluation of prosody.

11. The method of claim 8 wherein the analyzing the first voice for features includes analysis of vocal register and vocal resonance, pitch, speech loudness, or speech rate.

12. The method of claim 8 wherein the analyzing the first voice for features includes language analysis.

13. The method of claim 12 wherein the language analysis is dependent on language content.

14. The method of claim 1 wherein the learning is independent of language content.

15. The method of claim 1 wherein the learning further encompasses learning a second audio classifier.

16. The method of claim 1 wherein the cognitive content includes detection of one or more of sadness, stress, happiness, anger, frustration, confusion, disappointment, hesitation, cognitive overload, focusing, engagement, attention, boredom, exploration, confidence, trust, delight, disgust, skepticism, doubt, satisfaction, excitement, laughter, calmness, curiosity, humor, depression, envy, sympathy, embarrassment, poignancy, fatigue, drowsiness, or mirth.

17. The method of claim 1 wherein the identifying of the face includes detection of facial expressions.

18. The method of claim 1 further comprising determining a temporal audio signature for use with the further audio data.

19. The method of claim 1 further comprising associating the second voice, from the second person, with a second face within the video data.

20. The method of claim 1 further comprising manipulating a vehicle based on the analyzing of the further audio data.

21. The method of claim 20 wherein the manipulating the vehicle includes transfer into autonomous mode, transfer out of autonomous mode, locking out operation, recommending a break for an occupant, recommending a different route, recommending how far to drive, responding to traffic, adjusting seats, adjusting mirrors, climate control, lighting, music, audio stimuli, interior temperature, brake activation, and steering control.

22. The method of claim 1 wherein the learning the audio classifier is based on analyzing a plurality of faces within the video data.

23. The method of claim 1 wherein the learning further comprises:

synchronizing the audio data and the video data;

extracting an audio feature associated with the cognitive content that was analyzed from the face; and

abstracting an audio classifier based on the extracted audio feature.

24. A computer program product embodied in a non-transitory computer readable medium for audio analysis, the computer program product comprising code which causes one or more processors to perform operations of:

obtaining video data wherein the video data includes images of one or more people;

obtaining audio data corresponding to the video data;

identifying a face within the video data;

associating a first voice, from the audio data, with the face within the video data;

analyzing the face within the video data for cognitive content;

learning an audio classifier based on the analyzing of the face within the video data;

identifying and separating a second voice from the obtained audio data corresponding to the video data but not associated with the face associated with a first voice, wherein the second voice is included in the learning and wherein the second voice corresponds to a second person; and

analyzing further audio data using the audio classifier.

25. A computer system for audio analysis comprising:

a memory which stores instructions;

one or more processors attached to the memory wherein the one or more processors, when executing the instructions which are stored, are configured to:

obtain video data, on a first computing device, wherein the video data includes images of one or more people;

obtain audio data, on a second computing device, corresponding to the video data;

identify a face within the video data;

associate a first voice, from the audio data, with the face within the video data;

analyze the face within the video data for cognitive content;

learn an audio classifier, on a third computing device, based on the analyzing of the face within the video data;

identify and separate a second voice from the obtained audio data corresponding to the video data but not associated with the face associated with a first voice, wherein the second voice is included in learning and wherein the second voice corresponds to a second person; and

analyze further audio data using the audio classifier.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2018
From: MISHRA, TANIYA; EL KALIOUBY, RANA
To: AFFECTIVA, INC.
Reel/Frame 044801/0349 →
Continuity (54)
Continuation In Part 15670791 · Aug 7, 2017
Continuation In Part 14214918 · Mar 15, 2014
Continuation In Part 13153745 · Jun 6, 2011
Continuation In Part 15666048 · Aug 1, 2017
Continuation In Part 15395750 · Dec 30, 2016
Continuation In Part 15262197 · Sep 12, 2016
Continuation In Part 14796419 · Jul 10, 2015
Continuation In Part 13153745 · Jun 6, 2011
Continuation In Part 14460915 · Aug 15, 2014
Continuation In Part 13153745 · Jun 6, 2011
Provisional Application 62442325 · Jan 4, 2017
Provisional Application 62442291 · Jan 4, 2017
Provisional Application 62448448 · Jan 20, 2017
Provisional Application 62469591 · Mar 10, 2017
Provisional Application 62503485 · May 9, 2017
Provisional Application 62524606 · Jun 25, 2017
Provisional Application 62541847 · Aug 7, 2017
Provisional Application 62557460 · Sep 12, 2017
Provisional Application 62593449 · Dec 1, 2017
Provisional Application 62593440 · Dec 1, 2017
Provisional Application 62611780 · Dec 29, 2017
Provisional Application 62439928 · Dec 29, 2016
Provisional Application 61789038 · Mar 15, 2013
Provisional Application 61793761 · Mar 15, 2013
Provisional Application 61790461 · Mar 15, 2013
Provisional Application 61798731 · Mar 15, 2013
Provisional Application 61844478 · Jul 10, 2013
Provisional Application 61916190 · Dec 14, 2013
Provisional Application 61924252 · Jan 7, 2014
Provisional Application 61927481 · Jan 15, 2014
Provisional Application 61352166 · Jun 7, 2010
Provisional Application 61388002 · Sep 30, 2010
Provisional Application 61414451 · Nov 17, 2010
Provisional Application 61439913 · Feb 6, 2011
Provisional Application 61447089 · Feb 27, 2011
Provisional Application 61447464 · Feb 28, 2011
Provisional Application 61467209 · Mar 24, 2011
Provisional Application 62370421 · Aug 3, 2016
Provisional Application 62273896 · Dec 31, 2015
Provisional Application 62301558 · Feb 29, 2016
Provisional Application 62217872 · Sep 12, 2015
Provisional Application 62222518 · Sep 23, 2015
Provisional Application 62265937 · Dec 10, 2015
Provisional Application 62273896 · Dec 31, 2015
Provisional Application 62301558 · Feb 29, 2016
Provisional Application 62370421 · Aug 3, 2016
Provisional Application 62023800 · Jul 11, 2014
Provisional Application 62047508 · Sep 8, 2014
Provisional Application 62082579 · Nov 20, 2014
Provisional Application 62128974 · Mar 5, 2015
Provisional Application 61867007 · Aug 16, 2013
Provisional Application 61953878 · Mar 16, 2014
Provisional Application 61972314 · Mar 30, 2014
Related Publication 20180144746A1 · May 24, 2018
Cited By (3)
US 12,248,551 US 12,380,895 US 12,450,806