IP Library Granted Patent US 9,311,395
Granted Patent B2
US 9,311,395 · App. 13/156,780 · Granted Apr 12, 2016

Systems and methods for manipulating electronic content based on speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,311,395
App. No.
13/156,780
Granted
Apr 12, 2016
Kind
B2
Abstract

Systems and methods are disclosed for displaying electronic multimedia content to a user. One computer-implemented method for manipulating electronic multimedia content includes generating, using a processor, a speech model and at least one speaker model of an individual speaker. The method further includes receiving electronic media content over a network; extracting an audio track from the electronic media content; and detecting speech segments within the electronic media content based on the speech model. The method further includes detecting a speaker segment within the electronic media content and calculating a probability of the detected speaker segment involving the individual speaker based on the at least one speaker model.

Claims (69)

1. A computer-implemented method for manipulating electronic multimedia content, the method comprising:

generating, using a processor, a speech model, a non-speech model, at least one speaker model of an individual speaker, and a non-speaker speech model;

receiving electronic media content over a network;

extracting an audio track from the electronic media content;

detecting speech segments within the extracted audio track based on the speech model and the non-speech model, the speech segments containing speech from at least one of a plurality of speakers;

detecting a speaker segment within the detected speech segments based on the speaker model and the non-speaker speech model, the speaker segment containing speech from the individual speaker;

calculating a first probability of the detected speaker segment involving the individual speaker based on the at least one speaker speech model and the non-speaker speech model;

determining a ranking or filtration of the electronic media content relative to other electronic media content based on the first probability of the detected speaker segment;

detecting a face within a part of the electronic media content corresponding to the detected speaker segment and calculating a second probability of the detected face being a face of the individual speaker; and

adjusting the ranking or filtration of the electronic media content based on the second probability.

2. The computer-implemented method of claim 1 , further comprising:

displaying electronic media content to users based on the ranking or filtration.

3. The computer-implemented method of claim 1 , further comprises:

analyzing a query to generate a list of associated speakers; and

adjusting the ranking of electronic media content based on detected speech segments from speakers in the list.

4. The computer-implemented method of claim 1 , further comprises:

analyzing a query to generate a list of associated speakers; and

selecting electronic media content that have speech from speakers in the list.

5. The computer-implemented method of claim 1 , further comprising:

generating a plurality of speaker models for a subset of people, each speaker model corresponding to one person in the subset of people; and

calculating a probability of the speaker segment involving one of the people in the subset of people, based on the plurality of speaker models.

6. The computer implemented method of claim 1 , further comprising:

applying speaker segments and their probabilities to detect duplicated videos, among electronic media content.

7. The computer implemented method of claim 1 , further comprising:

applying speaker segments and their probabilities to detect words spoken by a particular individual speaker.

8. The computer-implemented method of claim 7 , further comprising:

applying detected words from the particular individual speaker to the ranking or filtration of electronic media content; and

displaying electronic media content to users based on the ranking or filtration.

9. The computer-implemented method of claim 1 , further comprising:

applying speaker segments and their probabilities to detect individual speakers represented in electronic media content.

10. The computer-implemented method of claim further comprising:

applying detected individual speakers to the ranking or filtration of electronic media content; and

displaying electronic media content to users based on the ranking or filtration.

11. The computer-implemented method of claim 1 , further comprising:

applying speaker segments and their probabilities to extract preview clips from electronic media content; and

displaying the extracted preview clips associated with electronic media content to users.

12. A system for manipulating electronic multimedia content, the system comprising:

a data storage device storing instructions for manipulating electronic multimedia content; and

a processor configured to execute the instructions stored in the data storage device for:

generating a speech model, a non-speech model, at least one speaker model of an individual speaker, and a non-speaker speech model;

receiving electronic media content over a network;

extracting an audio track from the electronic media content;

detecting speech segments within the extracted audio track based on the speech model and the non-speech model, the speech segments containing speech from at least one of a plurality of speakers;

detecting a speaker segment within the detected speech segments based on the speaker model and the non-speaker speech model, the speaker segment containing speech from the individual speaker;

calculating a first probability of the detected speaker segment involving the individual speaker based on the at least one speaker speech model and the non-speaker speech model;

determining a ranking or filtration of the electronic media content relative to other electronic media content based on the first probability of the detected speaker segment;

detecting a face within a part of the electronic media content corresponding to the detected speaker segment and calculating a second probability of the detected face being a face of the individual speaker; and

adjusting the ranking or filtration of the electronic media content based on the second probability.

13. The system of claim 12 , wherein the processor is further configured to execute instructions for:

displaying electronic media content to users based on the ranking or filtration.

14. The system of claim 12 , wherein the processor is further configured to execute instructions for:

analyzing a query to generate a list of associated speakers; and

adjusting the ranking of electronic media content based on detected speech segments from speakers in the list.

15. The system of claim 12 , wherein the processor is further configured to execute instructions for:

analyzing a query to generate a list of associated speakers; and

selecting electronic media content that have speech from speakers in the list.

16. The system of claim 12 , wherein the processor is further configured to execute instructions for:

generating a plurality of speaker models for a subset of people, each speaker model corresponding to one person in the subset of people; and

calculating a probability of the speaker segment involving one of the people in the subset of people, based on the plurality of speaker models.

17. The system of claim 12 , wherein the processor is further configured for:

applying speaker segments and their probabilities to detect duplicated videos, among electronic media content.

18. The system of claim 12 , wherein the processor is further configured to execute instructions for:

applying speaker segments and their probabilities to detect words spoken by a particular individual speaker.

19. The system of claim 18 , wherein the processor is further configured for:

applying detected words from the particular individual speaker to the ranking or filtration of electronic media content; and

displaying electronic media content to users based on the ranking or filtration.

20. The system of claim 12 , wherein the processor is further configured to execute instructions for:

applying speaker segments and their probabilities to extract preview clips from electronic media content; and

displaying the extracted preview clips associated with electronic media content to users.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2021
From: VERIZON MEDIA INC.
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 057453/0431 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
CHANGE OF NAME Recorded Aug 24, 2017
From: AOL INC.
To: OATH INC.
Reel/Frame 043672/0369 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS -RELEASE OF 030936/0011 Recorded Jul 1, 2015
From: JPMORGAN CHASE BANK, N.A.
To: AOL ADVERTISING INC.; AOL INC.; BUYSIGHT, INC.; MAPQUEST, INC.; PICTELA, INC.
Reel/Frame 036042/0053 →
SECURITY AGREEMENT Recorded Aug 2, 2013
From: AOL INC.; AOL ADVERTISING INC.; BUYSIGHT, INC.; MAPQUEST, INC.; PICTELA, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 030936/0011 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2011
From: KOCKS, PETER F.; HU, GUONING; WU, PING-HAO
To: AOL INC.
Reel/Frame 026933/0176 →