IP Library Granted Patent US 10,657,985
Granted Patent B2
US 10,657,985 · App. 16/014,178 · Granted May 19, 2020

Systems and methods for manipulating electronic content based on speech recognition

Inventors: Peter F. Kocks (San Francisco, CA); Guoning Hu (Fremont, CA); Ping-Hao Wu (San Francisco, CA)
Assignee: Oath Inc.
G10L25/57G06F16/784G06F16/7834G10L15/06G10L15/08G10L17/00G10L17/005H04N21/4394G06F16/433H04N21/4668
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,657,985
App. No.
16/014,178
Granted
May 19, 2020
Kind
B2
Abstract

Systems and methods are disclosed for displaying electronic multimedia content to a user. One computer-implemented method for manipulating electronic multimedia content includes generating, using a processor, a speech model and at least one speaker model of an individual speaker. The method further includes receiving electronic media content over a network; extracting an audio track from the electronic media content; and detecting speech segments within the electronic media content based on the speech model. The method further includes detecting a speaker segment within the electronic media content and calculating a probability of the detected speaker segment involving the individual speaker based on the at least one speaker model.

Claims (77)

1. A computer-implemented method comprising the following operations performed by at least one processor:

detecting speaker segments within a plurality of electronic media content items, each of the plurality of electronic media content items being associated with media metadata;

determining, by the processor, at least one individual speaker associated with each of the speaker segments based on a speaker speech fingerprint;

determining, by the processor, speaker metadata associated with the at least one individual speaker;

receiving a search query from a user requesting a ranking of one or more electronic media content items;

determining, by the processor, a first ranked list comprising a first subset of the plurality of electronic media content items, based on a correspondence between the search query and the media metadata;

determining, by the processor, a second ranked list comprising a second subset of the plurality of electronic media content items, based on a correspondence between the speaker metadata and the search query;

processing, by the processor, the first ranked list and the second ranked list to determine a final ranking value for each of a third subset of the plurality of electronic media content items;

generating a third ranked list based on the final ranking value for each of the third subset of the plurality of electronic media content items; and

transmitting and displaying the third ranked list to the user.

2. A computer-implemented method of claim 1 , further comprising:

determining whether the final ranking value of the electronic media content equals or exceeds a threshold; and

presenting the electronic media content, associated with the third ranked list, to the user if the final ranking value of the electronic media content equals or exceeds the threshold.

3. The computer-implemented method of claim 1 , further comprising:

determining, using linear discriminant analysis, whether a speaker segment of the speaker segments within at least one of the plurality of electronic media content items is based on a speech fingerprint or a non-speech fingerprint.

4. The computer-implemented method of claim 1 , further comprising:

determining, based on a visual content analysis, a probability of each of the detected speaker segments being associated with an individual speaker; and

adjusting the first ranking value based on the probability of each of the detected speaker segments.

5. The computer-implemented method of claim 1 , further comprising:

generating at least one speaker speech fingerprint of an individual speaker, and a non-speaker speech fingerprint that includes common characteristics from one or more speakers.

6. The computer-implemented method of claim 1 , further comprising:

generating a plurality of speaker speech fingerprints for a subset of people, each speaker speech fingerprint corresponding to one person in a subset of people; and

calculating a probability of the speaker speech fingerprint involving one of the people in the subset of people, based on the plurality of speaker speech fingerprints.

7. The computer-implemented method of claim 1 , further comprising:

detecting, based on the speaker speech fingerprint, duplicate videos among the electronic media content items.

8. The computer-implemented method of claim 1 , further comprising:

filtering a database for at least one additional electronic media content item based on the speaker speech fingerprint; and

presenting a preview of the additional electronic media content item to the user.

9. A system, comprising:

at least one processor; and

at least one memory storing executable instructions that, when executed by the at least one processor, causes the at least one processor to perform the following operations:

detecting speaker segments within a plurality of electronic media content items, each of the plurality of electronic media content items being associated with media metadata;

determining, by the processor, at least one individual speaker associated with each of the speaker segments based on a speaker speech fingerprint;

determining, by the processor, speaker metadata associated with the at least one individual speaker;

receiving a search query from a user requesting a ranking of one or more electronic media content items;

determining, by the processor, a first ranked list comprising a first subset of the plurality of electronic media content items, based on a correspondence between the search query and the media metadata;

determining, by the processor, a second ranked list comprising a second subset of the plurality of electronic media content items, based on a correspondence between the speaker metadata and the search query;

processing, by the processor, the first ranked list and the second ranked list to determine a final ranking value for each of a third subset of the plurality of electronic media content items;

generating a third ranked list based on the final ranking value for each of the third subset of the plurality of electronic media content items; and

transmitting and displaying the third ranked list to the user.

10. A computer-implemented method of claim 9 further comprising:

determining whether the final ranking value of the electronic media content equals or exceeds a threshold; and

presenting the electronic media content, associated with the third ranked list, to the user if the final ranking value of the electronic media content equals or exceeds the threshold.

11. The computer-implemented method of claim 9 , further comprising:

determining, using linear discriminant analysis, whether a speaker segment of the speaker segments within at least one of the plurality of electronic media content items is based on a speech fingerprint or a non-speech fingerprint.

12. The computer-implemented method of claim 9 , further comprising:

determining, based on a visual content analysis, a probability of each of the detected speaker segments being associated with an individual speaker; and

adjusting the first ranking value based on the probability of each of the detected speaker segments.

13. The computer-implemented method of claim 9 , further comprising:

generating at least one speaker speech fingerprint of an individual speaker, and a non-speaker speech fingerprint that includes common characteristics from one or more speakers.

14. The computer-implemented method of claim 9 , further comprising:

generating a plurality of speaker speech fingerprints for a subset of people, each speaker speech fingerprint corresponding to one person in a subset of people; and

calculating a probability of the speaker speech fingerprint involving one of the people in the subset of people, based on the plurality of speaker speech fingerprints.

15. The computer-implemented method of claim 9 , further comprising:

detecting, based on the speaker speech fingerprint, duplicate videos among the electronic media content items.

16. The computer-implemented method of claim 9 , further comprising:

filtering a database for at least one additional electronic media content item based on the speaker speech fingerprint; and

presenting a preview of the additional electronic media content item to the user.

17. A tangible, non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

detecting speaker segments within a plurality of electronic media content items, each of the plurality of electronic media content items being associated with media metadata;

determining, by the processor, at least one individual speaker associated with each of the speaker segments based on a speaker speech fingerprint;

determining, by the processor, speaker metadata associated with the at least one individual speaker;

receiving a search query from a user requesting a ranking of one or more electronic media content items;

determining, by the processor, a first ranked list comprising a first subset of the plurality of electronic media content items, based on a correspondence between the search query and the media metadata;

determining, by the processor, a second ranked list comprising a second subset of the plurality of electronic media content items, based on a correspondence between the speaker metadata and the search query;

processing, by the processor, the first ranked list and the second ranked list to determine a final ranking value for each of a third subset of the plurality of electronic media content items;

generating a third ranked list based on the final ranking value for each of the third subset of the plurality of electronic media content items; and

transmitting and displaying the third ranked list to the user.

18. A computer-implemented method of claim 17 , further comprising:

determining whether the final ranking value of the electronic media content equals or exceeds a threshold; and

presenting the electronic media content, associated with the third ranked list, to the user if the final ranking value of the electronic media content equals or exceeds the threshold.

19. The computer-implemented method of claim 17 , further comprising:

generating a plurality of speaker speech fingerprints for a subset of people, each speaker speech fingerprint corresponding to one person in a subset of people; and

calculating a probability of the speaker speech fingerprint involving one of the people in the subset of people, based on the plurality of speaker speech fingerprints.

20. The computer-implemented method of claim 17 , further comprising:

filtering a database for at least one additional electronic media content item based on the speaker speech fingerprint; and

presenting a preview of the additional electronic media content item to the user.

Assignments (8)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2021
From: VERIZON MEDIA INC.
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 057453/0431 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2020
From: KOCKS, PETER F.; HU, GUONING; WU, PING-HAO
To: AOL INC.
Reel/Frame 052367/0280 →
CHANGE OF NAME Recorded Apr 10, 2020
From: AOL INC.
To: OATH INC.
Reel/Frame 052373/0125 →
CHANGE OF NAME Recorded Jul 2, 2018
From: AMERICA ONLINE, INC.
To: AOL LLC
Reel/Frame 046467/0991 →
CHANGE OF NAME Recorded Jul 2, 2018
From: AOL INC.
To: OATH INC.
Reel/Frame 046468/0073 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2018
From: RENNER, W. KARL; MURPHY, STEPHEN VAUGHAN
To: AMERICA ONLINE, INC.
Reel/Frame 046252/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2018
From: AOL LLC
To: AOL INC.
Reel/Frame 047246/0587 →
Continuity (4)
Continuation 15057414 · Mar 1, 2016
Continuation 13156780 · Jun 9, 2011
Provisional Application 61353518 · Jun 10, 2010
Related Publication 20180301161A1 · Oct 18, 2018