IP Library Granted Patent US 11,790,933
Granted Patent B2
US 11,790,933 · App. 16/836,500 · Granted Oct 17, 2023

Systems and methods for manipulating electronic content based on speech recognition

Inventors: Peter F. Kocks (San Francisco, CA); Guoning Hu (Fremont, CA); Ping-Hao Wu (San Francisco, CA)
Assignee: Verizon Patent and Licensing Inc.
G10L25/57G06F16/784G06F16/7834G10L15/06G10L15/08G10L17/00H04N21/4394G06F16/433H04N21/4668
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,790,933
App. No.
16/836,500
Granted
Oct 17, 2023
Kind
B2
Abstract

Systems and methods are disclosed for displaying electronic multimedia content to a user. One computer-implemented method for manipulating electronic multimedia content includes generating, using a processor, a speech model and at least one speaker model of an individual speaker. The method further includes receiving electronic media content over a network; extracting an audio track from the electronic media content; and detecting speech segments within the electronic media content based on the speech model. The method further includes detecting a speaker segment within the electronic media content and calculating a probability of the detected speaker segment involving the individual speaker based on the at least one speaker model.

Claims (77)

1. A computer-implemented method for manipulating electronic multimedia content, the method comprising:

generating, using a processor, a speech model and at least one speaker model of an individual speaker;

receiving electronic media content over a network;

extracting an audio track from the electronic media content;

utilizing the speech model and the at least one speaker model to identify speakers in the audio track by:

detecting a plurality of speech segments within the electronic media content based on the speech model; and

identifying at least one speaker associated with at least one of the speech segments by determining a probability that the at least one of the speech segments contains the individual speaker based on the at least one speaker model;

determining a ranking of the electronic media content relative to other electronic media content based on (i) the determined probability that at least one of the speech segments contains a voice of the individual speaker, (ii) probabilities for speech segments containing the voice of the individual speaker within the other electronic media content, (iii) text relevancy between a user query and metadata of the individual speaker, and (iv) a normalized duration of speech segments from the individual speaker; and

generating a fingerprint of the electronic media content based on the ranking of the electronic media content.

2. The computer-implemented method of claim 1 , further comprising:

determining probabilities that each of the plurality of the speech segments contains the voice of the individual speaker; and

determining a ranking or filtration of each of the plurality of speech segments or electronic media content based on the determined probabilities that each of the plurality of speech segments contains the voice of the individual speaker.

3. The computer-implemented method of claim 2 , wherein determining the ranking or filtration comprises:

analyzing the user query to generate a list of associated speakers; and

adjusting the ranking or filtration based on detected speech segments from speakers in the list of associated speakers.

4. The computer-implemented method of claim 2 , wherein determining probabilities comprises:

selecting electronic media content containing speech from speakers in a list of associated speakers.

5. The computer-implemented method of claim 1 , further comprising:

generating a plurality of speaker models for a subset of people, each speaker model corresponding to one person in the subset of people; and

determining probabilities that at least one of the plurality of speech segments contains a voice of one of the people in the subset of people, based on the plurality of speaker models.

6. The computer-implemented method of claim 1 , further comprising:

determining probabilities that each of the plurality of speech segments contains a voice of any of a plurality of speakers by comparing the plurality of speech segments with a plurality of speaker models; and

determining duplicated media in the electronic media content based on the probabilities that each of the plurality of speech segments contains the voice of any of the plurality of speakers.

7. The computer-implemented method of claim 6 , further comprising:

detecting words in the plurality of speech segments; and

displaying the electronic media content to users based on the probabilities that each of the plurality of speech segments contains the voice of any of the plurality of speakers and the detected words.

8. The computer-implemented method of claim 1 , further comprising:

based on the probability that the at least one of the speech segments contains the voice of the individual speaker, extracting at least one preview clip from the electronic media content, the preview clip being associated with the individual speaker; and

displaying the at least one preview clip to a user associated with the user query.

9. A system for manipulating electronic multimedia content, the system comprising:

at least one data storage device storing instructions for manipulating electronic multimedia content; and

at least one processor configured to execute the instructions stored in the data storage device to perform operations comprising:

generating, using the at least one processor, a speech model and at least one speaker model of an individual speaker;

receiving electronic media content over a network;

extracting an audio track from the electronic media content;

utilizing the speech model and the at least one speaker model to identify speakers in the audio track by:

detecting a plurality of speech segments within the electronic media content based on the speech model; and

identifying at least one speaker associated with at least one of the speech segments by determining a probability that the at least one of the speech segments contains the individual speaker based on the at least one speaker model;

determining a ranking of the electronic media content relative to other electronic media content based on (i) the determined probability that at least one of the speech segments contains a voice of the individual speaker, (ii) probabilities for speech segments containing the voice of the individual speaker within the other electronic media content, (iii) text relevancy between a user query and metadata of the individual speaker, and (iv) a normalized duration of speech segments from the individual speaker; and

generating a fingerprint of the electronic media content based on the ranking of the electronic media content.

10. The system of claim 9 , the operations further comprising:

determining probabilities that each of the plurality of the speech segments contains the voice of the individual speaker; and

determining a ranking or filtration of each of the plurality of speech segments or electronic media content based on the determined probabilities that each of the plurality of speech segments contains the voice of the individual speaker.

11. The system of claim 10 , wherein determining the ranking or filtration comprises:

analyzing the user query to generate a list of associated speakers; and

adjusting the ranking or filtration based on detected speech segments from speakers in the list of associated speakers.

12. The system of claim 10 , wherein determining probabilities comprises:

selecting electronic media content containing speech from speakers in a list of associated speakers.

13. The system of claim 9 , the operations further comprising:

generating a plurality of speaker models for a subset of people, each speaker model corresponding to one person in the subset of people; and

determining probabilities that at least one of the plurality of speech segments contains a voice of one of the people in the subset of people, based on the plurality of speaker models.

14. The system of claim 9 , the operations further comprising:

determining probabilities that each of the plurality of speech segments contains a voice of any of a plurality of speakers by comparing the plurality of speech segments with a plurality of speaker models; and

determining duplicated media in the electronic media content based on the probabilities that each of the plurality of speech segments contains the voice of any of the plurality of speakers.

15. The system of claim 14 , the operations further comprising:

detecting words in the plurality of speech segments; and

displaying the electronic media content to users based on the probabilities that each of the plurality of speech segments contains the voice of any of the plurality of speakers and the detected words.

16. The system of claim 9 , the operations further comprising:

based on the probability that the at least one of the speech segments contains the voice of the individual speaker, extracting at least one preview clip from the electronic media content, the preview clip being associated with the individual speaker; and

displaying the at least one preview clip to a user associated with the user query.

17. A non-transitory computer-readable medium for manipulating electronic multimedia content, storing instructions to execute operations comprising:

generating, using a processor, a speech model and at least one speaker model of an individual speaker;

receiving electronic media content over a network;

extracting an audio track from the electronic media content;

utilizing the speech model and the at least one speaker model to identify speakers in the audio track by:

detecting a plurality of speech segments within the electronic media content based on the speech model; and

identifying at least one speaker associated with at least one of the speech segments by determining a probability that the at least one of the speech segments contains the individual speaker based on the at least one speaker model;

determining a ranking of the electronic media content relative to other electronic media content based on (i) the determined probability that at least one of the speech segments contains a voice of the individual speaker, (ii) probabilities for speech segments containing the voice of the individual speaker within the other electronic media content, (iii) text relevancy between a user query and metadata of the individual speaker, and (iv) a normalized duration of speech segments from the individual speaker; and

generating a fingerprint of the electronic media content based on the ranking of the electronic media content.

18. The computer-readable medium of claim 17 , the operations further comprising:

determining probabilities that each of the plurality of the speech segments contains the voice of the individual speaker; and

determining a ranking or filtration of each of the plurality of speech segments or electronic media content based on the determined probabilities that each of the plurality of speech segments contains the voice of the individual speaker.

19. The computer-readable medium of claim 18 , wherein determining the ranking or filtration comprises:

analyzing the user query to generate a list of associated speakers; and

adjusting the ranking or filtration based on detected speech segments from speakers in the list of associated speakers.

20. The computer-readable medium of claim 18 , wherein determining probabilities comprises:

selecting electronic media content containing speech from speakers in a list of associated speakers.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2021
From: VERIZON MEDIA INC.
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 057453/0431 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2020
From: KOCKS, PETER F.; HU, GUONING; WU, PING-HAO
To: AOL INC.
Reel/Frame 052367/0280 →
CHANGE OF NAME Recorded Apr 10, 2020
From: AOL INC.
To: OATH INC.
Reel/Frame 052373/0125 →
Continuity (5)
Continuation 16014178 · Jun 21, 2018
Continuation 15057414 · Mar 1, 2016
Continuation 13156780 · Jun 9, 2011
Provisional Application 61353518 · Jun 10, 2010
Related Publication 20200251128A1 · Aug 6, 2020