IP Library Granted Patent US 9,123,330
Granted Patent B1
US 9,123,330 · App. 13/875,001 · Granted Sep 1, 2015

Large-scale speaker identification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,123,330
App. No.
13/875,001
Granted
Sep 1, 2015
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for receiving audio data encoding ambient sounds, identifying media content that matches the audio data, and a timestamp corresponding to a particular portion of the identified media content, identifying a speaker associated with the particular portion of the identified media content corresponding to the timestamp, and providing information identifying the speaker associated with the particular portion of the identified media content for output.

Claims (94)

1. A computer-implemented method comprising:

receiving (i) a request to identify a speaker, and (ii) an audio data representation of ambient sounds;

transmitting, to a content recognition engine, a request to identify an item of media content that is associated with the ambient sounds;

obtaining, from the content recognition engine, a response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes (i) data identifying a particular item of media content that the content recognition engine has associated with the ambient sounds, and (ii) a timestamp associated with a portion of the particular item of media content that the content recognition engine has associated with the ambient sounds, and wherein the response to the request to identify an item of media content that is associated with the ambient sounds does not identify any speaker;

transmitting, to a speaker identification engine, a request to identify a speaker that is associated with the portion of the particular item of media content that is associated with the timestamp;

obtaining, from the speaker identification engine, a response to the request to identify a speaker that is associated with the portion of the particular item of media content that is associated with the timestamp, wherein the response identifies a particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp; and

providing, for output, information identifying the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp.

2. The method of claim 1 , wherein receiving the audio data representation of the ambient sounds comprises:

receiving a request to identify a speaker whose voice is included among the ambient sounds; and

receiving the audio data representation of the ambient sounds based on receiving the request to identify the speaker whose voice is included among the ambient sounds.

3. The method of claim 2 , wherein receiving the request to identify the speaker whose voice is included among the ambient sounds comprises detecting a spoken utterance input by a user.

4. The method of claim 2 , wherein receiving the request to identify the speaker whose voice is included among the ambient sounds comprises detecting a user selection of a control.

5. The method of claim 1 , wherein obtaining the response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds comprises:

obtaining one or more audio fingerprints of the audio data representation of the ambient sounds;

comparing one or more of the audio fingerprints of the audio data representation of the ambient sounds to one or more audio fingerprints of prerecorded versions of items of media content;

identifying the particular item of media content that the content recognition engine has associated with the ambient sounds based on determining that one or more of the audio fingerprints of the audio data representation of the ambient sounds match one or more audio fingerprints of the particular item of media content; and

identifying, as the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds, a particular timestamp associated with at least one of the one or more audio fingerprints of the particular item of media content.

6. The method of claim 1 , wherein obtaining the response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds comprises:

segmenting prerecorded versions of items of media content into one or more portions;

assigning timestamps to each of the one or more portions of the prerecorded versions of the items of media content;

obtaining one or more audio fingerprints of each of the one or more portions of the prerecorded versions of the items of media content;

obtaining one or more audio fingerprints of the audio data representation of the ambient sounds;

comparing one or more of the audio fingerprints of the audio data representation of the ambient sounds to one or more of the audio fingerprints of one or more of the portions of the prerecorded versions of the items of media content; and

identifying the timestamp associated with the portion of the particular item of media content based on detecting a match between one or more of the audio fingerprints of the audio data representation of the ambient sounds and one or more of the audio fingerprints of the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds.

7. The method of claim 1 , wherein obtaining the response to the request to identify a speaker that is associated with the portion of the particular item of media content that is associated with the timestamp, wherein the response identifies the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp comprises:

accessing information that identifies timestamps and speakers associated with those timestamps for various prerecorded versions of items of media content;

identifying, based on the information that identifies timestamps and speakers associated with those timestamps for various prerecorded versions of items of media content, a particular timestamp corresponding to the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds; and

identifying, as the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp, a particular speaker associated with the particular timestamp corresponding to the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds.

8. The method of claim 1 , wherein providing information identifying the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp comprises:

transmitting, to a speaker knowledge base, a request to identify information that is associated with one of the speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp or the particular item of media content that the content recognition engine has associated with the ambient sounds;

obtaining, from the speaker knowledge base, a response to the request to identify information that is associated with one of the speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp or the particular item of media content that the content recognition engine has associated with the ambient sounds, wherein the response identifies information relating to the speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp or the particular item of media content that the content recognition engine has associated with the ambient sounds; and

providing, for output, at least a portion of the information relating to the speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp or the particular item of media content that the content recognition engine has associated with the ambient sounds.

9. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving (i) a request to identify a speaker, and (ii) an audio data representation of ambient sounds;

transmitting, to a content recognition engine, a request to identify an item of media content that is associated with the ambient sounds;

obtaining, from the content recognition engine, a response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes (i) data identifying a particular item of media content that the content recognition engine has associated with the ambient sounds, and (ii) a timestamp associated with a portion of the particular item of media content that the content recognition engine has associated with the ambient sounds, and wherein the response to the request to identify an item of media content that is associated with the ambient sounds does not identify any speaker;

transmitting, to a speaker identification engine, a request to identify a speaker that is associated with the portion of the particular item of media content that is associated with the timestamp;

obtaining, from the speaker identification engine, a response to the request to identify a speaker that is associated with the portion of the particular item of media content that is associated with the timestamp, wherein the response identifies a particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp; and

providing, for output, information identifying the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp.

10. The system of claim 9 , wherein receiving the audio data representation of the ambient sounds comprises:

receiving a request to identify a speaker whose voice is included among the ambient sounds; and

receiving the audio data representation of the ambient sounds based on receiving the request to identify the speaker whose voice is included among the ambient sounds.

11. The system of claim 10 , wherein receiving the request to identify the speaker whose voice is included among the ambient sounds comprises detecting a spoken utterance input by a user.

12. The system of claim 10 , wherein receiving the request to identify the speaker whose voice is included among the ambient sounds comprises detecting a user selection of a control.

13. The system of claim 9 , wherein obtaining the response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds comprises:

obtaining one or more audio fingerprints of the audio data representation of the ambient sounds;

comparing one or more of the audio fingerprints of the audio data representation of the ambient sounds to one or more audio fingerprints of prerecorded versions of items of media content;

identifying the particular item of media content that the content recognition engine has associated with the ambient sounds based on determining that one or more of the audio fingerprints of the audio data representation of the ambient sounds match one or more audio fingerprints of the particular item of media content; and

identifying, as the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds, a particular timestamp associated with at least one of the one or more audio fingerprints of the particular item of media content.

14. The system of claim 9 , wherein obtaining the response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds comprises:

segmenting prerecorded versions of items of media content into one or more portions;

assigning timestamps to each of the one or more portions of the prerecorded versions of the items of media content;

obtaining one or more audio fingerprints of each of the one or more portions of the prerecorded versions of the items of media content;

obtaining one or more audio fingerprints of the audio data representation of the ambient sounds;

comparing one or more of the audio fingerprints of the audio data representation of the ambient sounds to one or more of the audio fingerprints of one or more of the portions of the prerecorded versions of the items of media content; and

identifying the timestamp associated with the portion of the particular item of media content based on detecting a match between one or more of the audio fingerprints of the audio data representation of the ambient sounds and one or more of the audio fingerprints of the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds.

15. The system of claim 9 , wherein obtaining the response to the request to identify a speaker that is associated with the portion of the particular item of media content that is associated with the timestamp, wherein the response identifies the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp comprises:

accessing information that identifies timestamps and speakers associated with those timestamps for various prerecorded versions of items of media content;

identifying, based on the information that identifies timestamps and speakers associated with those timestamps for various prerecorded versions of items of media content, a particular timestamp corresponding to the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds; and

identifying, as the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp, a particular speaker associated with the particular timestamp corresponding to the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds.

16. The system of claim 9 , wherein obtaining the response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds comprises:

identifying, based on the audio data representation of the ambient sounds, the particular item of media content; and

identifying the timestamp associated with the portion of the particular item of media content based on determining that the ambient sounds represented by the audio data representation of the ambient sounds correspond to audio of the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds.

17. A computer-readable storage device encoded with a computer program, the program comprising instructions that if executed by one or more computers cause the one or more computers to perform operations comprising:

receiving (i) a request to identify a speaker, and (ii) an audio data representation of ambient sounds;

transmitting, to a content recognition engine, a request to identify an item of media content that is associated with the ambient sounds;

obtaining, from the content recognition engine, a response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes (i) data identifying a particular item of media content that the content recognition engine has associated with the ambient sounds, and (ii) a timestamp associated with a portion of the particular item of media content that the content recognition engine has associated with the ambient sounds, and wherein the response to the request to identify an item of media content that is associated with the ambient sounds does not identify any speaker;

transmitting, to a speaker identification engine, a request to identify a speaker that is associated with the portion of the particular item of media content that is associated with the timestamp;

obtaining, from the speaker identification engine, a response to the request to identify a speaker that is associated with the portion of the particular item of media content that is associated with the timestamp, wherein the response identifies a particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp; and

providing, for output, information identifying the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp.

18. The device of claim 17 , wherein receiving the audio data representation of the ambient sounds comprises:

receiving a request to identify a speaker whose voice is included among the ambient sounds; and

receiving the audio data representation of the ambient sounds based on receiving the request to identify the speaker whose voice is included among the ambient sounds.

19. The device of claim 18 , wherein receiving the request to identify the speaker whose voice is included among the ambient sounds comprises detecting a spoken utterance input by a user.

20. The device of claim 17 , wherein obtaining the response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds comprises:

obtaining one or more audio fingerprints of the audio data representation of the ambient sounds;

comparing one or more of the audio fingerprints of the audio data representation of the ambient sounds to one or more audio fingerprints of prerecorded versions of items of media content;

identifying the particular item of media content that the content recognition engine has associated with the ambient sounds based on determining that one or more of the audio fingerprints of the audio data representation of the ambient sounds match one or more audio fingerprints of the particular item of media content; and

identifying, as the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds, a particular timestamp associated with at least one of the one or more audio fingerprints of the particular item of media content.

21. The device of claim 17 , wherein obtaining the response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds comprises:

segmenting prerecorded versions of items of media content into one or more portions;

assigning timestamps to each of the one or more portions of the prerecorded versions of the items of media content;

obtaining one or more audio fingerprints of each of the one or more portions of the prerecorded versions of the items of media content;

obtaining one or more audio fingerprints of the audio data representation of the ambient sounds;

comparing one or more of the audio fingerprints of the audio data representation of the ambient sounds to one or more of the audio fingerprints of one or more of the portions of the prerecorded versions of the items of media content; and

identifying the timestamp associated with the portion of the particular item of media content based on detecting a match between one or more of the audio fingerprints of the audio data representation of the ambient sounds and one or more of the audio fingerprints of the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds.

22. The device of claim 17 , wherein obtaining the response to the request to identify a speaker that is associated with the portion of the particular item of media content that is associated with the timestamp, wherein the response identifies the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp comprises:

accessing information that identifies timestamps and speakers associated with those timestamps for various prerecorded versions of items of media content;

identifying, based on the information that identifies timestamps and speakers associated with those timestamps for various prerecorded versions of items of media content, a particular timestamp corresponding to the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds; and

identifying, as the particular speaker that the speaker identification engine has associated with the portion of the particular item of media content that is associated with the timestamp, a particular speaker associated with the particular timestamp corresponding to the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds.

23. The method of claim 1 , wherein obtaining the response to the request to identify an item of media content that is associated with the ambient sounds, wherein the response includes the timestamp associated with the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds comprises:

identifying, based on the audio data representation of the ambient sounds, the particular item of media content; and

identifying the timestamp associated with the portion of the particular item of media content based on determining that the ambient sounds represented by the audio data representation of the ambient sounds correspond to audio of the portion of the particular item of media content that the content recognition engine has associated with the ambient sounds.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044334/0466 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2013
From: SHARIFI, MATTHEW; ROBLEK, DOMINIK
To: GOOGLE INC.
Reel/Frame 030627/0331 →