IP Library › Granted Patent US 10,565,996
Granted Patent B2
US 10,565,996 · App. 15/170,264 · Granted Feb 18, 2020

Speaker identification

Inventors: Matthew Sharifi (Kilchberg, CH); Ignacio Lopez Moreno (New York, NY); Ludwig Schmidt (Cambridge, MA)
Assignee: Google LLC
G10L17/02G10L17/005G10L17/08G10L17/18G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,565,996
App. No.
15/170,264
Filed
Jun 1, 2016
Granted
Feb 18, 2020
Kind
B2
Art Unit
2657
USPC
704/246
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for performing speaker identification. In some implementations, data identifying a media item including speech of a speaker is received. Based on the received data, one or more other media items that include speech of the speaker are identified. One or more search results are generated that each reference a respective media item of the one or more other media items that include speech of the speaker. The one or more search results are provided for display.

Claims (35)

1. A method performed by one or more computers, the method comprising:

receiving, by the one or more computers, a request from a client device, the request including data identifying a media item including speech of a speaker;

based on the received data identifying the media item including speech of the speaker, identifying, by the one or more computers, one or more other media items that include speech of the speaker;

generating, by the one or more computers, one or more search results that each reference a respective media item of the one or more other media items that include speech of the speaker; and

providing, by the one or more computers and to the client device, a response to the request that includes the one or more search results for display.

2. The method of claim 1 , wherein receiving the request comprises receiving a request that includes a URL that identifies (i) a video that includes speech of the speaker, or (ii) an audio recording that includes speech of the speaker.

3. The method of claim 1 , wherein receiving the request comprises receiving a request for other content containing speech of the speaker whose speech is included in the media item.

4. The method of claim 1 , wherein receiving the request comprises receiving the media item from the client device over a network, the received media item comprising (i) video data that includes speech of the speaker, or (ii) audio data that includes speech of the speaker.

5. The method of claim 1 , further comprising providing, for display with the one or more search results, a name of the speaker.

6. The method of claim 5 , further comprising determining the name of the speaker based on comparison of speech characteristics determined from speech in the media item with speech characteristics determined from speech in additional media items that include speech of the speaker.

7. The method of claim 1 , wherein receiving the request comprises receiving a request to determine an identity of the speaker whose speech is included in the media item.

8. The method of claim 1 , wherein the media item includes speech of multiple speakers; and

wherein identifying the one or more other media items comprises identifying one or more other media items that each include speech of each of the multiple speakers.

9. The method of claim 8 , further comprising providing, for display with the one or more search results, a name of each of the multiple speakers.

10. The method of claim 1 , wherein generating the one or more search results comprises generating one or more search results that each include a link to a media item that is available on the Internet and that includes speech of the speaker.

11. The method of claim 1 , wherein identifying the one or more other media items comprises identifying the one or more other media items based on audio characteristics of the identified media item.

12. The method of claim 1 , wherein identifying the one or more other media items comprises identifying multiple media items based on records of multiple different indexes that have records indexed using different indexing functions.

13. A system comprising:

one or more computers; and

one or more computer-readable media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving, by the one or more computers, a request from a client device, the request including data identifying a media item including speech of a speaker;

based on the received data identifying the media item including speech of the speaker, identifying, by the one or more computers, one or more other media items that include speech of the speaker;

generating, by the one or more computers, one or more search results that each reference a respective media item of the one or more other media items that include speech of the speaker; and

providing, by the one or more computers and to the client device, a response to the request that includes the one or more search results for display.

14. The system of claim 13 , wherein receiving the request comprises receiving a request that includes a URL that identifies (i) a video that includes speech of the speaker, or (ii) an audio recording that includes speech of the speaker.

15. The system of claim 13 , wherein receiving the request comprises receiving a request for other content containing speech of the speaker whose speech is included in the media item.

16. The system of claim 13 , wherein receiving the request comprises receiving the media item from the client device over a network, the received media item comprising (i) video data that includes speech of the speaker, or (ii) audio data that includes speech of the speaker.

17. One or more non-transitory computer-readable media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

receiving, by the one or more computers, a request from a client device, the request including data identifying a media item including speech of a speaker;

based on the received data identifying the media item including speech of the speaker, identifying, by the one or more computers, one or more other media items that include speech of the speaker;

generating, by the one or more computers, one or more search results that each reference a respective media item of the one or more other media items that include speech of the speaker; and

providing, by the one or more computers and to the client device, a response to the request that includes the one or more search results for display.

18. The one or more non-transitory computer-readable media of claim 17 , wherein receiving the request comprises receiving a request that includes a URL that identifies (i) a video that includes speech of the speaker, or (ii) an audio recording that includes speech of the speaker.

19. The one or more non-transitory computer-readable media of claim 17 , wherein receiving the request comprises receiving a request for other content containing speech of the speaker whose speech is included in the media item.

20. The one or more non-transitory computer-readable media of claim 17 , wherein receiving the request comprises receiving the media item from the client device over a network, the received media item comprising (i) video data that includes speech of the speaker, or (ii) audio data that includes speech of the speaker.

Assignments (2)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2016
From: SHARIFI, MATTHEW; MORENO, IGNACIO LOPEZ; SCHMIDT, LUDWIG
To: GOOGLE INC.
Reel/Frame 038766/0434 →
Continuity (3)
Continuation 14523198 · Oct 24, 2014
Provisional Application 61899434 · Nov 4, 2013
Related Publication 20160275953A1 · Sep 22, 2016