IP Library Granted Patent US 11,133,005
Granted Patent B2
US 11,133,005 · App. 16/397,004 · Granted Sep 28, 2021

Systems and methods for disambiguating a voice search query

Inventors: Ankur Aher (Maharashtra, IN); Sindhuja Chonat Sri (Tamil Nadu, IN); Aman Puniyani (Karnataka, IN); Nishchit Mahajan (Punjab, IN)
Assignee: Rovi Guides, Inc.
G10L15/22G06F16/635G06F16/638G06F16/683G10L15/08G10L25/51G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,133,005
App. No.
16/397,004
Granted
Sep 28, 2021
Kind
B2
Abstract

Systems and methods are described herein for disambiguating a voice search query that contains a command keyword by determining whether the user spoke a quotation from a content item and whether the user mimicked or approximated the way the quotation is spoken in the content item. The voice search query is transcribed into a string, and an audio signature of the voice search query is identified. Metadata of a quotation matching the string is retrieved from a database that includes audio signature information for the string as spoken within the content item. The audio signature of the voice search query is compared with the audio signature information in the metadata to determine whether the audio signature matches the audio signature information in the quotation metadata. If a match is detected, then a search result comprising an identifier of the content item from which the quotation comes is generated.

Claims (99)

1. A method for disambiguating a voice search query, the method comprising:

receiving a voice search query containing a command keyword;

transcribing the voice search query into a string comprising a plurality of words;

determining an audio signature of the voice search query;

querying a database with the string;

receiving, in response to the query, metadata of a quotation, the metadata comprising the string, audio signature information for the string as spoken within a particular content item, and an identifier of the particular content item;

comparing the audio signature of the voice search query with the audio signature information in the metadata of the quotation;

determining, based on the comparing, whether the voice search query mimics a voice that produced the quotation; and

in response to determining that the voice search query mimics a voice that produced the quotation, generating for display a search result comprising the identifier of the particular content item.

2. The method of claim 1 , further comprising:

determining a cadence of each word of the plurality of words;

retrieving metadata of at least one quotation, the metadata comprising a second string that is similar to the string and comprises a second plurality of words, and cadence information for each word of the second plurality of words; and

comparing a cadence of each word of the plurality of words with cadence information in the metadata for each corresponding word of the second plurality of words;

wherein determining whether the audio signature matches the audio signature information in the metadata of a quotation comprises determining, based on the comparing, whether the cadence of each word of the plurality of words matches the cadence information for each corresponding word of the second plurality of words.

3. The method of claim 1 , further comprising:

determining an emphasis of each word of the plurality of words;

retrieving metadata of at least one quotation, the metadata comprising a second string that is similar to the string and comprises a second plurality of words, and emphasis information for each word of the second plurality of words; and

comparing an emphasis of each word of the plurality of words with emphasis information in the metadata for each corresponding word of the second plurality of words;

wherein determining whether the audio signature matches the audio signature information in the metadata of a quotation comprises determining, based on the comparing, whether the emphasis of each word of the plurality of words matches the emphasis information for each corresponding word of the second plurality of words.

4. The method of claim 1 , further comprising:

determining a duration of each word of the plurality of words;

retrieving metadata of at least one quotation, the metadata comprising a second string that is similar to the string and comprises a second plurality of words, and duration information for each word of the second plurality of words; and

comparing a duration of each word of the plurality of words with duration information in the metadata for each corresponding word of the second plurality of words;

wherein determining whether the audio signature matches the audio signature information in the metadata of a quotation comprises determining, based on the comparing, whether the duration of each word of the plurality of words matches the duration information for each corresponding word of the second plurality of words.

5. The method of claim 1 , further comprising:

determining an overall rhythm of the plurality of words;

retrieving metadata of at least one quotation, the metadata comprising a second string that is similar to the string and comprises a second plurality of words, and rhythm information for the second plurality of words; and

comparing the overall rhythm of the plurality of words with rhythm information in the metadata for the second plurality of words;

wherein determining whether the audio signature matches the audio signature information in the metadata of a quotation comprises determining, based on the comparing, whether the overall rhythm of the plurality of words matches the rhythm information for the second plurality of words.

6. The method of claim 1 , wherein determining whether the audio signature matches audio signature information on the metadata of a quotation comprises:

establishing a lower threshold by negatively transposing the audio signature information in the metadata by a predetermined amount;

establishing an upper threshold by positively transposing the audio signature information in the metadata by the predetermined amount; and

determining whether the audio signature is between the lower threshold and the upper threshold.

7. The method of claim 2 , wherein determining whether the cadence of each word of the plurality of words matches the cadence information for each corresponding word of the second plurality of words comprises:

establishing a lower threshold by negatively modulating the cadence information for each word of the second plurality of words by a predetermined amount;

establishing an upper threshold by positively modulating the cadence information for each word of the second plurality of words by the predetermined amount; and

determining whether the cadence of each word of the plurality of words is between the lower threshold and the upper threshold for the corresponding word of the second plurality of words.

8. The method of claim 3 , wherein determining whether the emphasis of each word of the plurality of words matches the emphasis information for each corresponding word of the second plurality of words comprises:

determining a first plurality of relative emphasis levels corresponding to the relative emphasis between each word of the plurality of words;

determining a second plurality of relative emphasis levels corresponding to the relative emphasis between each word of the second plurality of words; and

determining, for each relative emphasis level of the first plurality of relative emphasis levels, whether the respective relative emphasis level is within a threshold amount of the corresponding relative emphasis level of the second plurality of emphasis levels.

9. The method of claim 4 , wherein determining whether the duration of each word of the plurality of words matches the duration information for each corresponding word of the second plurality of words comprises:

establishing, for each word of the second plurality of words, a lower threshold duration by reducing the duration information by a predetermined amount;

establishing, for each word of the second plurality of words, an upper threshold duration by increasing the duration information by the predetermined amount; and

determining, for each word of the plurality of words, whether the duration of each respective word is between the lower threshold duration and the upper threshold duration for the corresponding word of the second plurality of words.

10. The method of claim 5 , wherein determining whether the overall rhythm of the plurality of words matches the rhythm information for the second plurality of words comprises:

establishing a lower threshold rhythm by negatively modulating the rhythm information by a predetermined amount;

establishing an upper threshold rhythm by positively modulating the rhythm information by the predetermined amount; and

determining whether the overall rhythm of the plurality of words is between the lower threshold rhythm and the upper threshold rhythm.

11. A system for disambiguating a voice search query, the system comprising:

input circuitry configured to receive a voice search query containing a command keyword; and

control circuitry configured to:

transcribe the voice search query into a string comprising a plurality of words;

determine an audio signature of the voice search query;

query a database with the string;

receive, in response to the query, metadata of a quotation, the metadata comprising the string, audio signature information for the string as spoken in a particular content item, and an identifier of the particular content item;

compare the audio signature of the voice query with the audio signature information in the metadata of the quotation;

determine, based on the comparing, whether the voice search query mimics a voice that produced the quotation; and

in response to determining that the voice search query mimics a voice that produced the quotation, generate for display a search result comprising the identifier of the particular content item.

12. The system of claim 11 , wherein the control circuitry is further configured to:

determine a cadence of each word of the plurality of words;

retrieve metadata of at least one quotation, the metadata comprising a second string that is similar to the string and comprises a second plurality of words, and cadence information for each word of the second plurality of words; and

compare a cadence of each word of the plurality of words with cadence information in the metadata for each corresponding word of the second plurality of words;

wherein the control circuitry configured to determine whether the audio signature matches the audio signature information in the metadata of a quotation is further configured to determine, based on the comparing, whether the cadence of each word of the plurality of words matches the cadence information for each corresponding word of the second plurality of words.

13. The system of claim 11 , wherein the control circuitry is further configured to:

determine an emphasis of each word of the plurality of words;

retrieve metadata of at least one quotation, the metadata comprising a second string that is similar to the string and comprises a second plurality of words, and emphasis information for each word of the second plurality of words; and

compare an emphasis of each word of the plurality of words with emphasis information in the metadata for each corresponding word of the second plurality of words;

wherein the control circuitry configured to determine whether the audio signature matches the audio signature information in the metadata of a quotation is further configured to determine, based on the comparing, whether the emphasis of each word of the plurality of words matches the emphasis information for each corresponding word of the second plurality of words.

14. The system of claim 11 , wherein the control circuitry is further configured to:

determine a duration of each word of the plurality of words;

retrieve metadata of at least one quotation, the metadata comprising a second string that is similar to the string and comprises a second plurality of words, and duration information for each word of the second plurality of words; and

compare a duration of each word of the plurality of words with duration information in the metadata for each corresponding word of the second plurality of words;

wherein the control circuitry configured to determine whether the audio signature matches the audio signature information in the metadata of a quotation is further configured to determine, based on the comparing, whether the duration of each word of the plurality of words matches the duration information for each corresponding word of the second plurality of words.

15. The system of claim 11 , wherein the control circuitry is further configured to:

determine an overall rhythm of the plurality of words;

retrieve metadata of at least one quotation, the metadata comprising a second string that is similar to the string and comprises a second plurality of words, and rhythm information for the second plurality of words; and

compare the overall rhythm of the plurality of words with rhythm information in the metadata for the second plurality of words;

wherein the control circuitry configured to determine whether the audio signature matches the audio signature information in the metadata of a quotation is further configured to determine, based on the comparing, whether the overall rhythm of the plurality of words matches the rhythm information for the second plurality of words.

16. The system of claim 11 , wherein the control circuitry configured to determine whether the audio signature matches audio signature information on the metadata of a quotation is further configured to:

establish a lower threshold by negatively transposing the audio signature information in the metadata by a predetermined amount;

establish an upper threshold by positively transposing the audio signature information in the metadata by the predetermined amount; and

determine whether the audio signature is between the lower threshold and the upper threshold.

17. The system of claim 12 , wherein the control circuitry configured to determine whether the cadence of each word of the plurality of words matches the cadence information for each corresponding word of the second plurality of words is further configured to:

establish a lower threshold by negatively modulating the cadence information for each word of the second plurality of words by a predetermined amount;

establish an upper threshold by positively modulating the cadence information for each word of the second plurality of words by the predetermined amount; and

determine whether the cadence of each word of the plurality of words is between the lower threshold and the upper threshold for the corresponding word of the second plurality of words.

18. The system of claim 13 , wherein the control circuitry configured to determine whether the emphasis of each word of the plurality of words matches the emphasis information for each corresponding word of the second plurality of words is further configured to:

determine a first plurality of relative emphasis levels corresponding to the relative emphasis between each word of the plurality of words;

determine a second plurality of relative emphasis levels corresponding to the relative emphasis between each word of the second plurality of words; and

determine, for each relative emphasis level of the first plurality of relative emphasis levels, whether the respective relative emphasis level is within a threshold amount of the corresponding relative emphasis level of the second plurality of emphasis levels.

19. The system of claim 14 , wherein the control circuitry configured to determine whether the duration of each word of the plurality of words matches the duration information for each corresponding word of the second plurality of words is further configured to:

establish, for each word of the second plurality of words, a lower threshold duration by reducing the duration information by a predetermined amount;

establish, for each word of the second plurality of words, an upper threshold duration by increasing the duration information by the predetermined amount; and

determine, for each word of the plurality of words, whether the duration of each respective word is between the lower threshold duration and the upper threshold duration for the corresponding word of the second plurality of words.

20. The system of claim 15 , wherein the control circuitry configured to determine whether the overall rhythm of the plurality of words matches the rhythm information for the second plurality of words is further configured to:

establish a lower threshold rhythm by negatively modulating the rhythm information by a predetermined amount;

establish an upper threshold rhythm by positively modulating the rhythm information by the predetermined amount; and

determine whether the overall rhythm of the plurality of words is between the lower threshold rhythm and the upper threshold rhythm.

Assignments (7)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0231 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2019
From: AHER, ANKUR; SRI, SINDHUJA CHONAT; PUNIYANI, AMAN; MAHAJAN, NISHCHIT
To: ROVI GUIDES, INC.
Reel/Frame 049060/0280 →
Continuity (1)
Related Publication 20200342859A1 · Oct 29, 2020
Cited By (7)
US 12,328,497 US 12,393,627 US 12,395,369 US 12,406,665 US 12,580,988 US 12,670,909 US 12,694,046