IP Library Granted Patent US 11,626,113
Granted Patent B2
US 11,626,113 · App. 17/412,924 · Granted Apr 11, 2023

Systems and methods for disambiguating a voice search query

Inventors: Ankur Aher (Maharashtra, IN); Sindhuja Chonat Sri (Tamil Nadu, IN); Aman Puniyani (Karnataka, IN); Nishchit Mahajan (Punjab, IN)
Assignee: Rovi Guides, Inc.
G10L15/22G06F16/635G06F16/638G06F16/683G10L15/08G10L25/51G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,626,113
App. No.
17/412,924
Granted
Apr 11, 2023
Kind
B2
Abstract

Systems and methods are described herein for disambiguating a voice search query that contains a command keyword by determining whether the user spoke a quotation from a content item and whether the user mimicked or approximated the way the quotation is spoken in the content item. The voice search query is transcribed into a string, and an audio signature of the voice search query is identified. Metadata of a quotation matching the string is retrieved from a database that includes audio signature information for the string as spoken within the content item. The audio signature of the voice search query is compared with the audio signature information in the metadata to determine whether the audio signature matches the audio signature information in the quotation metadata. If a match is detected, then a search result comprising an identifier of the content item from which the quotation comes is generated.

Claims (70)

1. A method comprising:

determining an audio signature of a voice search query, wherein the audio signature comprises a plurality of audio characteristics;

retrieving, in response to the voice search query, metadata of a quotation, wherein the metadata comprises a plurality of audio characteristics characterized by audio signature information corresponding to the quotation and an identifier of a content item comprising the quotation;

comparing the audio signature of the voice search query to the audio signature information in the metadata of the quotation;

determining, based on the comparing, whether a portion of the plurality of audio characteristics of the audio signature of the voice search query match a corresponding portion of the plurality of audio characteristics characterized by the audio signature information; and

in response to determining the portion of the plurality of audio characteristics of the audio signature of the voice search query matches the corresponding portion of the plurality of audio characteristics characterized by the audio signature information, generating for display a search result comprising the identifier of the content item comprising the quotation.

2. The method of claim 1 , wherein the audio signature comprises a plurality of audio characteristics and a plurality of corresponding values for each of the plurality of audio characteristics.

3. The method of claim 2 , wherein the plurality of audio characteristics comprise at least one of a cadence, a spacing of each word spoken relative to other words, a rhythm, and a duration of each word spoken.

4. The method of claim 1 , wherein the plurality of audio characteristics characterized by audio signature information comprises at least one of a cadence, a spacing of each word spoken relative to other words, a rhythm, and a duration of each word spoken.

5. The method of claim 1 , further comprising:

retrieving a similarity threshold value for each of the plurality of audio characteristics, wherein the similarity threshold value corresponds to an allowable difference between a value for each of the plurality of audio characteristics of the audio signature of the voice search query and a value for each of the plurality of audio characteristics characterized by audio signature information corresponding to the quotation;

computing a difference between a value corresponding to the audio signature of the voice search query to a value corresponding to the audio signature information in the metadata of the quotation;

comparing a difference between values corresponding to each of the characteristic the audio signature of the voice search query to a corresponding characteristic of the audio signature information in the metadata of the quotation; and

in response to determining a difference between values corresponding to each of the characteristic the audio signature of the voice search query to the corresponding characteristic of the audio signature information in the metadata of the quotation is the same as or is below the similarity threshold value, determining, based on the comparing, that the audio signature of the voice search query matches the audio signature information corresponding to the quotation.

6. The method of claim 5 , further comprising:

establishing a lower threshold by negatively transposing the audio signature information in the metadata by a predetermined amount;

establishing an upper threshold by positively transposing the audio signature information in the metadata by the predetermined amount; and

determining whether the audio signature of the voice search query is between the lower threshold and the upper threshold.

7. The method of claim 1 , wherein determining, based on the comparing, whether the portion of the plurality of audio characteristics of the audio signature of the voice search query match the corresponding portion of the plurality of audio characteristics characterized by the audio signature information comprises:

establishing a lower threshold by negatively modulating cadence information for each word of the voice search query by a predetermined amount;

establishing an upper threshold by positively modulating the cadence information for each word of the voice search query by the predetermined amount; and

determining whether the cadence of each word of the voice search query is between the lower threshold and the upper threshold for a corresponding word of the quotation.

8. The method of claim 1 , wherein determining, based on the comparing, whether the portion of the plurality of audio characteristics of the audio signature of the voice search query match the corresponding portion of the plurality of audio characteristics characterized by the audio signature information comprises:

determining a first plurality of relative emphasis levels corresponding to a relative emphasis between each word of the voice search query;

determining a second plurality of relative emphasis levels corresponding to the relative emphasis between each word of the quotation; and

determining, for each relative emphasis level of the first plurality of relative emphasis levels, whether the respective relative emphasis level is within a threshold amount of the corresponding relative emphasis level of the second plurality of emphasis levels.

9. The method of claim 1 , wherein determining, based on the comparing, whether the portion of the plurality of audio characteristics of the audio signature of the voice search query match the corresponding portion of the plurality of audio characteristics characterized by the audio signature information comprises:

establishing, for each word of the quotation, a lower threshold duration by reducing the duration information by a predetermined amount;

establishing, for each word of the quotation, an upper threshold duration by increasing the duration information by the predetermined amount; and

determining, for each word of the voice search query, whether the duration of each respective word is between the lower threshold duration and the upper threshold duration for the corresponding word of the quotation.

10. The method of claim 1 , wherein determining, based on the comparing, whether the portion of the plurality of audio characteristics of the audio signature of the voice search query match the corresponding portion of the plurality of audio characteristics characterized by the audio signature information comprises:

establishing a lower threshold rhythm by negatively modulating rhythm information corresponding to the quotation by a predetermined amount;

establishing an upper threshold rhythm by positively modulating the rhythm information by the predetermined amount; and

determining whether the overall rhythm of the voice search query is between the lower threshold rhythm and the upper threshold rhythm.

11. A system comprising:

memory; and

control circuitry configured to:

determine an audio signature of a voice search query, wherein the audio signature comprises a plurality of audio characteristics;

retrieve, in response to the voice search query, metadata of a quotation, wherein the metadata comprises a plurality of audio characteristics characterized by audio signature information corresponding to the quotation and an identifier of a content item comprising the quotation;

compare the audio signature of the voice search query to the audio signature information in the metadata of the quotation;

determine, based on the comparing, whether a portion of the plurality of audio characteristics of the audio signature of the voice search query match a corresponding portion of the plurality of audio characteristics characterized by the audio signature information; and

in response to determining the portion of the plurality of audio characteristics of the audio signature of the voice search query matches the corresponding portion of the plurality of audio characteristics characterized by the audio signature information, generate for display a search result comprising the identifier of the content item comprising the quotation.

12. The system of claim 11 , wherein the control circuitry is configured to determine that the audio signature comprises a plurality of audio characteristics and a plurality of corresponding values for each of the plurality of audio characteristics.

13. The system of claim 12 , wherein the control circuitry is further configured to determine the plurality of audio characteristics comprise at least one of a cadence, a spacing of each word spoken relative to other words, a rhythm, and a duration of each word spoken.

14. The system of claim 11 , wherein the control circuitry is further configured to determine the plurality of audio characteristics characterized by audio signature information comprises at least one of a cadence, a spacing of each word spoken relative to other words, a rhythm, and a duration of each word spoken.

15. The system of claim 11 , wherein the control circuitry is further configured to:

retrieve a similarity threshold value for each of the plurality of audio characteristics, wherein the similarity threshold value corresponds to an allowable difference between a value for each of the plurality of audio characteristics of the audio signature of the voice search query and a value for each of the plurality of audio characteristics characterized by audio signature information corresponding to the quotation;

compute a difference between a value corresponding to the audio signature of the voice search query to a value corresponding to the audio signature information in the metadata of the quotation;

compare a difference between values corresponding to each of the characteristic the audio signature of the voice search query to a corresponding characteristic of the audio signature information in the metadata of the quotation; and

in response to determining a difference between values corresponding to each of the characteristic the audio signature of the voice search query to the corresponding characteristic of the audio signature information in the metadata of the quotation is the same as or is below the similarity threshold value, determine, based on the comparing, that the audio signature of the voice search query matches the audio signature information corresponding to the quotation.

16. The system of claim 15 , wherein the control circuitry is further configured to:

establish a lower threshold by negatively transposing the audio signature information in the metadata by a predetermined amount;

establish an upper threshold by positively transposing the audio signature information in the metadata by the predetermined amount; and

determine whether the audio signature of the voice search query is between the lower threshold and the upper threshold.

17. The system of claim 11 , wherein the control circuitry configured to determine, based on the comparing, whether the portion of the plurality of audio characteristics of the audio signature of the voice search query match the corresponding portion of the plurality of audio characteristics characterized by the audio signature information is further configured to:

establish a lower threshold by negatively modulating cadence information for each word of the voice search query by a predetermined amount;

establish an upper threshold by positively modulating the cadence information for each word of the voice search query by the predetermined amount; and

determine whether the cadence of each word of the voice search query is between the lower threshold and the upper threshold for a corresponding word of the quotation.

18. The system of claim 11 , wherein the control circuitry configured to determine, based on the comparing, whether the portion of the plurality of audio characteristics of the audio signature of the voice search query match the corresponding portion of the plurality of audio characteristics characterized by the audio signature information is further configured to:

determine a first plurality of relative emphasis levels corresponding to a relative emphasis between each word of the voice search query;

determine a second plurality of relative emphasis levels corresponding to the relative emphasis between each word of the quotation; and

determine, for each relative emphasis level of the first plurality of relative emphasis levels, whether the respective relative emphasis level is within a threshold amount of the corresponding relative emphasis level of the second plurality of emphasis levels.

19. The system of claim 11 , wherein the control circuitry configured to determine, based on the comparing, whether the portion of the plurality of audio characteristics of the audio signature of the voice search query match the corresponding portion of the plurality of audio characteristics characterized by the audio signature information is further configured to:

establish, for each word of the quotation, a lower threshold duration by reducing the duration information by a predetermined amount;

establish, for each word of the quotation, an upper threshold duration by increasing the duration information by the predetermined amount; and

determine, for each word of the voice search query, whether the duration of each respective word is between the lower threshold duration and the upper threshold duration for the corresponding word of the quotation.

20. The system of claim 11 , wherein the control circuitry configured to determine, based on the comparing, whether the portion of the plurality of audio characteristics of the audio signature of the voice search query match the corresponding portion of the plurality of audio characteristics characterized by the audio signature information is further configured to:

establish a lower threshold rhythm by negatively modulating rhythm information corresponding to the quotation by a predetermined amount;

establish an upper threshold rhythm by positively modulating the rhythm information by the predetermined amount; and

determine whether the overall rhythm of the voice search query is between the lower threshold rhythm and the upper threshold rhythm.

Assignments (3)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0231 →
SECURITY INTEREST Recorded May 19, 2023
From: ADEIA GUIDES INC.; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063707/0884 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2021
From: AHER, ANKUR; CHONAT SRI, SINDHUJA; PUNIYANI, AMAN; MAHAJAN, NISHCHIT
To: ROVI GUIDES, INC.
Reel/Frame 057334/0611 →
Continuity (2)
Continuation 16397004 · Apr 29, 2019
Related Publication 20210390954A1 · Dec 16, 2021
Cited By (1)
US 12,670,909