IP Library Granted Patent US 12670909
Granted Patent B2
US 12670909 · App. 18/828,772 · Granted Jun 30, 2026

Systems and methods for disambiguating a voice search query

Inventors: Ankur Aher (Kalyan, IN); Sindhuja Chonat Sri (Coimbatore, IN); Aman Puniyani (Bangalore, IN); Nishchit Mahajan (Amritsar, IN)
Assignee: Adeia Guides Inc.
G10L15/22G06F16/635G06F16/638G06F16/683G10L15/08G10L25/51G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670909
App. No.
18/828,772
Granted
Jun 30, 2026
Kind
B2
Abstract

Systems and methods are described herein for disambiguating a voice search query that contains a command keyword by determining whether the user spoke a quotation from a content item and whether the user mimicked or approximated the way the quotation is spoken in the content item. The voice search query is transcribed into a string, and an audio signature of the voice search query is identified. Metadata of a quotation matching the string is retrieved from a database that includes audio signature information for the string as spoken within the content item. The audio signature of the voice search query is compared with the audio signature information in the metadata to determine whether the audio signature matches the audio signature information in the quotation metadata. If a match is detected, then a search result comprising an identifier of the content item from which the quotation comes is generated.

Claims (44)

1 . A method comprising:

receiving a voice search query comprising a first plurality of words;

determining a first plurality of audio parameters for the voice search query, wherein a first audio parameter of the first plurality of audio parameters is associated with at least one word of the first plurality of words;

identifying metadata of a quotation based on the voice search query, wherein:

the quotation comprises a second plurality of words;

the metadata comprises a second plurality of audio parameters for the quotation; and

a first audio parameter of the second plurality of audio parameters is associated with at least one word of the second plurality of words;

determining that at least a portion of the voice search query matches at least a portion of the quotation in response to comparing the first plurality of audio parameters with the second plurality of audio parameters; and

generating for display a search result comprising an identifier of a content item associated with the quotation.

2 . The method of claim 1 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a volume level.

3 . The method of claim 1 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a duration.

4 . The method of claim 1 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to an emphasis level.

5 . The method of claim 1 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a tone.

6 . The method of claim 1 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a speed measurement.

7 . An apparatus, comprising:

control circuitry; and

at least one memory including computer program code for one or more programs, the at least one memory and the computer program code configured to, with the control circuitry, cause the apparatus to perform at least the following:

receive a voice search query comprising a first plurality of words;

determine a first plurality of audio parameters for the voice search query, wherein a first audio parameter of the first plurality of audio parameters is associated with at least one word of the first plurality of words;

identify metadata of a quotation based on the voice search query, wherein:

the quotation comprises a second plurality of words;

the metadata comprises a second plurality of audio parameters for the quotation; and

a first audio parameter of the second plurality of audio parameters is associated with at least one word of the second plurality of words;

determine that at least a portion of the voice search query matches at least a portion of the quotation in response to comparing the first plurality of audio parameters with the second plurality of audio parameters; and

generate for display a search result comprising an identifier of a content item associated with the quotation.

8 . The apparatus of claim 7 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a volume level.

9 . The apparatus of claim 7 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a duration.

10 . The apparatus of claim 7 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to an emphasis level.

11 . The apparatus of claim 7 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a tone.

12 . The apparatus of claim 7 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a speed measurement.

13 . A non-transitory computer-readable medium having instructions encoded thereon that, when executed by control circuitry, cause the control circuitry to:

receive a voice search query comprising a first plurality of words;

determine a first plurality of audio parameters for the voice search query, wherein a first audio parameter of the first plurality of audio parameters is associated with at least one word of the first plurality of words;

identify metadata of a quotation based on the voice search query, wherein:

the quotation comprises a second plurality of words;

the metadata comprises a second plurality of audio parameters for the quotation; and

a first audio parameter of the second plurality of audio parameters is associated with at least one word of the second plurality of words;

determine that at least a portion of the voice search query matches at least a portion of the quotation in response to comparing the first plurality of audio parameters with the second plurality of audio parameters; and

generate for display a search result comprising an identifier of a content item associated with the quotation.

14 . The non-transitory computer-readable medium of claim 13 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a volume level.

15 . The non-transitory computer-readable medium of claim 13 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a duration.

16 . The non-transitory computer-readable medium of claim 13 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to an emphasis level.

17 . The non-transitory computer-readable medium of claim 13 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a tone.

18 . The non-transitory computer-readable medium of claim 13 , wherein one or more of the first plurality of audio parameters and one or more of the second plurality of audio parameters correspond to a speed measurement.