IP Library Granted Patent US 12,332,937
Granted Patent B2
US 12,332,937 · App. 16/528,539 · Granted Jun 17, 2025

Systems and methods for managing voice queries using pronunciation information

Inventors: Ankur Aher (Kalyan, IN); Indranil Coomar Doss (Bengaluru, IN); Aashish Goyal (Bengaluru, IN); Aman Puniyani (Rohtak, IN); Kandala Reddy (Bangalore, IN); Mithun Umesh (Bangalore, IN)
Assignee: Adeia Guides Inc.
G06F16/635G06F16/686G06F40/295G10L15/187G10L15/22G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,332,937
App. No.
16/528,539
Filed
Jul 31, 2019
Granted
Jun 17, 2025
Kind
B2
Examiner
HU, XIAOQIN
Art Unit
2168
USPC
707/771
Abstract

The system receives a voice query at an audio interface and converts the voice query to text. The system can determine pronunciation information during conversion and generate metadata that indicates a pronunciation of one or more words of the query, include phonetic information in the text query, or both. A query includes one or more entities that may be more accurately identified based on pronunciation. The system searches for information, content, or both among one or more databases based on the generated text query, pronunciation information, user profile information, search histories or trends, and optionally other information. The system identifies one or more entities or content items that match the text query, and retrieves the identified information to provide to the user.

Claims (44)

1. A method for responding to voice queries, the method comprising:

generating, using a text-to-speech engine, an audio output based on a first text string describing an entity of a plurality of entities of a database;

transcribing, using one or more speech-to-text engines, the audio output to generate a second text string;

wherein the text-to-speech engine, the one or more speech-to-text engines, or both the text-to-speech engine and the one or more speech-to-text engines include one or more settings with which the audio output is generated, the one or more settings comprising: a language used to generate the audio output, an accent used to generate the audio output, a voice of a particular person used to generate the audio output, a gender used to generate the audio output, a playback speed of the audio output, a phonetic variation of the first text string, or any combination thereof;

storing the second text string as metadata for the entity when the second text string does not exactly match the first text string;

repeating the steps of generating, transcribing, and storing with a varying combination of settings of the one or more settings to generate one or more variations of the second text string to be associated with the entity;

receiving, via audio interface circuitry, a voice query input at an audio interface, wherein the second text string is stored as the metadata for the entity prior to the receiving of the voice query input at the audio interface;

transcribing, using the one or more speech-to-text engines, the voice query input from an electronic signal to an input text string;

identifying the entity of the plurality of entities of the database by executing a text-to-text search of the plurality of entities, wherein the text-to-text search of the plurality of entities identifies the entity by determining a match of the input text string with the first text string or the second text string stored in the metadata for the entity;

retrieving a content item associated with the entity; and

generating by control circuitry, based on the content item retrieved, an output comprising a search result responsive to the voice query input.

2. The method of claim 1 , wherein the identifying the entity is further based on user profile information.

3. The method of claim 2 , wherein the identifying the entity is based on a previously identified entity from a previous voice query input.

4. The method of claim 1 , wherein the identifying the entity is further based on popularity information associated with the entity.

5. The method of claim 1 , wherein the identifying the entity comprises:

identifying the plurality of entities, wherein respective metadata is stored for each entity of the plurality of entities,

determining a respective score for each respective entity of the plurality of entities based on comparing the respective metadata with the input text string; and

selecting the entity by determining a maximum score.

6. The method of claim 1 , wherein the entity is a first entity, and further comprising:

identifying a second entity among the plurality of entities based on the input text string and second metadata for the second entity, and wherein the content item is associated with the first entity and the second entity.

7. The method of claim 1 , further comprising:

determining that two or more textual representations exist for a selected keyword of the input text string;

determining pronunciation information for the selected keyword in response to the determining that two or more textual representations exist for the selected keyword.

8. A system for responding to voice queries, the system comprising:

an audio interface comprising audio interface circuitry configured to receive a voice query input; and

control circuitry coupled to the audio interface, the control circuitry configured to:

generate, using a text-to-speech engine, an audio output based on a first text string describing an entity of a plurality of entities of a database;

transcribe, using one or more speech-to-text engines, the audio output to generate a second text string;

wherein the text-to-speech engine, the one or more speech-to-text engines, or both the text-to-speech engine and the one or more speech-to-text engines include one or more settings with which the audio output is generated, the one or more settings comprising: a language used to generate the audio output, an accent used to generate the audio output, a voice of a particular person used to generate the audio output, a gender used to generate the audio output, a playback speed of the audio output, a phonetic variation of the first text string, or any combination thereof;

store the second text string as metadata for the entity when the second text string does not exactly match the first text string, wherein the second text string is stored as the metadata for the entity prior to the receiving of the voice query input at the audio interface;

repeating the steps of generate, transcribe, and store with a varying combination of settings of the one or more settings to generate one or more variations of the second text string to be associated with the entity;

transcribe, using the one or more speech-to-text engines, the voice query input from an electronic signal to an input text string;

identify the entity of the plurality of entities of the database by executing a text-to-text search of the plurality of entities, wherein the text-to-text search of the plurality of entities identifies the entity by determining a match of the input text string with the first text string or the second text string stored in the metadata for the entity;

retrieve a content item associated with the entity; and

generating by the control circuitry, based on the content item retrieved, an output comprising a search result responsive to the voice query input.

9. The system of claim 8 , wherein the control circuitry is further configured to identify the entity based on user profile information.

10. The system of claim 9 , wherein the control circuitry is further configured to identify the entity based on a previously identified entity from a previous voice query input.

11. The system of claim 8 , wherein the control circuitry is further configured to identify the entity based on popularity information associated with the entity.

12. The system of claim 8 , wherein the control circuitry is further configured to identify the entity by:

identifying the plurality of entities, wherein respective metadata is stored for each entity of the plurality of entities,

determining a respective score for each respective entity of the plurality of entities based on comparing respective metadata with the input text string; and

selecting the entity by determining a maximum score.

13. The system of claim 8 , wherein the entity is a first entity, wherein the control circuitry is further configured:

to identify a second entity among the plurality of entities based on the input text string and second metadata for the second entity, and wherein the content item is associated with the first entity and the second entity.

Assignments (7)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0231 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2020
From: AHER, ANKUR; DOSS, INDRANIL COOMAR; GOYAL, AASHISH; PUNIYANI, AMAN; REDDY, KANDALA; UMESH, MITHUN
To: ROVI GUIDES, INC.
Reel/Frame 051688/0835 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
Continuity (1)
Related Publication 20210034662A1 · Feb 4, 2021
References Cited (49)
US 5737485A · Flanagan et al. · 1998 [cited by applicant]
US 7181395B1 · Deligne et al. · 2007 [cited by applicant]
US 8239944B1 · Nachenberg et al. · 2012 [cited by applicant]
US 8423565B2 · Redlich et al. · 2013 [cited by applicant]
US 8630860B1 · Zhang et al. · 2014 [cited by applicant]
US 9043199B1 · Hayes · 2015 [cited by applicant]
US 9098551B1 · Fryz et al. · 2015 [cited by applicant]
US 9715877B2 · Sims, III · 2017 [cited by examiner]
US 9812120B2 · Takatsuka · 2017 [cited by applicant]
US 11157696B1 · Ramos · 2021 [cited by examiner]
US 11410656B2 · Aher et al. · 2022 [cited by applicant]
US 20110167053A1 · Lawler et al. · 2011 [cited by applicant]
US 20130132374A1 · Olstad · 2013 [cited by examiner]
US 20140359523A1 · Jang · 2014 [cited by examiner]
US 20150142812A1 · Ma · 2015 [cited by applicant]
US 20150269672A1 · Bhuyan · 2015 [cited by examiner]
US 20150371636A1 · Raedel et al. · 2015 [cited by applicant]
US 20160034458A1 · Choi · 2016 [cited by examiner]
US 20160098493A1 · Primke et al. · 2016 [cited by applicant]
US 20160180840A1 · Siddiq et al. · 2016 [cited by applicant]
US 20170068423A1 · Napolitano et al. · 2017 [cited by applicant]
US 20170147576A1 · Des Jardins · 2017 [cited by applicant]
US 20170264939A1 · Jang et al. · 2017 [cited by applicant]
US 20180024901A1 · Tankersley et al. · 2018 [cited by applicant]
US 20180166073A1 · Gandiga · 2018 [cited by applicant]
US 20190036856A1 · Bergenlid et al. · 2019 [cited by applicant]
US 20190037357A1 · Bijor et al. · 2019 [cited by applicant]
US 20190149987A1 · Moore · 2019 [cited by applicant]
US 20190295527A1 · Pore · 2019 [cited by examiner]
US 20190333499A1 · Li et al. · 2019 [cited by applicant]
US 20190384821A1 · Alders et al. · 2019 [cited by applicant]
US 20200042514A1 · Svonja · 2020 [cited by applicant]
US 20200175968A1 · Donati et al. · 2020 [cited by applicant]
US 20200184958A1 · Norouzi et al. · 2020 [cited by applicant]
US 20200193975A1 · Blau-Mccandliss et al. · 2020 [cited by applicant]
US 20200357390A1 · Bromand · 2020 [cited by applicant]
US 20200365136A1 · Candelore et al. · 2020 [cited by applicant]
US 20210011934A1 · Boxwell et al. · 2021 [cited by applicant]
US 20210026901A1 · Aher et al. · 2021 [cited by applicant]
US 20210034663A1 · Aher et al. · 2021 [cited by applicant]
US 20210035587A1 · Aher et al. · 2021 [cited by applicant]
US 20210241754A1 · Hiroya · 2021 [cited by applicant]
JP 2015526797A · 2015 [cited by applicant]
JP 2017010514A · 2017 [cited by applicant]
JP 2019032876A · 2019 [cited by applicant]
WO 2016167992A1 · 2016 [cited by applicant]
PCT International Search Report for International Application No. PCT/US2020/043131, dated Oct. 21, 2020 (14 pages). [cited by applicant]
U.S. Appl. No. 16/528,541, filed Jul. 31, 2019, Ankur Aher. [cited by applicant]
U.S. Appl. No. 16/528,550, filed Jul. 31, 2019, Ankur Aher. [cited by applicant]
Cited By (1)
US 12,609,995