IP Library Granted Patent US 11,494,434
Granted Patent B2
US 11,494,434 · App. 16/528,541 · Granted Nov 8, 2022

Systems and methods for managing voice queries using pronunciation information

Inventors: Ankur Aher (Kalyan, IN); Indranil Coomar Doss (Bengaluru, IN); Aashish Goyal (Bengaluru, IN); Aman Puniyani (Rohtak, IN); Kandala Reddy (Bangalore, IN); Mithun Umesh (Bangalore, IN)
Assignee: ROVI GUIDES, INC.
G06F16/635G06F16/24578G06F16/632G06F16/686G06F40/295G10L15/187G10L15/22G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,494,434
App. No.
16/528,541
Granted
Nov 8, 2022
Kind
B2
Abstract

The system receives a voice query at an audio interface and converts the voice query to text. The system can determine pronunciation information during conversion and generate metadata the indicates a pronunciation of one or more words of the query, include phonetic information in the text query, or both. A query includes one or more entities, which may be more accurately identified based on pronunciation. The system searches for information, content, or both among one or more databases based on the generated text query, pronunciation information, user profile information, search histories or trends, and optionally other information. The system identifies one or more entities or content items that match the text query, and retrieves the identified information to provide to the user.

Claims (45)

1. A method for responding to voice queries, the method comprising:

receiving a voice query at an audio interface;

extracting, using control circuitry, one or more keywords from the voice query;

generating, using the control circuitry, a text query based on the one or more keywords;

identifying an entity based on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of the entity based on pronunciation of an identifier associated with the entity, and wherein identifying the entity comprises:

identifying a plurality of entities, wherein respective metadata is stored for each entity of the plurality of entities;

determining a respective score for each respective entity of the plurality of entities based on comparing the respective one or more alternate text representations with the text query; and

selecting the entity by determining a maximum score; and

retrieving a content item associated with the entity.

2. The method of claim 1 , wherein the one more alternate text representations comprise a phonetic representation of the entity.

3. The method of claim 1 , wherein the one more alternate text representations comprise an alternate spelling of the entity based on pronunciation.

4. The method of claim 1 , wherein the one or more alternate text representations of the entity comprise a text string generated based on a previous speech-to-text conversion.

5. The method of claim 1 , wherein the one or more alternate text representations comprise a plurality of alternate text representations, and wherein each alternate text representation of the plurality of alternate text representations is generated by:

converting a first text representation to an audio file; and

converting the audio file to a second text representation, wherein the second text representation is not identical to the first text representation.

6. The method of claim 1 , wherein identifying the entity is further based on user profile information.

7. The method of claim 1 , wherein identifying the entity is further based on popularity information associated with the entity.

8. The method of claim 1 , further comprising generating a plurality of text queries, wherein the plurality of text queries comprises the text query, and wherein each text query of the plurality of text queries is generated based on a respective setting of a speech-to-text module of the control circuitry.

9. The method of claim 8 , further comprising:

identifying, based on a respective text query of the plurality of text queries, a respective entity;

determining a respective score for the respective entity based on a comparison of the respective text query to metadata associated with the respective entity; and

identifying the entity by selecting a maximum score of the respective scores.

10. A system for responding to voice queries, the system comprising:

an audio interface for receiving a voice query;

control circuitry configured to:

extract one or more keywords from the voice query;

generate a text query based on the one or more keywords;

identify an entity based on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of the entity based on pronunciation of an identifier associated with the entity, and wherein the control circuitry is further configured to identify the entity by:

identifying a plurality of entities, wherein respective metadata is stored for each entity of the plurality of entities,

determining a respective score for each respective entity of the plurality of entities based on comparing the respective one or more alternate text representations with the text query; and

selecting the entity by determining a maximum score; and

retrieve a content item associated with the entity.

11. The system of claim 10 , wherein the one more alternate text representations comprise a phonetic representation of the entity.

12. The system of claim 10 , wherein the one more alternate text representations comprise an alternate spelling of the entity based on pronunciation.

13. The system of claim 10 , wherein the one or more alternate text representations of the entity comprise a text string generated based on a previous speech-to-text conversion.

14. The system of claim 10 , wherein the one or more alternate text representations comprise a plurality of alternate text representations, and wherein the control circuitry is configured to generate each alternate text representation of the plurality of alternate text representations by:

converting a first text representation to an audio file; and

converting the audio file to a second text representation, wherein the second text representation is not identical to the first text representation.

15. The system of claim 10 , wherein the control circuitry is further configured to identify the entity based on user profile information.

16. The system of claim 10 , wherein the control circuitry is further configured to identify the entity based on popularity information associated with the entity.

17. The system of claim 10 , wherein the control circuitry is further configured to generate a plurality of text queries, wherein the plurality of text queries comprises the text query, wherein the control circuitry comprises a speech-to-text module, and wherein each text query of the plurality of text queries is generated based on a respective setting of a speech-to-text module.

18. The system of claim 17 , wherein the control circuitry is further configured to:

identify, based on a respective text query of the plurality of text queries, a respective entity;

determine a respective score for the respective entity based on a comparison of the respective text query to metadata associated with the respective entity; and

identify the entity by selecting a maximum score of the respective scores.

Assignments (7)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0231 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2020
From: AHER, ANKUR; DOSS, INDRANIL COOMAR; GOYAL, AASHISH; PUNIYANI, AMAN; REDDY, KANDALA; UMESH, MITHUN
To: ROVI GUIDES, INC.
Reel/Frame 051688/0844 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
Continuity (1)
Related Publication 20210034663A1 · Feb 4, 2021