IP Library Granted Patent US 11,410,656
Granted Patent B2
US 11,410,656 · App. 16/528,550 · Granted Aug 9, 2022

Systems and methods for managing voice queries using pronunciation information

Inventors: Ankur Aher (Kalyan, IN); Indranil Coomar Doss (Bengaluru, IN); Aashish Goyal (Bengaluru, IN); Aman Puniyani (Rohtak, IN); Kandala Reddy (Bangalore, IN); Mithun Umesh (Bangalore, IN)
Assignee: ROVI GUIDES, INC.
G10L15/26G10L13/02G10L15/187G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,410,656
App. No.
16/528,550
Granted
Aug 9, 2022
Kind
B2
Abstract

The system identifies one or more entities or content items among a plurality of stored information. The system generates an audio file based on a first text string that represents the entity or content item. Based on the first text string and at least one speech criterion, the system generating, using a speech-to-text module a second text string based on the audio file. The system then compares the text strings and stores the second text string if it is not identical to the first text string. The system generates metadata that includes results from text-speech-text conversions to forecast possible misidentifications when responding to voice queries during search operations. The metadata includes alternative representations of the entity.

Claims (55)

1. A method for generating entity metadata for voice queries, the method comprising:

receiving a search query comprising a plurality of entities;

identifying an entity of the plurality of entities;

generating, using a text-to-speech module, an audio file comprising speech content based on a first text string and at least one speech criterion, wherein the first text string describes the entity, and wherein the at least one speech criterion comprises a pronunciation setting for generating the speech content;

generating, using a speech-to-text module, a second text string based on the speech content in the audio file;

comparing the second text string to the first text string;

in response to determining that the second text string is not identical to the first text string, storing, in metadata associated with the entity, the second text string and pronunciation information, based on the pronunciation setting, for the second text string; and

disambiguating a subsequent search query comprising the entity based on the metadata and on the search query.

2. The method of claim 1 , wherein the at least one speech criterion comprises a language setting.

3. The method of claim 1 , wherein the at least one speech criterion comprises a plurality of speech criterion, the method further comprising:

generating, using the text-to-speech module, a respective audio file based on a first text string and a respective speech criterion;

generating, using the speech-to-text module, a respective second text string based on the respective audio file;

comparing the respective second text string to the first text string; and

storing, in metadata associated with the entity, the respective second text string if it is not identical to the first text string.

4. The method of claim 1 , further comprising updating the metadata based on one or more text queries.

5. The method of claim 1 , further comprising storing, in metadata associated with the entity, a phonetic representation of the first text string.

6. The method of claim 1 , wherein generating the audio file based on the first text string comprises:

converting the first text string to a first audio signal;

generating speech at a speaker based on the audio signal;

detecting the speech using a microphone to generate a second audio signal; and

processing the audio signal to generate the audio file.

7. The method of claim 6 , wherein generating the speech at the speaker is further based on at least one speech setting of the text-to-speech module.

8. The method of claim 1 , wherein generating the second text string based on the audio file comprises:

generating a playback of the audio file at a speaker;

detecting the playback using a microphone to generate an audio signal; and

converting the audio signal to the second text string by identifying one or more words.

9. The method of claim 8 , wherein converting the audio signal to the second text string is based on at least one text setting of the speech-to-text module.

10. A system for generating entity metadata for voice queries, the system comprising:

control circuitry configured to:

receive a search query comprising a plurality of entities;

identify an entity of the plurality of entities;

generate an audio file, using an audio interface coupled to the control circuitry, comprising speech content based on a first text string and at least one speech criterion, wherein the first text string describes the entity, and wherein the at least one speech criterion comprises a pronunciation setting for generating the speech content;

generate, using the audio interface, a second text string based on the speech content in the audio file;

compare the second text string to the first text string;

in response to determining that the second text string is not identical to the first text string, store, in metadata associated with the entity, the second text string and pronunciation information, based on the pronunciation setting, for the second text string; and

disambiguate a subsequent search query comprising the entity based on the metadata and on the search query.

11. The system of claim 10 , wherein the at least one speech criterion comprises a language setting.

12. The system of claim 10 , wherein the at least one speech criterion comprises a plurality of speech criterion, and wherein the control circuitry is further configured to:

generate, using the audio equipment, a respective audio file based on a first text string and a respective speech criterion;

generating, using the audio equipment, a respective second text string based on the respective audio file;

compare the respective second text string to the first text string; and

store, in metadata associated with the entity, the respective second text string if it is not identical to the first text string.

13. The system of claim 10 , wherein the control circuitry is further configured to update the metadata based on one or more text queries.

14. The system of claim 10 , wherein the control circuitry is further configured to store, in metadata associated with the entity, a phonetic representation of the first text string.

15. The system of claim 10 , wherein the audio equipment comprises a speaker and a microphone, and wherein the control circuitry is further configured to generate the audio file based on the first text string by:

converting the first text string to a first audio signal;

generating speech at the speaker based on the audio signal;

detecting the speech using the microphone to generate a second audio signal; and

processing the audio signal to generate the audio file.

16. The system of claim 15 , wherein the control circuitry is further configured to generate the speech at the speaker based on at least one speech setting.

17. The system of claim 10 , wherein the audio equipment comprises a speaker and a microphone, and wherein the control circuitry is further configured to generate the second text string based on the audio file by:

generating a playback of the audio file at the speaker;

detecting the playback at the microphone to generate an audio signal; and

converting the audio signal to the second text string by identifying one or more words.

18. The system of claim 17 , wherein the control circuitry is further configured to convert the audio signal to the second text string based on at least one text setting of the speech-to-text module.

Assignments (7)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0231 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2020
From: AHER, ANKUR; DOSS, INDRANIL COOMAR; GOYAL, AASHISH; PUNIYANI, AMAN; REDDY, KANDALA; UMESH, MITHUN
To: ROVI GUIDES, INC.
Reel/Frame 051688/0897 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
Continuity (1)
Related Publication 20210035587A1 · Feb 4, 2021
Cited By (2)
US 12,332,937 US 12,603,079