IP Library Granted Patent US 9,721,564
Granted Patent B2
US 9,721,564 · App. 14/448,308 · Granted Aug 1, 2017

Systems and methods for performing ASR in the presence of heterographs

Inventors: Akshat Agarwal (Delhi, IN); Rakesh Barve (Bangalore, IN)
Assignee: Rovi Guides, Inc.
G10L15/187G10L15/193G10L15/1815
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,721,564
App. No.
14/448,308
Filed
Jul 31, 2014
Granted
Aug 1, 2017
Kind
B2
Examiner
YANG, QIAN
Art Unit
2674
USPC
704/254
Abstract

Systems and methods for performing ASR in the presence of heterographs are provided. Verbal input is received from the user that includes a plurality of utterances. A first of the plurality of utterances is matched to a first word. It is determined that a second utterance in the plurality of utterances matches a plurality of words that is in a same heterograph set. It is identified which one of the plurality of words is associated with a context of the first word. A function is performed based on the first word and the identified one of the plurality of words.

Claims (53)

1. A method for performing automatic speech recognition (ASR) when a heterographic word is present, the method comprising:

receiving verbal input from a user that comprises a plurality of utterances;

matching a first of the plurality of utterances to a first word;

determining a word that describes the context for the first word;

determining that a second utterance in the plurality of utterances matches a plurality of words that are in a same heterograph set;

combining a second word chosen from the plurality of words with the word that describes the context for the first word to generate a first combined set of words;

storing a first value representing a distance between words in the first combined set of words;

combining a third word chosen from the plurality of words with the word that describes the context for the first word to generate a second combined set of words;

storing a second value representing a distance between words in the second combined set of words;

in response to determining that the second value is smaller than the first value, performing a media guidance application function on an available media asset based on the second combined set of words.

2. The method of claim 1 further comprising:

storing a knowledge graph of a relationship between words, wherein a distance between words in the knowledge graph is indicative of strength in relationship between the words; and

calculating the first value and the second value based on the distance between the words in the first combined set of words and the distance between the words in the second combined set of words.

3. The method of claim 2 further comprising:

identifying positions, in the knowledge graph, of the context of the first word and each of the plurality of words; and

computing, based on the identified positions, a distance between the context of the first word and each of the plurality of words.

4. The method of claim 1 , wherein the first word is a name of a competitor in a sporting event, further comprising:

setting the context to be the sporting event; and

determining which of the plurality of words corresponds to the sporting event, wherein the third word corresponds to another competitor in the sporting event.

5. The method of claim 1 , wherein the plurality of words that are in the same heterograph set are phonetically similar to each other.

6. The method of claim 1 further comprising generating a recommendation based on the first word and the third word.

7. The method of claim 1 , wherein matching the first of the plurality of utterances to the first word comprises determining that the first utterance phonetically corresponds to the first word.

8. The method of claim 1 , wherein the first word is a name of an actor in a media asset, further comprising:

setting the context to be the media asset; and

determining which of the plurality of words corresponds to the media asset, wherein the third word corresponds to another actor in the media asset.

9. The method of claim 1 further comprising determining the context based on a conjunction between two of the plurality of utterances.

10. A system for performing automatic speech recognition (ASR) when a heterographic word is present, the system comprising:

control circuitry configured to:

receive verbal input from a user that comprises a plurality of utterances;

match a first of the plurality of utterances to a first word;

determine a word that describes the context for the first word;

determine that a second utterance in the plurality of utterances matches a plurality of words that are in a same heterograph set;

combine a second word chosen from the plurality of words with the word that describes the context for the first word to generate a first combined set of words;

store a first value representing a distance between words in the first combined set of words;

combine a third word chosen from the plurality of words with the word that describes the context for the first word to generate a second combined set of words;

store a second value representing a distance between words in the second combined set of words; and

in response to determining that the second value is smaller than the first value, perform a media guidance application function on an available media asset based on the second combined set of words.

11. The system of claim 10 , wherein the control circuitry is further configured to:

store a knowledge graph of a relationship between words, wherein a distance between words in the knowledge graph is indicative of strength in relationship between the words; and

calculate the first value and the second value based on a distance between the words in the first combined set of words and the words in the second combined set of words.

12. The system of claim 11 , wherein the control circuitry is further configured to:

identify positions, in the knowledge graph, of the first word and each of the plurality of words; and

compute, based on the identified positions, a distance between the first word and each of the plurality of words.

13. The system of claim 10 , wherein the first word is a name of a competitor in a sporting event, and wherein the control circuitry is further configured to:

set the context to be the sporting event;

determine which of the plurality of words corresponds to the sporting event, wherein the third word corresponds to another competitor in the sporting event.

14. The system of claim 10 , wherein the plurality of words that are in the same heterograph set are phonetically similar to each other.

15. The system of claim 10 , wherein the control circuitry is further configured to generate a recommendation based on the first word and the third word.

16. The system of claim 10 , wherein the control circuitry is further configured to match the first of the plurality of utterances to the first word by determining that the first utterance phonetically corresponds to the first word.

17. The system of claim 10 , wherein the first word is a name of an actor in a media asset, and wherein the control circuitry is further configured to:

set the context to be the media asset; and

determine which of the plurality of words corresponds to the media asset, wherein the third word corresponds to another actor in the media asset.

18. The system of claim 10 , wherein the control circuitry is further configured to determine the context based on a conjunction between two of the plurality of utterances.

Assignments (10)
CHANGE OF NAME Recorded Oct 2, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069085/0697 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
MERGER Recorded May 28, 2015
From: UNITED VIDEO PROPERTIES, INC.
To: UV CORP.
Reel/Frame 035729/0921 →
MERGER Recorded May 28, 2015
From: TV GUIDE, INC.
To: ROVI GUIDES, INC.
Reel/Frame 035729/0971 →
MERGER Recorded May 28, 2015
From: UV CORP.
To: TV GUIDE, INC.
Reel/Frame 035729/0950 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2014
From: AGARWAL, AKSHAT; BARVE, RAKESH
To: UNITED VIDEO PROPERTIES, INC.
Reel/Frame 033487/0909 →
Continuity (1)
Related Publication 20160035347A1 · Feb 4, 2016