IP Library Granted Patent US 11,151,184
Granted Patent B2
US 11,151,184 · App. 16/265,932 · Granted Oct 19, 2021

Method and system for voice based media search

Inventors: Mukesh Patel (Fremont, CA); Lu Silverstein (Alviso, CA); Srinivas Jandhyala (Alviso, CA)
Assignee: TiVo Solutions Inc.
G06F16/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,151,184
App. No.
16/265,932
Granted
Oct 19, 2021
Kind
B2
Abstract

Voice-based input is used to operate a media device and/or to search for media content. Voice input is received by a media device via one or more audio input devices and is translated into a textual representation of the voice input. The textual representation of the voice input is used to search one or more cache mappings between input commands and one or more associated device actions and/or media content queries. One or more natural language processing techniques may be applied to the translated text and the resulting text may be transmitted as a query to a media search service. A media search service returns results comprising one or more content item listings and the results may be presented on a display to a user.

Claims (74)

1. A method comprising:

receiving voice input data at a media device;

transducing the voice input data into non-textual digital data;

determining whether the voice input corresponds to one of a plurality of commands by analyzing the transduced non-textual digital data;

in response to determining that the voice input does not correspond to one of the plurality of commands, determining a textual representation of at least a portion of the voice input data;

generating a signature based on at least a portion of the textual representation;

identifying a particular data entry of a set of data entries that corresponds to the generated signature, each data entry of the set of data entries specifying an association between a given signature and one or more media device actions;

in response to identifying the particular data entry that corresponds to the generated signature, performing a media device action associated with the particular data entry; and

generating a display based on the media device action.

2. The method of claim 1 , wherein the textual representation of the at least a portion of the voice input data is a textual representation of an entire voice input data.

3. The method of claim 1 , wherein generating the signature based on the at least a portion of the textual representation comprises generating the signature based on the entire textual representation.

4. The method of claim 1 , further comprising:

determining whether an unknown word exists in the at least a portion of the textual representation;

generating for display the at least a portion of the textual representation; and

in response to determining that the unknown word exists in the at least a portion of the textual representation, visually distinguishing the unknown word from other words.

5. The method of claim 1 , further comprising:

generating for display the at least a portion of the textual representation;

receiving a user selection of a portion of the at least a portion of the textual representation;

receiving additional voice input data;

transmitting the additional voice input data to the speech-to-text service;

receiving from the speech-to-text service, an additional textual representation of the additional voice input data; and

replacing the selected portion of the at least a portion of the textual representation with the additional textual representation.

6. The method of claim 1 , wherein identifying the particular data entry of the set of data entries that corresponds to the generated signature further comprises determining that the particular data entry matches the generated signature based on a respective probability assigned to each of a plurality of data entries of the set of data entries.

7. The method of claim 1 , further comprising:

determining whether an extraneous portion of the at least a portion of the textual representation exists; and

in response to determining that the extraneous portion exists, removing the extraneous portion from the at least a portion of the textual representation, wherein the generated signature is not based on the removed extraneous portion.

8. A system comprising:

control circuitry configured to:

receive voice input data at a media device;

transduce the voice input data into non-textual digital data;

determine whether the voice input corresponds to one of a plurality of commands by analyzing the transduced non-textual digital data;

in response to determining that the voice input does not correspond to one of the plurality of commands, determine a textual representation of at least a portion of the voice input data;

generate a signature based on at least a portion of the textual representation;

identify a particular data entry of a set of data entries that corresponds to the generated signature, each data entry of the set of data entries specifying an association between a given signature and one or more media device actions;

in response to the identification of the particular data that corresponds to the generated signature, perform a media device action associated with the particular data entry; and

generate a display based on the media device action.

9. The system of claim 8 , wherein the textual representation of the at least a portion of the voice input data is a textual representation of an entire voice input data.

10. The system of claim 8 , wherein to generate the signature based on the at least a portion of the textual representation the control circuitry is configured to generate the signature based on the entire textual representation.

11. The system of claim 8 , wherein the control circuitry is further configured to:

determine whether an unknown word exists in the at least a portion of the textual representation;

generate for display the at least a portion of the textual representation; and

in response to the determination that the unknown word exists in the at least a portion of the textual representation, visually distinguish the unknown word from other words.

12. The system of claim 8 , wherein the control circuitry is further configured to:

generate for display the at least a portion of the textual representation;

receive a user selection of a portion of the at least a portion of the textual representation;

receive additional voice input data;

transmit the additional voice input data to the speech-to-text service;

receive from the speech-to-text service, an additional textual representation of the additional voice input data; and

replace the selected portion of the at least a portion of the textual representation with the additional textual representation.

13. The system of claim 8 , wherein the identification of the particular data entry of the set of data entries that corresponds to the generated signature further comprises determining that the particular data entry matches the generated signature based on a respective probability assigned to each of a plurality of data entries of the set of data entries.

14. The system of claim 8 , wherein the control circuitry is further configured to:

determine whether an extraneous portion of the at least a portion of the textual representation exists; and

in response to determining that the extraneous portion exists, remove the extraneous portion from the at least a portion of the textual representation, wherein the generated signature is not based on the removed extraneous portion.

15. The method of claim 1 , wherein generating the signature based on at least a portion of the textual representation comprises:

applying a hash function to the textual representation and computing a hash value; and

determining a cache entry based on the computed hash value.

16. The method of claim 1 , wherein transducing the voice input data into non-textual digital data further comprises:

receiving a plurality of voice inputs in close proximity of time to each other;

determining the time stamp of each input received; and

selecting the voice input associated with the earliest time stamp to transduce into non-textual digital data.

17. The method of claim 1 , wherein transducing the voice input data into non-textual digital data further comprises:

receiving a voice input and a media input in close proximity to each other, wherein the media input is associated with ambient noise from a media device;

distinguishing the voice input from the media input; and

selecting the voice input for transducing into non-textual digital data.

18. The method of claim 1 , wherein, in response to determining that the voice input does not correspond to one of the plurality of commands:

comparing the voice input to plurality of sampled voice data; and

selecting a closest matched sampled voice data from the plurality of sampled voice data, wherein the closest matched sampled voice data is mapped to one of the plurality of commands.

19. The system of claim 8 , wherein transducing the voice input data into non-textual digital data further comprises, the control circuitry configured to:

receive a plurality of voice inputs in close proximity of time to each other;

determine the time stamp of each input received; and

select the voice input associated with the earliest time stamp to transduce into non-textual digital data.

20. The system of claim 8 , wherein, in response to determining that the voice input does not correspond to one of the plurality of commands, the control circuitry further configured to:

compare the voice input to plurality of sampled voice data; and

select a closest matched sampled voice data from the plurality of sampled voice data, wherein the closest matched sampled voice data is mapped to one of the plurality of commands.

Assignments (8)
CHANGE OF NAME Recorded Sep 27, 2024
From: TIVO SOLUTIONS INC.
To: ADEIA MEDIA SOLUTIONS INC.
Reel/Frame 069067/0504 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2019
From: PATEL, MUKESH; SILVERSTEIN, LU; JANDHYALA, SRINIVAS
To: TIVO INC.
Reel/Frame 048232/0585 →
CHANGE OF NAME Recorded Feb 4, 2019
From: TIVO INC.
To: TIVO SOLUTIONS INC.
Reel/Frame 048232/0673 →
Cited By (4)
US 12,339,894 US 12,475,162 US 12,505,152 US 12,505,153