IP Library Granted Patent US 12,505,153
Granted Patent B2
US 12,505,153 · App. 17/951,921 · Granted Dec 23, 2025

Method and system for voice based media search

Inventors: Mukesh Patel (Fremont, CA); Lu Silverstein (Alviso, CA); Srinivas Jandhyala (Alviso, CA)
G06F16/48G06F16/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,153
App. No.
17/951,921
Granted
Dec 23, 2025
Kind
B2
Abstract

Voice-based input is used to operate a media device and/or to search for media content. Voice input is received by a media device via one or more audio input devices and is translated into a textual representation of the voice input. The textual representation of the voice input is used to search one or more cache mappings between input commands and one or more associated device actions and/or media content queries. One or more natural language processing techniques may be applied to the translated text and the resulting text may be transmitted as a query to a media search service. A media search service returns results comprising one or more content item listings and the results may be presented on a display to a user.

Claims (49)

1 . A method comprising:

receiving an audio signal at a server;

generating non-textual data based at least in part on the received audio signal;

processing the non-textual data to determine whether the non-textual data matches keyword data; and

based at least in part on determining that the non-textual data matches the keyword data:

processing, at the server, the received audio signal by performing a speech-to-text translation of the received audio signal;

determining whether the speech-to-text translation of the audio signal corresponds to an electronic device action; and

based at least in part on determining that the speech-to-text translation of the audio signal corresponds to the electronic device action, executing the corresponding electronic device action.

2 . The method of claim 1 , further comprising:

receiving an audio input at an electronic device; and

receiving the audio signal at the server, wherein the audio signal corresponds to the audio input received at the electronic device, and wherein the audio signal was transmitted from the electronic device to the server.

3 . The method of claim 2 , wherein the electronic device is a separate device from the server and is communicatively connected to the server via a network.

4 . The method of claim 1 , wherein the audio signal includes a first portion and a second portion.

5 . The method of claim 4 , wherein the first portion is related to the keyword data and the second portion is related to the electronic device action.

6 . The method of claim 1 , wherein the speech-to-text translation is based at least in part on a user profile.

7 . The method of claim 1 , further comprising:

receiving one or more voice samples at the server; and

training the server for performing the speech-to-text translation by using the one or more received voice samples.

8 . The method of claim 1 , further comprising, not performing a speech-to-text translation of the audio signal based at least in part on determining that the non-textual data does not match the keyword data.

9 . The method of claim 1 , further comprising:

performing natural language processing on the translated speech-to-text translation;

determining a context of the based at least in part on the natural language processing; and

generating a modified textual representation of the translated speech-to-text translation based at least in part on the determined context.

10 . The method of claim 9 , further comprising, determining the electronic device action based at least in part on the modified textual representation.

11 . A system comprising:

control circuitry configured to:

receive an audio signal at a server;

generate non-textual data based at least in part on the received audio signal;

process the non-textual data to determine whether the non-textual data matches keyword data; and

based at least in part on determining that the non-textual data matches keyword data:

process, at the server, the received audio signal by performing a speech-to-text translation of the received audio signal;

determine whether the speech-to-text translation of the audio signal corresponds to an electronic device action; and

based at least in part on determining that the speech-to-text translation of the audio signal corresponds to the electronic device action, execute the corresponding electronic device action.

12 . The system of claim 11 , wherein the control circuitry is further configured to:

receive an audio input at an electronic device; and

receive the audio signal at the server, wherein the audio signal corresponds to the audio input received at the electronic device, and wherein the audio signal was transmitted from the electronic device to the server.

13 . The system of claim 12 , wherein the electronic device is a separate device from the server and is communicatively connected to the server via a network.

14 . The system of claim 11 , wherein the audio signal includes a first portion and a second portion.

15 . The system of claim 14 , wherein the first portion is related to the keyword data and the second portion is related to the electronic device action.

16 . The system of claim 11 , wherein the speech-to-text translation is based at least in part on a user profile.

17 . The system of claim 11 , wherein the control circuitry is further configured to:

receive one or more voice samples at the server; and

train the server for performing the speech-to-text translation by using the one or more received voice samples.

18 . The system of claim 11 , wherein the control circuitry is further configured to not perform the speech-to-text translation of the audio signal based at least in part on determining that the non-textual data does not match the keyword data.

19 . The system of claim 11 , wherein the control circuitry is further configured to:

perform natural language processing on the translated speech-to-text translation;

determine a context of the based at least in part on the natural language processing; and

generate a modified textual representation of the translated speech-to-text translation based at least in part on the determined context.

20 . The system of claim 19 , wherein the control circuitry is further configured to determine the electronic device action based at least in part on the modified textual representation.

Assignments (4)
CHANGE OF NAME Recorded Sep 27, 2024
From: TIVO SOLUTIONS INC.
To: ADEIA MEDIA SOLUTIONS INC.
Reel/Frame 069067/0504 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2022
From: PATEL, MUKESH; SILVERSTEIN, LU; JANDHYALA, SRINIVAS
To: TIVO INC.
Reel/Frame 061242/0858 →
CHANGE OF NAME Recorded Sep 28, 2022
From: TIVO INC.
To: TIVO SOLUTIONS INC.
Reel/Frame 061242/0888 →