IP Library Granted Patent US 12,475,162
Granted Patent B2
US 12,475,162 · App. 17/951,962 · Granted Nov 18, 2025

Method and system for voice based media search

Inventors: Mukesh Patel (Fremont, CA); Lu Silverstein (Alviso, CA); Srinivas Jandhyala (Alviso, CA)
Assignee: Adeia Media Solutions Inc.
G06F16/48G06F16/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,162
App. No.
17/951,962
Granted
Nov 18, 2025
Kind
B2
Abstract

Voice-based input is used to operate a media device and/or to search for media content. Voice input is received by a media device via one or more audio input devices and is translated into a textual representation of the voice input. The textual representation of the voice input is used to search one or more cache mappings between input commands and one or more associated device actions and/or media content queries. One or more natural language processing techniques may be applied to the translated text and the resulting text may be transmitted as a query to a media search service. A media search service returns results comprising one or more content item listings and the results may be presented on a display to a user.

Claims (45)

1 . A method comprising:

storing personalized command word data;

receiving an audio signal;

generating non-textual data representing the received audio signal;

processing the generated non-textual data, based at least in part on the audio signal, wherein the processing comprises comparing the generated non-textual data with the personalized command word data; and

based at least in part on the processing:

generating a textual representation of the received audio signal, wherein the textual representation is generated based at least in part on a speech-to-text transformation of the audio signal;

determining whether the textual representation corresponds to an electronic device action; and

based at least in part on determining that the textual representation corresponds to the electronic device action, executing a command of the electronic device action that corresponds to the textual representation.

2 . The method of claim 1 , wherein comparing the generated non-textual data with the personalized command word data further comprises:

receiving a voice input;

generating a voice profile based at least in part on the received voice input;

determining the personalized command word data based at least in part on the voice profile; and

comparing the generated non-textual data with the determined personalized command word data.

3 . The method of claim 2 , wherein the voice input received is based at least in part on a user speaking into a microphone of a electronic device.

4 . The method of claim 1 , further comprising determining a spatial location of a user based at least in part on the audio signal received.

5 . The method of claim 4 , wherein the spatial location is determined by applying a beam forming technique.

6 . The method of claim 4 , further comprising, identifying the user based at least in part on a closest match between the determined spatial location and a voice input received subsequent to receiving the audio signal.

7 . The method of claim 1 , further comprising filtering out ambient noise while the audio signal is being received.

8 . The method of claim 7 , wherein the ambient noise is generated by an audio output of a electronic device.

9 . The method of claim 1 , further comprising receiving a second audio signal subsequent to receiving the audio signal.

10 . The method of claim 9 , further comprising associating back-to-back receipt of the audio signal and the second audio signal with multiple electronic device actions.

11 . The method of claim 9 , wherein the audio signal is received from a first user and the second audio signal is received from a second user.

12 . The method of claim 9 , wherein the audio signal and the second audio signal is received from a same user.

13 . A system comprising:

control circuitry configured to:

store personalized command word data;

receive an audio signal;

generate non-textual data representing the received audio signal;

process the generated non-textual data, based at least in part on the audio signal, wherein the processing comprises comparing the generated non-textual data with the personalized command word data; and

based at least in part on the processing:

generate a textual representation of the received audio signal, wherein the textual representation is generated based at least in part on a speech-to-text transformation of the audio signal;

determine whether the textual representation corresponds to an electronic device action; and

based at least in part on determining that the textual representation corresponds to the electronic device action, execute a command of the electronic device action that corresponds to the textual representation.

14 . The system of claim 13 , wherein comparing the generated non-textual data with the personalized command word data further comprises, the control circuitry configured to:

receive a voice input;

generate a voice profile based at least in part on the received voice input;

determine the personalized command word data based at least in part on the voice profile; and

compare the generated non-textual data with the determined personalized command word data.

15 . The system of claim 14 , wherein the voice input received is based at least in part on a user speaking into a microphone of a electronic device.

16 . The system of claim 13 , further comprising, the control circuitry configured to determine a spatial location of a user based at least in part on the audio signal received.

17 . The system of claim 16 , further comprising, the control circuitry configured to identify the user based at least in part on a closest match between the determined spatial location and a voice input received subsequent to receiving the audio signal.

18 . The system of claim 13 , further comprising, the control circuitry configured to filter out ambient noise while the audio signal is being received.

19 . The system of claim 13 , further comprising, the control circuitry configured to receive a second audio signal subsequent to receiving the audio signal.

20 . The system of claim 19 , further comprising, the control circuitry configured to associate back-to-back receipt of the audio signal and the second audio signal with multiple electronic device actions.

Assignments (4)
CHANGE OF NAME Recorded Sep 27, 2024
From: TIVO SOLUTIONS INC.
To: ADEIA MEDIA SOLUTIONS INC.
Reel/Frame 069067/0504 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2022
From: PATEL, MUKESH; SILVERSTEIN, LU; JANDHYALA, SRINIVAS
To: TIVO INC.
Reel/Frame 061242/0858 →
CHANGE OF NAME Recorded Sep 28, 2022
From: TIVO INC.
To: TIVO SOLUTIONS INC.
Reel/Frame 061242/0888 →
Continuity (6)
Continuation 17481831 · Sep 22, 2021
Continuation 16265932 · Feb 1, 2019
Continuation 15949754 · Apr 10, 2018
Continuation 15645526 · Jul 10, 2017
Continuation 13665735 · Oct 31, 2012
Related Publication 20230016510A1 · Jan 19, 2023
References Cited (44)
US 6374214B1 · Friedland et al. · 2002 [cited by applicant]
US 6823493B2 · Baker · 2004 [cited by applicant]
US 8428227B2 · Angel et al. · 2013 [cited by applicant]
US 8564544B2 · Jobs et al. · 2013 [cited by applicant]
US 8875021B2 · Lee et al. · 2014 [cited by applicant]
US 9734151B2 · Patel · 2017 [cited by examiner]
US 9971772B2 · Patel · 2018 [cited by examiner]
US 10242005B2 · Patel · 2019 [cited by examiner]
US 11151184B2 · Patel · 2021 [cited by examiner]
US 20040102959A1 · Estrin · 2004 [cited by applicant]
US 20040220926A1 · Lamkin et al. · 2004 [cited by applicant]
US 20070150275A1 · Garner et al. · 2007 [cited by applicant]
US 20080005688A1 · Najdenovski · 2008 [cited by applicant]
US 20090228277A1 · Bonforte et al. · 2009 [cited by applicant]
US 20090299752A1 · Rodriguez et al. · 2009 [cited by applicant]
US 20090326938A1 · Marila et al. · 2009 [cited by applicant]
US 20110098917A1 · Lebeau et al. · 2011 [cited by applicant]
US 20110173539A1 · Rottler et al. · 2011 [cited by applicant]
US 20110208524A1 · Haughay · 2011 [cited by applicant]
US 20110286584A1 · Angel et al. · 2011 [cited by applicant]
US 20120016678A1 · Gruber et al. · 2012 [cited by applicant]
US 20120035924A1 · Jitkoff et al. · 2012 [cited by applicant]
US 20120201362A1 · Crossan et al. · 2012 [cited by applicant]
US 20120245936A1 · Treglia · 2012 [cited by applicant]
US 20120259924A1 · Patil et al. · 2012 [cited by applicant]
US 20130332168A1 · Kim et al. · 2013 [cited by applicant]
US 20140115465A1 · Lee et al. · 2014 [cited by applicant]
US 20140122059A1 · Patel et al. · 2014 [cited by applicant]
US 20180025001A1 · Patel et al. · 2018 [cited by applicant]
US 20180053507A1 · Wang et al. · 2018 [cited by applicant]
US 20180232368A1 · Patel et al. · 2018 [cited by applicant]
US 20190027131A1 · Zajac · 2019 [cited by applicant]
US 20190080685A1 · Johnson · 2019 [cited by applicant]
US 20190206405A1 · Gillespie et al. · 2019 [cited by applicant]
US 20190236089A1 · Patel et al. · 2019 [cited by applicant]
US 20220012275A1 · Patel et al. · 2022 [cited by applicant]
US 20230012940A1 · Patel · 2023 [cited by examiner]
US 20230016510A1 · Patel · 2023 [cited by examiner]
US 20230017928A1 · Patel et al. · 2023 [cited by applicant]
WO 2010119288A1 · 2010 [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US13/67602, mailed on Jan. 30, 2014, 7 pages. [cited by applicant]
Takuo Henmi, Shengyang Huang and Fuji Ren, “Wisdom media “CAIWA Channel” based on natural language interface agent,” Proceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineeri… [cited by applicant]
U.S. Appl. No. 17/951,905, filed Sep. 23, 2022, Mukesh Patel. [cited by applicant]
U.S. Appl. No. 17/951,921, filed Sep. 23, 2022, Mukesh Patel. [cited by applicant]