IP Library Granted Patent US 12,524,461
Granted Patent B2
US 12,524,461 · App. 18/239,401 · Granted Jan 13, 2026

Systems and methods for using conjunctions in a voice input to cause a search application to wait for additional inputs

Inventors: Susanto Sen (Karnataka, IN); Charishma Chundi (Andhra Pradesh, IN)
Assignee: Adeia Guides Inc.
G06F16/632G06F16/433G06F16/438G06V10/80G06V40/20G06V40/28G10L15/19G10L15/22G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,461
App. No.
18/239,401
Granted
Jan 13, 2026
Kind
B2
Abstract

A search is performed based on a voice input combined with user selection of entities displayed on a display screen as well as real-world entities. A voice input is received from the user by a media device, as well as a selection of a first entity being displayed on the media device. A conjunction spoken in the voice input triggers the media device to wait for selection of a second entity before performing the search. After receiving selection of the second entity, a search query is constructed based on the voice input, the first entity, and the second entity. The search query is transmitted to a database and, in response, the media device receives at least one identifier of a least one content item. The at least one identifier is then generated for display to the user.

Claims (95)

1 . A computer-implemented method, comprising:

receiving input from a user via a user input interface of a media device;

processing the input to identify a particular pronoun;

identifying a gesture made by the user;

processing an image associated with the gesture to determine a plurality of identities of a plurality of entities, respectively, in the image, displayed on a display of the media device;

determining, based on the plurality of identities, a respective pronoun for each entity of the plurality of entities;

determining, from among the plurality of entities, a particular entity having a pronoun that corresponds to the particular pronoun identified based on the input, wherein the particular entity is displayed on the display of the media device;

querying a database based on the input and the particular entity;

based on querying the database, receiving at least one identifier of at least one content item; and

generating for presentation, using the media device, the at least one identifier of the at least one content item.

2 . The method of claim 1 , wherein the input is a voice input, the particular entity is a second entity, and the method further comprises:

receiving, at the media device, a selection of a first entity currently being displayed on the display of the media device; and

processing the voice input to identify a search operator,

wherein querying the database comprises constructing a search query based on the identified search operator, the first entity and the second entity.

3 . The method of claim 1 , wherein the image is captured by an imaging sensor, and identifying the gesture made by the user comprises determining, based on the image, motion of the user.

4 . The method of claim 1 , wherein the image is captured by an imaging sensor, and determining the particular entity associated with the gesture further comprises:

determining a direction of the gesture; and

determining that the direction of the gesture corresponds to the image, the image depicting a real-world scene proximate to the user.

5 . The method of claim 4 , wherein determining the particular entity associated with the gesture further comprises:

performing image processing of the image to determine a plurality of portions of the image that respectively correspond to the plurality of entities;

based on the direction of the gesture, extrapolating a path of the gesture to a portion of the image; and

determining as the particular entity an entity of the plurality of entities associated with the portion of the image that the path intersects.

6 . The method of claim 4 , wherein the imaging sensor is a first imaging sensor, the image captured by the first imaging sensor is a first image, and the method further comprises:

determining that a second image captured by a second imaging sensor depicts a different perspective of the real-world scene than a perspective of the real-world scene depicted in the first image;

extrapolating a first path from the direction of the gesture in the first image;

extrapolating a second path from the direction of the gesture in the second image;

identifying a point at which the first path crosses the second path; and

determining as the particular entity an entity of the plurality of entities associated with the point of the image at which the first path crosses the second path.

7 . The method of claim 4 , wherein the imaging sensor is a first imaging sensor, the image captured by the first imaging sensor is a first image, and the method further comprises:

determining that a second image, captured by a second imaging sensor that is facing a second direction, depicts an area in which the user made the gesture;

performing image processing of the second image to identify the gesture;

extrapolating a first path from the direction of the gesture in the second image;

calculating, based on a position and an angle of the first imaging sensor and a position and an angle of the second imaging sensor, a second path in the first image corresponding to the first path;

performing image processing of the first image to determine a plurality of portions of the first image that respectively correspond to the plurality of entities; and

determining as the particular entity an entity of the plurality of entities associated with the portion of the image that the second path intersects.

8 . The method of claim 1 , wherein the media device is a first media device, the method further comprising:

generating, for presentation at a second media device proximate to the first media device, a content item,

wherein the image associated with the gesture corresponds to a frame of the content item being presented at the second media device.

9 . The method of claim 8 , wherein the particular entity is a first entity, the content item is a first content item, and the method further comprises:

generating, for presentation at the first media device while the first content item is being presented at the second media device, a second content item,

wherein the input is a voice input, and the voice input comprises a reference to a second entity depicted in a frame of the second media device, the querying of the database being further based on the second entity.

10 . The method of claim 1 , wherein determining, based on the plurality of identities, the respective pronoun for each entity of the plurality of entities comprises:

determining a first pronoun for a first entity of the plurality of entities; and

determining a second pronoun for a second entity of the plurality of entities;

wherein determining the particular entity having a pronoun that corresponds to the particular pronoun is based on the first pronoun for the first entity of the plurality of entities and the second pronoun for a second entity of the plurality of entities.

11 . A computer-implemented system, comprising:

input/output (I/O) circuitry configured to:

receive input from a user via a user input interface of a media device;

control circuitry configured to:

process the input to identify a particular pronoun;

identify a gesture made by the user;

process an image associated with the gesture to determine a plurality of identities of a plurality of entities, respectively, in the image, the image displayed on a display of the media device;

determine, based on the plurality of identities, a respective pronoun for each entity of the plurality of entities;

determine, from among the plurality of entities, a particular entity having a pronoun that corresponds to the particular pronoun identified based on the input, wherein the particular entity is displayed on the display of the media device; and

query a database based on the input and the particular entity,

wherein the I/O circuitry is further configured to:

based on querying the database, receive at least one identifier of at least one content item; and

generate for presentation, using the media device, the at least one identifier of the at least one content item.

12 . The system of claim 11 , wherein:

the input is a voice input, the particular entity is a second entity;

the I/O circuitry is further configured to receive, at the media device, a selection of a first entity currently being displayed on the display of the media device; and

the control circuitry is further configured to:

process the voice input to identify a search operator; and

query the database by constructing a search query based on the identified search operator, the first entity and the second entity.

13 . The system of claim 11 , wherein the image is captured by an imaging sensor, and the control circuitry is configured to identify the gesture made by the user by determining, based on the image, motion of the user.

14 . The system of claim 11 , wherein the image is captured by an imaging sensor, and the control circuitry is configured to determine the particular entity associated with the gesture by:

determining a direction of the gesture; and

determining that the direction of the gesture corresponds to the image, the image depicting a real-world scene proximate to the user.

15 . The system of claim 14 , wherein the control circuitry is configured to determine the particular entity associated with the gesture by:

performing image processing of the image to determine a plurality of portions of the image that respectively correspond to the plurality of entities;

based on the direction of the gesture, extrapolating a path of the gesture to a portion of the image; and

determining as the particular entity an entity of the plurality of entities associated with the portion of the image that the path intersects.

16 . The system of claim 14 , wherein the imaging sensor is a first imaging sensor, the image captured by the first imaging sensor is a first image, and the control circuitry is further configured to:

determine that a second image captured by a second imaging sensor depicts a different perspective of the real-world scene than a perspective of the real-world scene depicted in the first image;

extrapolate a first path from the direction of the gesture in the first image;

extrapolate a second path from the direction of the gesture in the second image;

identify a point at which the first path crosses the second path; and

determine as the particular entity an entity of the plurality of entities associated with the point of the image at which the first path crosses the second path.

17 . The system of claim 14 , wherein the imaging sensor is a first imaging sensor, the image captured by the first imaging sensor is a first image, and the control circuitry is further configured to:

determine that a second image, captured by a second imaging sensor that is facing a second direction, depicts an area in which the user made the gesture;

perform image processing of the second image to identify the gesture;

extrapolate a first path from the direction of the gesture in the second image;

calculate, based on a position and an angle of the first imaging sensor and a position and an angle of the second imaging sensor, a second path in the first image corresponding to the first path;

perform image processing of the first image to determine a plurality of portions of the first image that respectively correspond to the plurality of entities; and

determine as the particular entity an entity of the plurality of entities associated with the portion of the image that the second path intersects.

18 . The system of claim 11 , wherein the media device is a first media device, and the control circuitry is further configured to:

generate, for presentation at a second media device proximate to the first media device, a content item,

wherein the image associated with the gesture corresponds to a frame of the content item being presented at the second media device.

19 . The system of claim 18 , wherein the particular entity is a first entity, the content item is a first content item, and the control circuitry is further configured to:

generate, for presentation at the first media device while the first content item is being presented at the second media device, a second content item,

wherein the input is a voice input, and the voice input comprises a reference to a second entity depicted in a frame of the second media device, the querying of the database being further based on the second entity.

20 . The system of claim 11 , wherein the control circuitry is configured, when determining the respective pronoun for each entity of the plurality of entities, to:

determine a first pronoun for a first entity of the plurality of entities; and

determine a second pronoun for a second entity of the plurality of entities;

wherein determining the particular entity having a pronoun that corresponds to the particular pronoun is based on the first pronoun for the first entity of the plurality of entities and the second pronoun for a second entity of the plurality of entities.

Assignments (2)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0238 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2023
From: SEN, SUSANTO; CHUNDI, CHARISHMA
To: ROVI GUIDES, INC.
Reel/Frame 064738/0169 →
Continuity (3)
Continuation 17750532 · May 23, 2022
Continuation 16736076 · Jan 7, 2020
Related Publication 20230409632A1 · Dec 21, 2023
References Cited (33)
US 5715468A · Budzinski · 1998 [cited by applicant]
US 6685477B1 · Goldman et al. · 2004 [cited by applicant]
US 7069215B1 · Bangalore · 2006 [cited by examiner]
US 9031840B2 · Sharifi · 2015 [cited by examiner]
US 9256396B2 · Monson et al. · 2016 [cited by applicant]
US 9390726B1 · Smus · 2016 [cited by examiner]
US 10515121B1 · Setlur et al. · 2019 [cited by applicant]
US 10542114B2 · Cunico et al. · 2020 [cited by applicant]
US 10802673B2 · Kim et al. · 2020 [cited by applicant]
US 10845956B2 · Chand · 2020 [cited by applicant]
US 11367444B2 · Sen et al. · 2022 [cited by applicant]
US 11604830B2 · Sen et al. · 2023 [cited by applicant]
US 20120115112A1 · Purushotma et al. · 2012 [cited by applicant]
US 20120197857A1 · Huang et al. · 2012 [cited by applicant]
US 20130218896A1 · Palay · 2013 [cited by applicant]
US 20140214415A1 · Klein · 2014 [cited by examiner]
US 20150348430A1 · Kasbar · 2015 [cited by applicant]
US 20160109954A1 · Harris et al. · 2016 [cited by applicant]
US 20160124706A1 · Vasilieff et al. · 2016 [cited by applicant]
US 20160170710A1 · Kim · 2016 [cited by examiner]
US 20160239259A1 · Lenchner · 2016 [cited by examiner]
US 20170068423A1 · Napolitano et al. · 2017 [cited by applicant]
US 20170083214A1 · Furesjöet al. · 2017 [cited by applicant]
US 20180046851A1 · Kienzle et al. · 2018 [cited by applicant]
US 20180136834A1 · Tumwattana · 2018 [cited by applicant]
US 20200258515A1 · Suzuki · 2020 [cited by examiner]
US 20210118442A1 · Poddar · 2021 [cited by examiner]
US 20210168293A1 · Yang · 2021 [cited by examiner]
US 20210209171A1 · Sen et al. · 2021 [cited by applicant]
US 20210210084A1 · Sen et al. · 2021 [cited by applicant]
US 20210295833A1 · Rastrow et al. · 2021 [cited by applicant]
US 20230011143A1 · Sen et al. · 2023 [cited by applicant]
PCT International Search Report for International Application No. PCT/US2020/065540, dated Apr. 9, 2021 (15 pages). [cited by applicant]