IP Library Granted Patent US 11,367,444
Granted Patent B2
US 11,367,444 · App. 16/736,076 · Granted Jun 21, 2022

Systems and methods for using conjunctions in a voice input to cause a search application to wait for additional inputs

Inventors: Susanto Sen (Karnataka, IN); Charishma Chundi (Andhra Pradesh, IN)
Assignee: ROVI GUIDES, INC.
G10L15/22G06F16/433G06F16/438G06V40/20G10L15/19G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,367,444
App. No.
16/736,076
Granted
Jun 21, 2022
Kind
B2
Abstract

A search is performed based on a voice input combined with user selection of entities displayed on a display screen as well as real-world entities. A voice input is received from the user by a media device, as well as a selection of a first entity being displayed on the media device. A conjunction spoken in the voice input triggers the media device to wait for selection of a second entity before performing the search. After receiving selection of the second entity, a search query is constructed based on the voice input, the first entity, and the second entity. The search query is transmitted to a database and, in response, the media device receives at least one identifier of a least one content item. The at least one identifier is then generated for display to the user.

Claims (76)

1. A method for searching for media content, the method comprising:

receiving, at a media device, a selection of a first entity currently being displayed on a display of the media device;

receiving, at the media device, a voice input from a user;

detecting, by processing the voice input, a conjunction;

in response to detecting the conjunction, waiting for a selection of at least one additional entity;

receiving the selection of at least one additional entity;

constructing a search query based on the conjunction, the first entity, and the at least one additional entity;

transmitting the search query to a database;

receiving, in response to transmitting the search query, at least one identifier of at least one content item; and

generating for display, on the media device, the at least one identifier.

2. The method of claim 1 , wherein the conjunction is a coordinating conjunction and wherein constructing the search query further comprises:

determining a type of the coordinating conjunction;

identifying a logical operator corresponding to the type of coordinating conjunction; and

generating a search string comprising the first entity and each additional entity separated by the logical operator.

3. The method of claim 1 , wherein the conjunction is a subordinating conjunction and wherein constructing the search query further comprises:

identifying a search parameter corresponding to the type of subordinating conjunction; and

generating a search string comprising the identified search parameter, the first entity, and the at least one additional entity.

4. The method of claim 1 , further comprising:

processing the voice input to identify a search operator;

wherein the search query is constructed based on the conjunction, the first entity, the at least one additional entity, and the identified search operator.

5. The method of claim 1 , wherein the database is stored at a remote server.

6. The method of claim 1 , wherein the media device is a virtual reality display.

7. The method of claim 1 , wherein receiving the selection of at least one additional entity further comprises:

identifying a gesture made by the user; and

determining an additional entity associated with the gesture, wherein the additional entity is not being displayed on the display of the media device.

8. The method of claim 7 , wherein identifying the gesture made by the user comprises capturing, using a camera, a motion of the user.

9. The method of claim 7 , wherein determining the additional entity associated with the gesture comprises:

determining a direction of the gesture;

capturing, using a camera, an image representing an area corresponding to the direction of the gesture; and

identifying an entity in the image associated with the gesture.

10. The method of claim 9 , further comprising:

processing the voice input to identify a pronoun corresponding to the additional entity;

performing image processing to identify a plurality of entities in the image; and

determining, based on the identity of each respective entity of the plurality of entities, a respective pronoun corresponding to each respective entity of the plurality of entities;

wherein identifying the entity in the image associated with the gesture comprises:

comparing the respective pronoun of each respective entity of the plurality of entities with the identified pronoun; and

selecting an entity of the plurality of entities having a respective pronoun that matches the identified pronoun.

11. A system for searching for media content, the system comprising:

a display; and

control circuitry configured to:

receive, at a media device, a selection of a first entity currently being displayed on the display;

receive, at the media device, a voice input from a user;

detect, by processing the voice input, a conjunction;

in response to detecting the conjunction, wait for a selection of at least one additional entity;

receive the selection of at least one additional entity;

construct a search query based on the conjunction, the first entity, and the at least one additional entity;

transmit the search query to a database;

receive, in response to transmitting the search query, at least one identifier of at least one content item; and

generate for display, on the display, the at least one identifier.

12. The system of claim 11 , wherein the conjunction is a coordinating conjunction and wherein the control circuitry configured to construct the search query is further configured to:

determine a type of the coordinating conjunction;

identify a logical operator corresponding to the type of coordinating conjunction; and

generate a search string comprising the first entity and each additional entity separated by the logical operator.

13. The system of claim 11 , wherein the conjunction is a subordinating conjunction and wherein the control circuitry configured to construct the search query is further configured to:

identify a search parameter corresponding to the type of subordinating conjunction; and

generate a search string comprising the identified search parameter, the first entity, and the at least one additional entity.

14. The system of claim 11 , wherein the control circuitry is further configured to:

process the voice input to identify a search operator;

wherein the search query is constructed based on the conjunction, the first entity, the at least one additional entity, and the identified search operator.

15. The system of claim 11 , wherein the database is stored at a remote server.

16. The system of claim 11 , wherein the display is a virtual reality display.

17. The system of claim 11 , wherein the control circuitry configured to receive the selection of at least one additional entity is further configured to:

identify a gesture made by the user; and

determine an additional entity associated with the gesture, wherein the additional entity is not being displayed on the display of the media device.

18. The system of claim 17 , further comprising a camera, and wherein the control circuitry configured to identify the gesture made by the user is configured to capture, using the camera, a motion of the user.

19. The system of claim 17 , further comprising a camera, and wherein the control circuitry configured to determine the additional entity associated with the gesture is further configured to:

determine a direction of the gesture;

capture, using the camera, an image representing an area corresponding to the direction of the gesture; and

identify an entity in the image associated with the gesture.

20. The system of claim 19 , wherein the control circuitry is further configured to:

process the voice input to identify a pronoun corresponding to the additional entity;

perform image processing to identify a plurality of entities in the image; and

determine, based on the identity of each respective entity of the plurality of entities, a respective pronoun corresponding to each respective entity of the plurality of entities;

wherein the control circuitry configured to identify the entity in the image associated with the gesture is further configured to:

compare the respective pronoun of each respective entity of the plurality of entities with the identified pronoun; and

select an entity of the plurality of entities having a respective pronoun that matches the identified pronoun.

Assignments (3)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0238 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2020
From: SEN, SUSANTO; CHUNDI, CHARISHMA
To: ROVI GUIDES, INC.
Reel/Frame 051447/0533 →