IP Library Granted Patent US 11,849,193
Granted Patent B2
US 11,849,193 · App. 17/969,651 · Granted Dec 19, 2023

Methods, systems, and apparatuses to respond to voice requests to play desired video clips in streamed media based on matched close caption and sub-title text

Inventor: Mayank Verma (Bangalore, IN)
H04N21/4828G06F16/735G06F16/7834G06F16/7867G10L15/183G10L15/1815G10L15/1822G10L15/22G10L15/30H04N21/4394H04N21/47202H04N21/4882G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,849,193
App. No.
17/969,651
Granted
Dec 19, 2023
Kind
B2
Abstract

Methods, Systems, and Apparatuses are described to implement voice search in media content for requesting media content of a video clip of a scene contained in the media content streamed to the client device; for capturing the voice request for the media content of the video clip to display at the client device wherein the streamed media content is a selected video streamed from a video source; for applying a NLP solution to convert the voice request to text for matching to a set of one or more words contained in at least close caption text of the selected video; for associating matched words to close caption text with a start index and an end index of the video clip contained in the selected video; and for streaming the video clip to the client device based on the start index and the end index associated with matched closed caption text.

Claims (34)

1. A method for searching content, comprising:

applying, by a server, natural language processing (NLP) to convert a voice request to text, wherein the voice request is captured by a local device in communication with the server;

creating a first search handler to execute a search in a media database based on the text, wherein the search identifies media content in response to contextual data associated with the media content matching the text;

tagging the media content identified by the first search handler in response to the search; and

playing a video scene at the local device in response to a selection of the video scene from the tagged media content.

2. The method of claim 1 , wherein the contextual data comprises metadata.

3. The method of claim 2 , further comprising creating a second search handler to execute a second search, wherein the second search identifies the media content in response to subtitle data matching the text.

4. The method of claim 3 , further comprising creating a third search handler to execute a third search to identify the media content in response to close caption data associated with the media content matching the text.

5. The method of claim 1 , further comprising creating a searchable index of media content based on the tagged media content.

6. The method of claim 1 , further comprising creating a presentation layer including thumbnail images of scenes identified by the first search handler.

7. The method of claim 1 , further comprising executing a script associated with the video scene in response to the selection of the video scene at the local device.

8. A system for processing voice requests for identifying media content in streaming media for display, comprising:

a local device; and

a server in communication with the local device and configured to perform operations, the operations comprising:

applying natural language processing (NLP) to convert a voice request to text, wherein the voice request is captured by the local device;

creating a first search handler to execute a search in a media database based on the text, wherein the search identifies media content in response to contextual data associated with the media content matching the text;

tagging the media content identified by the first search handler in response to the search; and

playing a video scene at the local device in response to a selection of the video scene from the tagged media content.

9. The system of claim 8 , wherein the operations further comprise creating a second search handler to execute a second search, wherein the second search identifies the media content in response to subtitle data matching the text.

10. The system of claim 9 , wherein the operations further comprise creating a third search handler to execute a third search to identify the media content in response to close caption data associated with the media content matching the text.

11. The system of claim 8 , wherein the operations further comprise creating a searchable index of media content based on the tagged media content.

12. The system of claim 8 , wherein the operations further comprise creating a presentation layer comprising thumbnail images of scenes identified by the first search handler.

13. The system of claim 8 , further comprising executing a script associated with the video scene in response to the selection of the video scene at the local device.

14. The system of claim 8 , wherein the contextual data comprises metadata.

15. A streaming system comprising a processor in communication with a non-transitory memory configured to store instructions that, when executed by the processor, cause the streaming system to perform operations, the operations comprising:

applying natural language processing (NLP) to convert a voice request to text, wherein the voice request is captured by a local device in communication with the streaming system;

creating a first search handler to execute a search in a media database based on the text, wherein the search identifies media content in response to contextual data associated with the media content matching the text;

tagging the media content identified by the first search handler in response to the search; and

playing a video scene at the local device in response to a selection of the video scene from the tagged media content.

16. The streaming system of claim 15 , wherein the contextual data comprises metadata.

17. The streaming system of claim 15 , wherein the contextual data comprises subtitles and closed caption data.

18. The streaming system of claim 15 , wherein the operations further comprise creating a second search handler to execute a second search, wherein the second search identifies the media content in response to subtitle data matching the text.

19. The streaming system of claim 15 , wherein the operations further comprise creating a second search handler to execute a second search to identify the media content in response to close caption data associated with the media content matching the text.

20. The streaming system of claim 15 , wherein the operations further comprise creating a searchable index of media content based on the tagged media content.

Assignments (3)
CHANGE OF NAME Recorded Jul 28, 2026
From: DISH NETWORK TECHNOLOGIES INDIA PRIVATE LIMITED
To: DISH NETWORK TECHNOLOGIES INDIA PRIVATE LIMITED
Reel/Frame 075427/0343 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2026
From: DISH NETWORK TECHNOLOGIES INDIA PRIVATE LIMITED
To: DISH NETWORK L.L.C.
Reel/Frame 075427/0479 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2022
From: VERMA, MAYANK
To: DISH NETWORK TECHNOLOGIES INDIA PRIVATE LIMITED
Reel/Frame 061487/0461 →
Continuity (3)
Continuation 17329679 · May 25, 2021
Continuation 16791347 · Feb 14, 2020
Related Publication 20230037744A1 · Feb 9, 2023
Cited By (1)
US 12,425,694