IP Library Granted Patent US 11,032,580
Granted Patent B2
US 11,032,580 · App. 15/845,704 · Granted Jun 8, 2021

Systems and methods for facilitating a personalized viewing experience

Inventors: Swapnil Anil Tilaye (Broomfield, CO); Rima Shah (Thornton, CO)
Assignee: DISH Network L.L.C.
H04N21/21805H04N21/23418H04N21/472H04N21/4728H04N21/8549
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,032,580
App. No.
15/845,704
Granted
Jun 8, 2021
Kind
B2
Abstract

Embodiments are related to processing of a source video stream for generation of a target video stream that includes an object of interest to a viewer. In some embodiments, the target video stream may exclusively or primarily include the performance of the object of interest to the viewer, without including other persons in that video. This allows a viewer to focus on an object of his or her interest and not necessarily have to view the performances of other objects in the source video stream.

Claims (61)

1. A method for facilitating a personalized viewing experience in connection with an object of interest included in a source video content comprising:

a first phase for identifying features specific to the object of interest using training video content including:

receiving the training video content including the object of interest, wherein the training video content is distinct from the source video content;

processing the training video content to identify one or more audio features and one or more video features, wherein the one or more audio features includes a change in a pitch or a frequency and the one or more video features includes a geometric attribute, a color, or an indication of a change in form;

extracting, from the one or more audio features and the one or more video features based upon processing the training video content, at least one audio feature or at least one video feature specific to the object of interest;

generating a list of objects including the object of interest in the source video content;

receiving a user selection of the object of interest from the list of objects of interest;

a second phase for identifying the object of interest using the source video content including:

receiving the source video content including multiple objects distributed in a plurality of frames, the multiple objects including the object of interest;

identifying, in the plurality of frames, the object of interest in response to detecting the at least one audio feature or the at least one video feature specific to the object of interest extracted during the first phase;

segmenting the source video content into multiple chunks, wherein each chunk includes at least one frame having the object of interest; and

generating a target video content by combining the multiple chunks, the target video content including sequentially-arranged frames having the object of interest.

2. The method of claim 1 , wherein the target video content is a first target video content, further comprising:

creating a second target video content with a zoomed-in or a zoomed-out view of the object of interest.

3. The method of claim 1 , wherein the source video content includes at least one of: a personal video recording, a TV show, a pre-recorded sports event, a movie, a music video, a documentary, or streaming content from a content provider.

4. The method of claim 1 , further comprising:

receiving a selection of the object of interest via at least one mechanism allowing the selection of the object of interest, the at least one mechanism including: a remote control, an online web portal, a mobile application configured to run on a mobile device, or an application program included within a voice control system.

5. The method of claim 1 , further comprising:

re-encoding the source video content for generating the target video content.

6. The method of claim 5 , wherein the source video content has a higher quality than the target video content.

7. A non-transitory computer-readable storage medium storing instructions configured for facilitating a personalized viewing experience in connection with an object of interest included in a source video content to perform a method comprising:

a first phase for identifying features specific to the object of interest using training video content including:

receive the training video content including the object of interest, wherein the training video content is distinct from the source video content;

process the training video content to identify one or more audio features and one or more video features, wherein the one or more audio features includes a change in a pitch or a frequency and the one or more video features includes a geometric attribute, a color, or an indication of a change in form;

extract, from the one or more audio features and the one or more video features based upon processing the training video content, at least one audio feature or at least one video feature specific to the object of interest;

generating a list of objects including the object of interest in the source video content;

receiving a user selection of the object of interest from the list of objects of interest;

a second phase for identifying the object of interest using the source video content including:

receive the source video content including multiple objects distributed in a plurality of frames, the multiple objects including the object of interest;

identify, in the plurality of frames, the object of interest in response to detecting the at least one audio feature or the at least one video feature specific to the object of interest extracted during the first phase;

segment the source video content into multiple chunks, wherein each chunk includes at least one frame having the object of interest; and

generate a target video content by combining the multiple chunks, the target video content including sequentially-arranged frames having the object of interest.

8. The computer-readable storage medium of claim 7 , wherein the target video content is a first target video content, wherein the method further comprises:

create a second target video content with a zoomed-in or a zoomed-out view of the object of interest.

9. The computer-readable storage medium of claim 7 , wherein the source video content includes at least one of: a personal video recording, a TV show, a pre-recorded sports event, a movie, a music video, a documentary, or streaming content from a content provider.

10. The computer-readable storage medium of claim 7 , wherein the method further comprises:

receive a selection of the object of interest via at least one mechanism allowing the selection of the object of interest, the at least one mechanism including: a remote control, an online web portal, a mobile application configured to run on a mobile device, or an application program included within a voice control system.

11. The computer-readable storage medium of claim 7 , wherein the method further comprises:

re-encode the source video content for generating the target video content.

12. The computer-readable storage medium of claim 11 , wherein the source video content has a higher quality than the target video content.

13. An apparatus for facilitating a personalized viewing experience in connection with an object of interest included in a source video content comprising:

a memory;

one or more processors electronically coupled to the memory and configured for:

a first phase for identifying features specific to the object of interest using training video content including:

receiving the training video content including the object of interest, wherein the training video content is distinct from the source video content;

processing the training video content to identify one or more audio features and one or more video features, wherein the one or more audio features includes a change in a pitch or a frequency and the one or more video features includes a geometric attribute, a color, or an indication of a change in form;

extracting, from the one or more audio features and the one or more video features based upon processing the training video content, at least one audio feature or at least one video feature specific to the object of interest;

generating a list of objects including the object of interest in the source video content;

receiving a user selection of the object of interest from the list of objects of interest;

a second phase for identifying the object of interest using the source video content including:

receiving the source video content including multiple objects distributed in a plurality of frames, the multiple objects including the object of interest;

identifying, in the plurality of frames, the object of interest in response to detecting the at least one audio feature or the at least one video feature specific to the object of interest extracted during the first phase;

segmenting the source video content into multiple chunks, wherein each chunk includes at least one frame having the object of interest; and

generating a target video content by combining the multiple chunks, the target video content including sequentially-arranged frames having the object of interest.

14. The apparatus of claim 13 , wherein the target video content is a first target video content, further comprising:

creating a second target video content with a zoomed-in or a zoomed-out view of the object of interest.

15. The apparatus of claim 13 , wherein the source video content includes at least one of: a personal video recording, a TV show, a pre-recorded sports event, a movie, a music video, a documentary, or streaming content from a content provider.

16. The apparatus of claim 13 , wherein the memory is further configured for:

receiving a selection of the object of interest via at least one mechanism allowing the selection of the object of interest, the at least one mechanism including: a remote control, an online web portal, a mobile application configured to run on a mobile device, or an application program included within a voice control system.

17. The apparatus of claim 13 , wherein the memory is further configured for:

re-encoding the source video content for generating the target video content.

Assignments (2)
SECURITY INTEREST Recorded Nov 30, 2021
From: DISH BROADCASTING CORPORATION; DISH NETWORK L.L.C.; DISH TECHNOLOGIES L.L.C.
To: U.S. BANK, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 058295/0293 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2018
From: TILAYE, SWAPNIL ANIL; SHAH, RIMA
To: DISH NETWORK, L.L.C.
Reel/Frame 044564/0014 →