IP Library Granted Patent US 12,323,651
Granted Patent B2
US 12,323,651 · App. 16/642,628 · Granted Jun 3, 2025

Tracked video zooming

Inventors: Louis Kerofsky (San Diego, CA); Eduardo Asbun (Santa Clara, CA)
Assignee: InterDigital VC Holdings, Inc.
H04N21/4316H04N5/45H04N21/44008H04N21/4728
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,323,651
App. No.
16/642,628
Granted
Jun 3, 2025
Kind
B2
Abstract

Systems, methods, and instrumentalities are disclosed for dynamic picture-in-picture (PIP) by a client. The client may reside on any device. The client may receive video content from a server, and identify an object within the video content using at least one of object recognition or metadata. The metadata may include information that indicates a location of an object within a frame of the video content. The client may receive a selection of the object by a user, and determine positional data of the object across frames of the video content using at least one of object recognition or metadata. The client may display an enlarged and time-delayed version of the object within a PIP window across the frames of the video content. Alternatively or additionally, the location of the PIP window within each frame may be fixed or may be based on the location of the object within each frame.

Claims (45)

1. A method for generating a dynamic picture-in-picture for displaying on a display device, the method comprising:

receiving video content from a server;

determining a first position of a tracked object among a plurality of objects within a first frame of the video content based on object recognition or metadata;

determining a position of a first window based on the first position of the tracked object, wherein the position of the first window is overlapping with the tracked object, wherein the first window comprises a first portion of the first frame, wherein the first portion of the first frame is visually enlarged, and wherein the tracked object, within the first portion of the first frame, is visually enlarged;

generating the first window within the first frame for displaying on the display device;

determining a second position of the tracked object among the plurality of objects within a second frame of the video content based on object recognition or metadata, wherein the second frame is temporally subsequent to the first frame, and wherein the second position of the tracked object is different than the first position of the tracked object;

determining a position of a second window based on the second position of the tracked object, wherein the position of the second window is overlapping with the tracked object, wherein the second window comprises a second portion of the second frame, wherein the second portion of the second frame is visually enlarged, and wherein the tracked object, within the second portion of the second frame, is visually enlarged; and

generating the second window within the second frame for displaying on the display device.

2. The method of claim 1 , further comprising determining that the first window or the second window is overlapping with the tracked object.

3. The method of claim 1 , further comprising:

determining a third position of the tracked object among the plurality of objects within a third frame of the video content based on object recognition or metadata; and

generating a third window in a predetermined location within a fourth frame for displaying on the display device, wherein a third portion of the third frame is within the third window, wherein the third portion of the third frame is visually enlarged, wherein the tracked object, within the third portion of the third frame is visually enlarged, and wherein the fourth frame is temporally subsequent to the third frame.

4. The method of claim 1 , wherein the first window comprises the first portion of the first frame based on a user selection of the tracked object.

5. The method of claim 1 , further comprising:

identifying the plurality of objects within an earlier frame of the video content, the plurality of objects comprising the tracked object;

generating a plurality of windows within the earlier frame for displaying on the display device, each of the plurality of windows comprising a respective tracked object among the plurality of objects, wherein each of the plurality of windows provides an indication of the respective tracked object; and

cycling through a window of focus of the plurality of windows based on user input.

6. The method of claim 5 , further comprising:

receiving a user selection of the tracked object among the plurality of objects; and

enlarging the tracked object within the first window based on the user selection.

7. The method of claim 1 , wherein metadata comprises information indicating a location of the tracked object within a frame of the video content.

8. The method of claim 1 , further comprising generating information for displaying on the display device relating to the tracked object within the first frame or the second frame.

9. A device comprising:

a processor configured to at least:

receive video content from a server;

determine a first position of a tracked object among a plurality of objects within a first frame of the video content based on object recognition or metadata;

determine a position of a first window based on the first position of the tracked object, wherein the position of the first window is overlapping with the tracked object, wherein the first window comprises a first portion of the first frame, wherein the first portion of the first frame is visually enlarged, and wherein the tracked object, within the first portion of the first frame, is visually enlarged;

generate the first window within the first frame for displaying on a display device;

determine a second position of the tracked object among the plurality of objects within a second frame of the video content based on object recognition or metadata, wherein the second frame is temporally subsequent to the first frame, and wherein the second position of the tracked object is different than the first position of the tracked object;

determine a position of a second window based on the second position of the tracked object, wherein the position of the second window is overlapping with the tracked object, wherein the second window comprises a second portion of the second frame, wherein the second portion of the second frame is visually enlarged, and wherein the tracked object, within the second portion of the second frame, is visually enlarged; and

generate the second window within the second frame for displaying on the display device.

10. The device of claim 9 , wherein the processor is configured to determine that the first window or the second window is overlapping with the tracked object.

11. The device of claim 9 , wherein the processor is configured to:

determine a third position of the tracked object among the plurality of objects within a third frame of the video content based on object recognition or metadata; and

generate a third window in a predetermined location within a fourth frame for displaying on the display device, wherein a third portion of the third frame is within the third window, wherein the third portion of the third frame is visually enlarged, wherein the tracked object, within the third portion of the third frame is visually enlarged, and wherein the fourth frame is temporally subsequent to the third frame.

12. The device of claim 9 , wherein the first window comprises the first portion of the first frame based on a user selection of the tracked object.

13. The device of claim 9 , wherein the processor is configured to:

identify the plurality of objects within an earlier frame of the video content, the plurality of objects comprising the tracked object;

generate a plurality of windows within the earlier frame for displaying on the display device, each of the plurality of windows comprising a respective tracked object among the plurality of objects, wherein each of the plurality of windows provides an indication of the respective tracked object; and

cycle through a window of focus of the plurality of windows based on user input.

14. The device of claim 13 , wherein the processor is configured to:

receive a user selection of the tracked object among the plurality of objects; and

enlarge the tracked object within the first window based on the user selection.

15. The device of claim 9 , wherein metadata comprises information indicating a location of the tracked object within a frame of the video content.

16. The device of claim 9 , wherein the processor is configured to generate information for displaying on the display device relating to the tracked object within the first frame or the second frame.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2024
From: KEROFSKY, LOUIS; ASBUN, EDUARDO
To: VID SCALE, INC.
Reel/Frame 069217/0120 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
Continuity (2)
Provisional Application 62552032 · Aug 30, 2017
Related Publication 20200351543A1 · Nov 5, 2020
References Cited (62)
US 5852474A · Nakagaki et al. · 1998 [cited by applicant]
US 8164690B2 · Fukui · 2012 [cited by examiner]
US 8331760B2 · Butcher · 2012 [cited by applicant]
US 8358907B2 · Inagaki · 2013 [cited by examiner]
US 9020259B2 · Bhagavathy · 2015 [cited by examiner]
US 9571734B2 · Kwak · 2017 [cited by examiner]
US 9648269B2 · Park · 2017 [cited by examiner]
US 9681165B1 · Gupta · 2017 [cited by examiner]
US 10019782B2 · Yeo et al. · 2018 [cited by applicant]
US 10375312B2 · Choi et al. · 2019 [cited by applicant]
US 10459532B2 · Hirata et al. · 2019 [cited by applicant]
US 20060045381A1 · Matsuo · 2006 [cited by examiner]
US 20060061602A1 · Schmouker · 2006 [cited by examiner]
US 20070191098A1 · An · 2007 [cited by applicant]
US 20090067723A1 · Yamazaki et al. · 2009 [cited by applicant]
US 20110271227A1 · Takahashi · 2011 [cited by applicant]
US 20110299832A1 · Butcher · 2011 [cited by examiner]
US 20140218611A1 · Park et al. · 2014 [cited by applicant]
US 20150062434A1 · Deng et al. · 2015 [cited by applicant]
US 20150179219A1 · Gao et al. · 2015 [cited by applicant]
US 20150268822A1 · Waggoner et al. · 2015 [cited by applicant]
US 20160057508A1 · Borcherdt · 2016 [cited by applicant]
US 20170094184A1 · Gao et al. · 2017 [cited by applicant]
US 20170147174A1 · Olejniczak · 2017 [cited by examiner]
US 20170302719A1 · Chen · 2017 [cited by examiner]
US 20180262708A1 · Lee · 2018 [cited by examiner]
CN 1758714A · 2006 [cited by applicant]
CN 1960479A · 2007 [cited by applicant]
CN 101170683A · 2008 [cited by applicant]
CN 102014248A · 2011 [cited by applicant]
CN 102106145A · 2011 [cited by applicant]
CN 103826095A · 2014 [cited by applicant]
CN 103886322A · 2014 [cited by applicant]
CN 106204653A · 2016 [cited by applicant]
CN 106384359A · 2017 [cited by applicant]
CN 106416223A · 2017 [cited by applicant]
CN 107066990A · 2017 [cited by applicant]
JP H0965225A · 1997 [cited by applicant]
JP 2006087098A · 2006 [cited by applicant]
JP 2006099404A · 2006 [cited by applicant]
JP 2009069185A · 2009 [cited by applicant]
JP 2016538601A · 2016 [cited by applicant]
JP 2017508192A · 2017 [cited by applicant]
KR 1020080010633A · 2008 [cited by applicant]
RU 2015102840A · 2016 [cited by applicant]
RU 2602778C2 · 2016 [cited by applicant]
WO 2015197815A1 · 2015 [cited by applicant]
WO 2016014537A1 · 2016 [cited by applicant]
WO 2016186254A1 · 2016 [cited by applicant]
WO 2017058665A1 · 2017 [cited by applicant]
Allen et al., “Object Tracking using CamShift Algorithm and Multiple Quantized Feature Spaces”, Proceedings of the Pan-Sydney Area Workshop on Visual Information Processing, Australian Computer Society, Inc., Jun. 2004,… [cited by applicant]
Exner et al., “Fast and Robust CAMShift Tracking”, IEEE, Computer Vision and Pattern Recognition Workshops (CVPRW), 2010, pp. 9-16. [cited by applicant]
ITU-T, “Advance Video Coding for Generic Audiovisual Services”, H.264, Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Coding of Moving Video, Nov. 2007, 564 pages. [cited by applicant]
ITU-T, “High Efficiency Video Coding”, H.265-v2, Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Coding of Moving Video, Oct. 2014, 540 pages. [cited by applicant]
Lin et al., “Outside-In: Visualizing Out-of-Sight Regions-of-Interest in a 360 Video Using Spatial Picture-in-Picture Previews”, Available at <https://www.cmlab.csie.ntu.edu.tw/˜robin/docs/uist17_lin.pdf>, Oct. 22-25, 2… [cited by applicant]
Needham et al., “Tracking Multiple Sports Players Through Occlusion, Congestion and Scale”, Proceedings of the British Machine Vision Conference, Jan. 1, 2001, pp. 93-102. [cited by applicant]
Sanchez et al., “Compressed Domain Video Processing for Tile Based Panoramic Streaming using SHVC”, Immersive Media Experiences, ACM, Brisbane, Australia, Oct. 30, 2015, 6 pages. [cited by applicant]
Software RT, “How to Overlay Photos and Videos & Create Picture-in Picture Videos using Filmora?”, Available at <https://www.softwarert.com/overlay-photos-videos-picture-in-picture-filmora/>, pp. 1-7. [cited by applicant]
Trang, “Video Overlay—Creating a Picture-in-Picture Effect”, Available at <https://atomisystems.com/tutorials/ap7/video-overlay-creating-picture-in-picture-effect/>, Mar. 9, 2018, pp. 1-5. [cited by applicant]
Wikipedia, “Mathematical Morphology”, Available at <https://en.wikipedia.org/wiki/Mathematical_morphology>, pp. 1-9. [cited by applicant]
Whitney, “LG Touts Zoom, On-Screen Remote in Upcoming Smart TVs”, Available at <https://www.cnet.com/news/lg-touts-zoom-on-screen-remote-in-upcoming-smart-tvs/>, Dec. 22, 2015, pp. 1-2. [cited by applicant]
Zebra & The NFL, “Official On-Field Player-Tracking”, Available at <https://www.zebra.com/us/en/nfl.html>, pp. 1-5. [cited by applicant]